Two creatives land next to each other in the weekly report. A is at a $60 CPA, B is at $50. Somebody says "let's push B," everyone nods, budget moves. But A ran two weeks across a holiday weekend on $12,000, and B ran four days mid-month on $800. B has sixteen conversions. That comparison told you nothing, and a budget decision just came out of it.
This guide covers why most creative comparisons are wrong, what has to match before one means anything, and what to do about the differences you can't remove.
Why creative comparison is hard
The problem is not the metric. It's the conditions. Two creatives rarely run under identical conditions — one launched earlier, one got more budget, one skewed to a different placement. Compare their numbers anyway and you're reading the conditions, not the creative.
| Creative A | Creative B | |
|---|---|---|
| Days live | 14 (holiday weekend) | 4 (mid-month) |
| Spend | $12,000 | $800 |
| Conversions | 200 | 16 |
| CPA | $60 | $50 |
To claim B won you have to claim three things at once: the holiday weekend didn't move CPA, B holds its efficiency at fifteen times the spend, and sixteen conversions is a stable reading. Nothing supports any of them.
The error also hides itself. Move budget to B, it starts running under A's conditions, and its CPA drifts up toward A's. The team then says the creative got tired — not that it was never a $50 creative. Bad comparisons get written off as fatigue.
The five things that have to match
If any of these rows is badly off, what you have is an observation, not a comparison.
| Condition | What goes wrong when it doesn't match |
|---|---|
| Length and calendar position | Payday, a holiday or a sale shifts the conversion baseline. Different lengths also make learning a different share of each creative's data |
| Budget scale | Efficiency degrades as spend rises. A low-spend creative's good CPA usually means it has only eaten the easy audience |
| Audience and placement | Retargeting and prospecting differ several times over on CVR. A different placement mix makes CTR comparison meaningless |
| Funnel stage the metric measures | CTR measures the hook, CVR persuasion after the click, CPA both plus price. Mix stages and the winner flips |
| Product being sold | CPA isn't comparable across order values. A $40 CPA on a $120 product is worse than a $180 CPA on a $900 one |
All five matching is rare. The goal isn't a perfect comparison — it's one where you know which rows are off.
Budget scale is the row teams skip most. Delivery systems front-load the people most likely to respond, so raising spend means working past the easy audience. A small new creative beating your main one on CPA is close to a weekly event, and most of the time it isn't the creative. Setting conditions up front is what A/B testing ad copy is for.
Rate metrics vs. total metrics
When spend differs, totals — clicks, conversions, revenue — are a budget ranking. A creative on $12,000 gets more clicks than one on $800. That's arithmetic, not evidence. So you compare rates, which divide the scale back out.
| Metric | Creative A | Creative B |
|---|---|---|
| Impressions | 1,600,000 | 90,000 |
| Clicks | 20,800 | 1,530 |
| CTR | 1.30% | 1.70% |
| Conversions | 200 | 16 |
| CVR | 0.96% | 1.05% |
| Spend | $12,000 | $800 |
| CPA | $60 | $50 |
But rates mislead in their own way: a small denominator makes a rate unstable.
A third creative spent $450 for 3 conversions — a $150 CPA. One conversion fewer and it's $225; one more and it's $112.50. The number moves by half its own value on a single order, yet "$150 CPA" looks as solid on a dashboard as "$60 CPA on 200 conversions."
CTR stabilises early because impressions arrive by the tens of thousands per day; CVR and CPA stabilise late because conversions arrive a handful per day. So when CTR says A and CPA says B, trust the metric with the bigger denominator first.
Totals still matter for scale — the account's best CPA on three conversions cannot fill a monthly target. Rank on rates, then use totals to check the ranking is happening at a size worth acting on.
How many conversions is enough depends on your category, order value and conversion definition, so don't import a threshold from a blog post. Rank last quarter's creatives as they stood at day 3, day 5 and day 7, and find where the ranking first matched the final one. That's your baseline.
Normalise what you can, flag what you can't
Some of the mismatch can be corrected after the fact.
- Trim to a common window. Keep only the days both were live. You lose volume, and most of the seasonality with it.
- Convert to per-1,000-impression figures. Same unit regardless of budget.
- Exclude the learning period. The first days reflect the algorithm exploring, not the creative performing. Cut what the platform marks as learning, or the first two to three days.
- Divide by order value. Different products, use ROAS instead of raw CPA.
| Per 1,000 impressions | Creative A | Creative B |
|---|---|---|
| Clicks | 13.0 | 17.0 |
| Spend | $7.50 | $8.89 |
| Conversions | 0.125 | 0.178 |
B earns more clicks and conversions per impression, but its impressions cost more to buy. Strip the scale out and B still doesn't look bad — that is the entire claim this table supports. Sixteen conversions are still sixteen conversions.
Then the rule that matters most: differences you can't normalise away should be marked, not smoothed over.
A report that always ranks creatives first to last is comfortable, and the price is that a broken comparison and a sound one look identical on screen. A tool that refuses to declare a winner when conditions differ is more useful than one that always shows a leaderboard. "Not comparable — different windows and budgets" isn't a missing answer; it's the absence of a wrong one.
In practice, put conditions next to numbers. "B — $50 CPA (4 days / $800 / 16 conversions)" gets discounted correctly by whoever reads it. "B — $50 CPA" does not. Three extra columns is the whole cost.
Rolling creative numbers up to the appeal
One creative's number is noisy by nature — a thumbnail, a line of copy and some luck in early delivery move it around. But the question you care about is usually the idea inside it, and "does leading with material and detail sell?" is not something one creative can answer. Six creatives carrying the same appeal across three channels is different: individually they wobble, together they read the idea.
| Creative | Channel | Spend | Impressions | Clicks | CTR | Conv. | CPA |
|---|---|---|---|---|---|---|---|
| 1 | Meta | $7,200 | 900,000 | 11,700 | 1.30% | 120 | $60.00 |
| 2 | Meta | $1,600 | 220,000 | 2,640 | 1.20% | 25 | $64.00 |
| 3 | TikTok | $1,200 | 200,000 | 2,200 | 1.10% | 15 | $80.00 |
| 4 | TikTok | $600 | 100,000 | 900 | 0.90% | 6 | $100.00 |
| 5 | Google Display | $900 | 300,000 | 2,100 | 0.70% | 8 | $112.50 |
| 6 | Google Display | $500 | 180,000 | 1,440 | 0.80% | 4 | $125.00 |
| Appeal total | $12,000 | 1,900,000 | 20,980 | 1.10% | 178 | $67.42 |
A $67.42 CPA on 178 conversions looks far steadier than any single row. Here is the trap: creative 1 is 60% of the spend and 67% of the conversions. The rollup isn't a reading of the appeal, it's a reading of creative 1 rounded off by five small ones. If creative 1 worked because of its photograph rather than its argument, you're about to judge an idea by a picture.
Check the spread before you trust the total.
- What share of spend sits in the largest creative? Over half, and the total is close to that creative's own result.
- Do the channels point the same way? Meta runs 1.2–1.3%, TikTok 0.9–1.1%, Display 0.7–0.8%. The gaps between channels exceed the gaps between creatives, so this table is partly reading placement.
- Does the conclusion survive dropping a creative? Without creative 1, the other five come to $4,800 over 58 conversions — an $82.76 CPA. Same direction, different level.
The rollup only exists if every creative carries a tag for the appeal it came from. Tracking creatives by ID covers keeping that link intact; what an ad appeal is covers the layer separation it depends on.
Cumulative vs. per-period
Cumulative numbers tell you how efficient a creative has been over its life. Per-period numbers tell you what it's doing now. They often disagree.
| Week | Spend | Impressions | Clicks | CTR | Conv. | CPA |
|---|---|---|---|---|---|---|
| 1 | $1,600 | 200,000 | 3,200 | 1.60% | 32 | $50.00 |
| 2 | $1,600 | 210,000 | 3,150 | 1.50% | 32 | $50.00 |
| 3 | $1,600 | 220,000 | 2,860 | 1.30% | 25 | $64.00 |
| 4 | $1,600 | 230,000 | 2,760 | 1.20% | 20 | $80.00 |
| 5 | $1,600 | 250,000 | 2,250 | 0.90% | 16 | $100.00 |
| 6 | $1,600 | 260,000 | 1,820 | 0.70% | 10 | $160.00 |
| Cumulative | $9,600 | 1,370,000 | 16,040 | 1.17% | 135 | $71.11 |
Cumulatively this creative is unremarkable and fine — 1.17% CTR, $71.11 CPA, probably near account average. Nothing asks for attention.
Week by week it's a different creative. CTR falls in one direction from 1.60% to 0.70%, and the last two weeks spent $3,200 for 26 conversions between them — a $123.08 CPA. A good start is hiding a bad present. Impressions rise every week while clicks fall, the standard shape of a creative pushed into audiences that don't respond to it.
It runs the other way too: a creative that spiked during learning and then settled looks worse cumulatively than it now is, and killing it on the lifetime number kills a working creative. Rank on cumulative, read direction on per-period — then write the decision down with a date. The creative review workflow covers that log.
A comparison checklist
Run this before declaring any winner.
- Did both creatives run over the same date range? If not, have you trimmed to the overlap?
- Does that range contain a payday, holiday or promotion only one of them caught?
- Have you excluded the learning period for both?
- Is spend within a reasonable multiple? If not, have you converted to per-1,000-impression figures?
- Are the audience and placement mix the same?
- Are you comparing metrics from the same funnel stage?
- Is the order value the same? If not, are you using ROAS rather than CPA?
- How many conversions sit under each side? Would one order flip the ranking?
- For an appeal rollup, does one creative hold more than half the spend?
- Have you written down which of the above failed?
Ten is the one that gets dropped. If any row failed, the honest sentence is not "B wins" but "B is ahead, but the windows differ, so this isn't settled." A report that states unsettled things as settled costs more than one that admits what it doesn't know.
What this looks like in a tool
copy hound records performance per creative, per period. Six weeks of one creative stay as six rows instead of collapsing into a lifetime average, and creatives from the same appeal roll up to that appeal's line. Optional creative IDs let you follow a creative and its derivatives across channels as one lineage.
The comparison view doesn't produce a leaderboard. It shows the window, the spend and the conversion count next to each number, so a comparison built on four days and sixteen conversions doesn't look like one built on two months. The judgement stays with you, and you log it as stop, keep or scale into a dated history. There is no ad-platform sync — numbers go in by hand. What you get for that is numbers on the same line as the appeal, the copy and the creative that produced them. Ad platforms remember your creatives; they don't remember what idea each was testing.
FAQ
Which metric should I compare on?
Match the metric to what the creative is responsible for: CTR for thumbnails and hooks, CVR or CPA for whether the argument survives the click. In practice, shortlist on CTR and decide on CPA — scaling on CTR alone gets you creatives that collect clicks and don't sell. If the creatives push products at different price points, use ROAS; CPA silently assumes every conversion is worth the same.
How long should a creative run before I compare it?
Judge by denominator, not calendar. Three days or ten, what matters is how many conversions have landed — a CPA built on 3 of them swings by half its value on one order. Look back at previous creatives in your own account and find where their ranking stopped changing. That's your minimum, and it shifts with category, order value and how far down-funnel your conversion event sits, so don't borrow someone else's number.
A creative has great CTR but bad CPA. What now?
Usually the ad and the landing page promise different things — the creative sets an expectation the page doesn't pick up. Before retiring it, align the top of the landing page with the same appeal and re-measure. If CPA still won't come down, it's a creative problem. The reverse case, low CTR and good CPA, often deserves more impressions rather than fewer.
Can I compare creatives across different channels?
Not on absolute values. A 1.2% CTR on Meta and 0.8% on Display are drawn against different baselines, so the gap tells you about placement, not creative. Compare relative rank within each channel instead. An appeal in the top group on Meta and on TikTok is a strong appeal; top on Meta only is a good placement fit, which is a much smaller claim.
Read next
Tracking ad performance by creative ID
The platform's ad ID and the name your team says out loud never meet, so the monthly report gets rebuilt from memory. Here is a naming convention that survives contact with reality, where to store the ID, and how to link derived creatives so one idea reads as one thing.
How to design an A/B test for ad copy
Most of what gets called an A/B test in an ad account is a ranking, not an experiment. Here is how to decide what you're testing, how much data you actually need, and the seven failure modes that quietly break ad tests.
Somewhere to actually run this
copy hound builds your appeals up as assets, keeps an improvement trail through derived appeals, and rolls per-channel creative performance back up to the appeal. Runs in the browser — nothing to install.
Start free