copy hound Guides Free tools
KOEN
Open the app · free
copy houndGuides

How to design an A/B test for ad copy

Most of what gets called an A/B test in an ad account is a ranking, not an experiment. Here is how to decide what you're testing, how much data you actually need, and the seven failure modes that quietly break ad tests.

Updated 2026-08-25 · 8 min read · ad copy a/b test · ab testing ads · creative testing · ad experiment design

"We A/B tested it and B won." You hear it constantly. Ask why B won and the room goes quiet.

This guide is about designing ad copy tests that leave behind a reason, not just a ranking — no statistics lecture, no invented confidence.

What most "A/B tests" in ad accounts actually are

The usual shape: two creatives go live the same day. A leads with a discount over a product shot, on an interest audience. B leads with a use-case line over a lifestyle shot, on a lookalike. Four days later B's CTR is higher, so A is switched off and B is the winner.

You now have a ranking. What you don't have is a reason — was it the copy, the image, or the audience? Three things changed at once.

A ranking is worthless for the next build. Lifestyle shots from now on? Always lookalikes? Nobody can say.

The point of a test isn't choosing which creative to pause this week — that's just running the account. The point is knowing what to build next.

Decide what you're testing before you write anything

The first question isn't "how should A and B differ?" but "which layer should they differ on?" — the same three layers as in what an ad appeal is.

Layer under testThe question it asksSize of differenceBudget needed
AppealDoes this reason persuade at all?LargeLarge
Expression (copy)Does this wording say the reason better?ModerateModerate
Creative (image)Does this image carry it better?ModerateModerate

An appeal test asks "price or craftsmanship?" The variants tell genuinely different stories, so the gap can be wide — easier to read on a small budget. The cost is up front: every appeal needs its own copy and image.

An expression test asks "is it 'lightweight' or 'you stop noticing it in your bag'?" The reason is identical, so the gap is narrow, and narrow gaps need a lot of data.

The common mistake is running an expression test while believing you ran an appeal test. "Emotional appeals don't work for us" is usually the much smaller finding that one emotional line didn't work. Judging an appeal from a single piece of copy is nearly always wrong.

The one-variable rule, and where it's genuinely impractical

Change one thing. Testing copy means image, audience, budget, flight window, placement and landing page all stay fixed.

Being honest: on a small account, strict single-variable testing is often unaffordable. Hold everything constant and your variants differ so little you can't afford to detect the difference. Add variants and the per-variant budget falls below the readable level.

The workable compromise:

  • When you can't hold every element constant, hold the layer constant. In a copy test the images needn't be identical, but they must be the same kind of image — no product shot against a lifestyle shot.
  • Never compare across layers. Don't rank A and B (copy differs) alongside C (appeal differs). If C wins, that is not a result from your copy test.
  • Write down what you couldn't control. One line — "B's crop was tighter" — gives you somewhere to look when the result won't reproduce.

Narrowing what you're willing to conclude beats pretending to perfect control.

Sample size, without the maths lecture

You don't need a formula, just one instinct: small differences need large numbers, and the requirement grows much faster than the difference shrinks. Roughly, halve the difference you want to detect and you need about four times the data. That is why expression tests cost more than appeal tests.

The table below calibrates scale. It is not a formula and not a threshold to copy.

Relative difference you hope to detectRoughly, impressions per variant
50% (1.0% → 1.5%)thousands
30% (1.0% → 1.3%)tens of thousands
20% (1.0% → 1.2%)tens of thousands
10% (1.0% → 1.1%)over a hundred thousand
5% (1.0% → 1.05%)several hundred thousand

Those orders of magnitude assume a 1% baseline CTR and the conventional error tolerances most calculators default to — neither is a law of nature. Change the baseline and the whole table moves: at a 0.3% baseline every row grows substantially, and for conversion rate what you must accumulate is conversions, not impressions. Take the spacing from this table, not the numbers.

If you can't afford enough data to detect a 10% difference, don't design a test that hinges on a 10% difference. Test something that will move more than that — which usually means testing the appeal rather than the wording.

The smaller your budget, the bolder your test has to be.

The mistakes that break ad tests specifically

These don't appear in generic A/B testing advice — they're artefacts of how ad platforms work.

1. Judging inside the learning period. Platforms spend the opening days exploring delivery, so CPA and CTR in that window are exploration costs, not performance. Call a winner early and you've measured which creative stabilised first.

2. Unequal delivery. The most commonly missed one. The optimiser notices an early lead and pours budget into it, so by day four A has 90% of the spend and B has 10%. That isn't A versus B; it's a big A versus a small B. Force the allocation with separate campaigns or budgets, and at close check spend per variant. If one spent more than roughly twice the other, don't read the result.

3. Different flight windows. A ran Thursday to Sunday, B ran Monday to Wednesday. You may have measured the weekend, not the copy. Every variant should start and stop at the same moment and cover at least one full cycle — normally seven days.

4. Audience overlap. Point two similar audiences at two variants and the same people see both. Whichever they see second is helped or hurt by the first, and the variants cannibalise each other. Use split-test tooling or mutually exclusive audiences.

5. Peeking. Statistically the most damaging habit. Stopping the moment one variant is ahead badly inflates your chance of crowning a random fluctuation. Fix the stopping point before launch, then look but don't touch.

6. Judging on a metric that doesn't test the hypothesis. If the hypothesis is "this line gets read," CPA is the wrong judge — when the landing page is the bottleneck, every piece of copy returns a similarly bad CPA and you conclude copy doesn't matter. Attention → CTR. Persuasion → landing conversion rate. Audience fit → CPA.

7. Winner's curse. The winner of a small test won on some real effect plus some luck. The luck doesn't scale, so performance settles below what the test showed. That's expected, not failure. Judge further scaling on the numbers measured after scaling, not the original margin.

A test you can actually run on a small budget

A shape that satisfies all of the above on a modest budget: 3 appeals × 1 creative each, equal budget, fixed window. Appeals rather than wording, for the reason above — a small budget can only resolve large differences.

Then, before launch, fill in this table and save it somewhere durable. Writing it down is what prevents mistake 6 (picking a flattering metric afterwards) and mistake 5 (stopping early).

FieldEntry
HypothesisThis product sells on material detail rather than on price
VariantsA price · value / B material · detail / C occasion · use
Held constantAudience, placement, landing page, image style, budget (one third each)
MetricCTR — the hypothesis is about attention
Minimum data30,000 impressions per variant, minimum 7 days
Decision ruleAdopt the top variant only if it leads the bottom by 30% or more. Anything narrower is recorded as "no difference"
Validity checkSpend per variant within 20% of each other at close

The decision rule is the load-bearing row. Without "30% or more" set in advance, a 12% gap gets called a win.

Reading the result honestly

Every finished test lands in one of three states.

1. A clear winner. The gap cleared your pre-registered margin and spend was even. Adopt it — but don't multiply the budget by five overnight; scale in steps and re-measure. Then push the winning appeal with derived appeals to find its ceiling.

2. No difference. All three landed within noise of each other. Most teams bin this — it's a real finding. This category doesn't respond to which reason you give, so the next budget belongs to another layer: image, offer, landing page. Record it, or someone tests the same three appeals in six months.

3. Inconclusive because the test broke. Spend split 8:2, one variant caught the weekend, someone edited a budget mid-flight. The only honest conclusion is "we don't know." The pull to pick a winner anyway is strong — but that winner is the one that reverses next month. File it as a design note, fix the flaw, run it again.

Keeping the record

The value of a test shows up six months later, in a meeting. But results kept in the ad account vanish with the creatives, results in chat aren't searchable, and results in a spreadsheet drift away from what they were about.

The minimum unit worth keeping is what was tested + what changed + the numbers for that period + the verdict, and those four have to stay on one line.

copy hound keeps each test as a derived appeal under its root, with a change note for what was varied. Underneath sit the per-channel copy, creative images, performance by period, and the stop/keep/scale verdict with its date — so the result stays attached to the thing tested, not to a campaign name. See comparing creative performance for cutting creatives by period, and the creative review workflow for who signs off.

FAQ

How many days should a test run?

Decide on data, not dates — but as a floor, clear both one full cycle (usually seven days) and the platform's learning period. Four days is almost always too short; seven days with too few impressions isn't finished either. Two weeks still short of target volume means the test was chasing a difference too small for the budget.

How many variants can I run?

Your budget decides. More variants means less data each, and it raises the odds one leads by chance. On a small account, start with two or three — five in two rounds beats five at once and reading none of them.

Should I use the platform's built-in A/B test feature?

Generally yes — it handles delivery splitting and audience de-duplication, covering most of mistakes 2 and 4. But the "winner" it declares follows its own criteria, so check the metric matches your hypothesis. It doesn't remove the learning period or winner's curse.

The result wasn't significant. Did the test fail?

No — but separate two things. "There is no difference" and "we didn't gather enough data to see a difference" are different outcomes. The first is information: this layer isn't where the leverage is, so move to another. The second means redesigning — fewer variants, or a bigger swing. Which you have depends on whether you hit the minimum data you wrote down before launch.

Somewhere to actually run this

copy hound builds your appeals up as assets, keeps an improvement trail through derived appeals, and rolls per-channel creative performance back up to the appeal. Runs in the browser — nothing to install.

Start free