Open the account and there are forty creatives in it. Ask which are running on purpose and nobody can answer. That's not a reporting problem — nothing in there was ever formally killed.
This guide covers the three decisions that make a review work: when you judge, what you judge on, and what happens after the verdict.
The failure mode: creatives accumulate
Nobody wants to be the person who turns a creative off. Someone made it, someone approved it, and it's still getting a little delivery. So anything with ambiguous numbers stays, and two things follow.
Spend drifts. Automated delivery pushes impressions toward whatever got early traction. Not wrong in itself, but the team can no longer tell a creative that's alive because you chose it from one that's alive because nobody switched it off. Allocation becomes a residue rather than a decision.
Decisions evaporate. Next quarter someone new relaunches a creative killed six months ago, because the only thing that survived was the file. You buy the same failure twice, at full price.
Why three states and not five
Give people a five-point scale and most things land on 3. A 3 isn't a judgment — it's a deferral wearing a number, and nothing follows from it. Three states work because three force you to pick an action.
| State | What it means | Budget | Derivatives |
|---|---|---|---|
| 🔴 Stop | This creative is done | Goes to zero | None |
| 🟡 Keep | Leave it exactly as is | Unchanged | None |
| 🟢 Scale | Push it further | Raise it | Derive from it |
What matters is what each commits you to.
🔴 Stop does not mean "underwhelming." It means budget goes to zero, and the creative is not relaunched without a new reason — a concrete change like a new season, a rebuilt landing page, a different audience. "We didn't spend enough last time" is the same bet placed again.
🟡 Keep is a decision to do nothing: no extra budget, no derivatives, no rework. It's the state people misread most. Keep isn't "it's fine, we'll get to it" — it's "the team is not spending its hours here this cycle."
🟢 Scale means you raise budget, or build derived appeals from it, or both. Mark something scale and produce nothing from it and either the verdict was wrong or the follow-up got dropped.
When to judge
The most common error is judging too early. Kill a creative after two days and you didn't evaluate the creative — you evaluated noise.
What follows is rules of thumb to calibrate against your own account, not universal thresholds. Category, price point and channel all move the right numbers a long way.
- Count data, not days. "It's been a week" is a worse trigger than "impressions have reached a level where the difference means something." A creative getting a hundred impressions a day has nothing to judge after a week.
- Be careful with single-digit conversions. The gap between three conversions and five is usually chance. Before stopping something on CPA, you want a comfortably larger count behind it than feels necessary.
- Don't judge inside the learning phase. While delivery is still redistributing impressions, the numbers describe the allocation, not the creative. Read the stretch after it settles.
- Match the weekday mix. A creative launched Friday and reviewed Tuesday is mostly weekend. Cutting on 7-day boundaries fixes the mix for free.
Then the rule that outranks every threshold: judge on a fixed cadence, not when somebody happens to look. Ad-hoc reviews get triggered by bad news, so the sample is biased before anyone opens a dashboard. Same day, same time, every creative.
What to judge on
Different creatives deserve different metrics, because the appeal each carries is tested on different ground.
| Primary appeal type | What's actually on trial | Primary metric | Supporting |
|---|---|---|---|
| Image-led | Did it stop the scroll? | CTR, thumb-stop rate | 3-second views |
| Copy-led | Does the argument survive the click? | CTR, landing engagement | Scroll depth, bounce |
| Offer / price | Did it make someone buy? | Conversion rate, CPA | Add-to-cart, ROAS |
| Trust / proof | Did it remove the hesitation? | Conversion rate | Pre-sale questions, returns |
The expensive mistake here is killing an image appeal on CPA alone.
An image appeal's job ends at the stop. If CTR is clearly above your account norm and CPA is bad, what failed is probably the landing page — the ad made a promise the page didn't keep. Mark that creative stop and you've thrown away a good asset while leaving the bottleneck in place, and the next image fails the same way.
So read it in order: CTR first, for whether the creative did its job; landing metrics second, for whether the page did its. Comparing creative performance covers getting creatives onto the same period and the same footing first.
The note is the whole point
A verdict with no note is nearly worthless six months later. All a bare 🔴 tells the next person is that somebody, once, disagreed with this creative — so they test it again. Two sentences is enough: what you concluded, and what you'd try next. Longer doesn't get written.
| Date | Creative | Verdict | Note |
|---|---|---|---|
| 2026-08-04 | Nacre close-up A | 🟢 Scale | CTR 1.8× the worn-shot version, so the close-up itself is doing the work. Two derivatives with different light angles. |
| 2026-08-11 | Three-year warranty banner | 🔴 Stop | CTR and conversion both below account average — warranty may be a weak appeal in this category. One more copy variant before retiring the appeal itself. |
| 2026-08-18 | Price comparison B | 🟡 Keep | Conversion rate average, CPA stable. Nothing wrong with it, no reason to push it. Revisit at the next promo. |
Each note ends with a next action. "Underperformed" is not a note. Give each creative a stable ID and the note travels with every derivative made from it — tracking creatives by ID covers how.
What to do after each verdict
🔴 Stop. Zero the budget, write the reason, then ask: did the appeal fail, or the execution? Only when the same appeal has failed across two or three genuinely different images or copy angles do you have grounds to retire the appeal itself — what an ad appeal is explains why they're separate layers needing separate verdicts.
🟡 Keep. Do nothing, on purpose. A keep verdict converts "we should probably look at this" into "we've decided not to." It comes back next cycle.
🟢 Scale. Raise the budget, build derivatives, watch for fatigue. If a winner's CTR slides week over week while frequency climbs, the creative is tiring and you want the replacement already built. Derive while it's still winning; start after the decline and you're producing under pressure with no runway.
A weekly 20-minute review
A process only survives attached to a meeting. Twenty minutes is enough.
| Time | What happens |
|---|---|
| 0–5 min | Last 7 days on screen, per creative. Pulled beforehand, never during |
| 5–13 min | Mark verdicts. 30 seconds each, hard cap. Anything longer is a keep |
| 13–17 min | List what gets derived from the scale verdicts, with owner and date |
| 17–20 min | Walk the stop list, confirm the budgets are actually at zero |
Two rules keep it inside the box. If a verdict is contested, mark it keep and revisit next week — the argument costs more than the week does. And when the meeting ends, every creative carries one of the three states. Leaving nothing unjudged is the only real output.
Keeping the record
A spreadsheet is a fine place to start. It stops working when the verdict lives somewhere other than the creative, the appeal and the numbers — nobody opens it.
copy hound puts exactly this on each creative: a stop / keep / scale dropdown with a note field, saved as a dated log entry rather than overwriting the last. The most recent verdict shows on the creative and older entries stay collapsed underneath, so the whole history — stopped in March, relaunched in June after the page rebuild, scaled in August — reads on one screen. And because verdicts are filterable, "every creative currently marked scale" is one click when you're choosing what to derive.
FAQ
How many days should a creative run before I judge it?
Don't set it in days. Set it on whether impressions and conversions have accumulated to a level where the difference between two creatives means anything. Cutting on 7-day boundaries is still a good habit — it keeps the weekday mix comparable. Low-delivery creatives may need two or three weeks.
Can I change a verdict later?
Yes — a verdict is a conclusion at a point in time, not a permanent label. What matters is that changing it doesn't erase the previous entry. If something was stopped in August and relaunched in October, both need to survive: the useful information is what changed in between.
What if most of my creatives end up in keep?
If keep is more than about half your library, the bar is too soft. Put a ceiling on how long something can sit there — after three consecutive keep verdicts it has to move to stop or scale. Three cycles without finding a reason to push it is itself a conclusion.
The creative is weak but I still believe in the appeal. What then?
Stop the creative, keep the appeal — separate layers, separate verdicts. Test the same appeal with two or three genuinely different executions before retiring it, and write "appeal holds, execution replaced" in the note so the next person doesn't read the stop as a verdict on the idea.
Read next
How to compare ad creative performance fairly
Two creatives almost never run under the same conditions, so putting their numbers side by side usually produces a wrong answer. Here are the five conditions that have to match, why rates beat totals until the denominator gets small, and how to roll creative numbers up to the appeal.
Derived appeals: keeping an improvement trail on your creatives
When an ad appeal wins, the useful next move is neither running it to death nor starting over. Here is how derived appeals work, the one-variable rule that keeps them readable, and why the resulting tree is worth more than any single creative in it.
Somewhere to actually run this
copy hound builds your appeals up as assets, keeps an improvement trail through derived appeals, and rolls per-channel creative performance back up to the appeal. Runs in the browser — nothing to install.
Start free