PricingOpen the demo

Creative testing that proves something

“Is this creative actually better, or did I get lucky?”

8 min readUpdated August 2026No signup required

Two ads run for four days. One gets nine purchases, the other gets six. The nine wins, the six is paused, and a conclusion enters the company's folklore. With those volumes, that gap is well inside what pure chance produces — you would see it regularly from two identical ads.

The sample size problem, in plain terms

Conversion is a rare event. The rarer the event, the more of it you need before a difference means anything. The intuition worth carrying: to reliably detect a genuine 20% improvement in conversion rate, you need hundreds of conversions per variant — not dozens.

Conversions per variantWhat you can detectWhat it is good for
Under 30Almost nothingChecking the ad renders and the link works
30–100Only enormous differences (roughly 2×)Killing obvious failures, not picking winners
100–300Large differences (~30–50%)Deciding between genuinely different concepts
300+Moderate differences (~20%)Iterating on a working concept
Approximate and deliberately so — the precise threshold depends on your baseline rate, but the order of magnitude is what people get wrong.
How much could you even see?

Only enormous differences are visible. With 40 conversions per side, a difference smaller than roughly 44% is indistinguishable from chance.

Indicative, not a power calculation — the exact figure depends on your baseline rate

Change one thing

A test between a new hook, a new format and a new offer tells you that the bundle won. It does not tell you which part to keep, so nothing is learnt and the next test starts from zero. Vary one element and you build knowledge that compounds.

  • Hook — the first two seconds, or the headline. Usually the largest single lever.
  • Format — static against video, aspect ratio, whether it looks native to the feed.
  • Offer — what is actually promised. Changes economics as well as performance, so test it separately.
  • Proof — reviews, demonstration, before-and-after.
  • Call to action — reliably the smallest effect of the five, and the most tested.

Why the platform's own test tool flatters winners

Platform split tests optimise delivery inside each variant while comparing them. The winner is partly the ad and partly the audience the algorithm found for it — which is genuinely useful information, but it is not a clean read on the creative. Expect the measured gap to be larger than the gap you will see when the loser is retired and the winner has to carry the whole budget.

Judge on profit, not on the leaderboard

Creative profit efficiency

(orders × contribution margin per order) ÷ spend on that creative

Creative A: 120 orders at DKK 480 margin on DKK 42,000 spend = 1.37. Creative B: 95 orders at DKK 610 margin on DKK 38,000 = 1.53. A won on volume and ROAS; B won on money.

This happens whenever creatives sell different products. A discount-led ad wins on order count and loses on margin, and any leaderboard sorted by ROAS will promote it. Sort by profit and the ranking often inverts.

A protocol you can actually follow

  1. 1Write the hypothesis and the expected direction down first. “A shorter hook will raise hold rate and cost per order will fall.”
  2. 2Decide the decision threshold before launch — what result makes you keep it, kill it, or run it again.
  3. 3Change one element. Keep audience, placement, budget and bidding identical.
  4. 4Run for at least one full week, and past the learning phase. Never conclude on a weekend alone.
  5. 5Compare on profit efficiency, not ROAS and not clicks.
  6. 6Re-run the winner against the incumbent once more before believing it.

In short

  • Under about 100 conversions per variant, you are reading noise.
  • Change one element per test, or you learn nothing reusable.
  • Platform split tests overstate the gap; expect the winner to regress.
  • Rank creatives by profit per krone spent, not by ROAS or volume.
  • Write down the decision threshold before launch. That single habit does most of the work.

Where this method runs out

Everything above works in a spreadsheet. Keeping it current, and matching every order back to the ad that actually caused it, is the part that does not. That is what Kepra does — and the demo runs on sample data with no signup, so you can judge it before believing any of this.

Open the demo →

Read next