Blog

Why published hit rates run from 2% to 33%, and how many ads it takes to find one winner.

Each figure counts a different kind of win, and only one of them is about ads. Then the arithmetic for your own account.

Cover: How many ads to find one winner, beside three finished ads in different formats.

Search for a creative hit rate and you get answers that refuse to agree. In AppsFlyer's 2024 report, 2% of ad variations took 68% of spend. Optimizely found 12% of web experiments produced a significant lift, and Microsoft put its experiment success rate at 33% (Kohavi, Deng and Vermeer, 2022).

If your own number sits somewhere else, the account probably isn't broken. These figures disagree because each one counts a different thing as a win, and only the first is about ads. The more useful question is how many ads you need to launch for a fair chance at one winner.

What is a creative hit rate?

Short answer

A creative hit rate is the share of ads launched that meet a stated win condition, and published hit rates disagree because each source counts a different thing as a win: a share of spend, a CPA held at target, or a statistically significant lift.

The arithmetic is winners divided by ads launched. The hard part is the top of that fraction, because Ads Manager doesn't decide for you what a winner is.

Each win condition gives a different number for the same account. A spend-share definition that counts the top slice of ads as winners reports a hit rate close to the size of that slice, whatever you launch. A CPA definition moves with your target: tighten it and the rate falls. A significance definition depends on traffic, so small budgets produce fewer winners even when the ads are good.

So the hit rate is a choice before it's a result. Whether one winner is real is a separate check: see what statistical significance means for one test.

What counts as a win, and what rate does each definition give?

Here are the published figures people quote, lined up by what each one counted as a win. Only the AppsFlyer rows are about ads, and those are mobile app installs, not DTC brands. The other rows come from website and product experiments. They're included because ad testers quote them as if they were ad benchmarks.

Published hit rates by win definition. Sources: AppsFlyer press releases (15 May 2024, 29 Apr 2025), Optimizely (127,000 experiments), Kohavi, Deng and Vermeer (KDD 2022). None is a DTC Meta benchmark.
What counts as a winPublished rateSourceWhat it measures
Top slice of spend2% of ad variations ("one in fifty")AppsFlyer, 15 May 2024Where budget went across 220,000 creative variations from over 2,000 apps
Top slice of spend, by categoryTop 2% of creatives took 53% of spend in Gaming and 43% in Non-GamingAppsFlyer, 29 Apr 2025Spend concentration across 1.1 million video creatives from 1,300 apps
Significant lift on the primary metric12% of experimentsOptimizely, 127,000 experimentsWebsite and feature tests, not ads
Company-reported experiment success8% (Airbnb Search) to 33% (Microsoft)Kohavi, Deng and Vermeer, KDD 2022, Table 2Each company's own accounting of product and website experiments, not ads

How to read it: a published rate only means something next to its definition. AppsFlyer measured spend concentration in mobile app marketing. Optimizely counted experiments that beat control on their primary metric. Kohavi and his co-authors collected what each company reported as a success. None of these is a DTC Meta benchmark, and we didn't find a published figure that is.

According to Kohavi, Deng and Vermeer (KDD 2022), published experiment success rates range from 8% at Airbnb Search to 33% at Microsoft, and the authors call them ballpark figures because companies count success differently.

Their own caveat: "These numbers may involve different accounting schemes, and we never know the true rates, but they suffice as ballpark estimates." A single figure lifted from that table can't tell you whether your account is on track.

Why a top-slice-of-spend definition sets its own hit rate

When a winner is defined as the top slice of spend, the hit rate is fixed by the definition: AppsFlyer's 2024 report found that 2% of ad variations took 68% of spend, which it described as one in fifty ads being a winner.

AppsFlyer's press release describes the dataset as "an anonymous aggregate of proprietary global data from over 2,000 apps, looking at 220,000 creative variations, and 720 million marketing-driven installs." The full report is gated, so the definition here comes from the release. Ad platforms sent budget to the ads that engaged best, and the ads at the top of that spend curve counted as winners.

Here's the catch. Under a relative cut-off, winners aren't independent draws. If the top 2% by spend are the winners, 100 ads produce about two and 1,000 ads produce about twenty (illustrative). The count of winners moves with volume. The rate stays pinned near the cut-off, because the cut-off is the rate.

So 2% describes where budget went across 220,000 app-install variations. It isn't a prediction for your next ten ads, and it says nothing about whether the ads that got the budget were profitable. How many new ads to launch each week at your spend level is a different question with its own answer.

One honest note about this post's title. The 33% end of the range comes from Microsoft's product experiments, not ads, and the 2% end comes from app-install ads. "2% to 33%" is the spread of published definitions of a win. It isn't the spread of ad results you should expect.

What the experiment research says about winners

The careful work on win rates comes from website and product experiments, but it answers a question every ad tester has: how often does a tested change do anything at all?

Rarely. According to a 2022 study by Berman and Van den Bulte in Management Science, which analysed 4,964 effects from 2,766 experiments, about 70% of tested changes had no true effect. Most ideas don't move the metric.

That sounds bleak until you read a second paper. Azevedo and colleagues, in a 2020 study in the Journal of Political Economy built on Microsoft Bing's experimentation platform, found that idea quality is fat-tailed. Most ideas do little and a rare few do a lot. Under that distribution, their analysis favours a "lean" approach: test more ideas with smaller samples rather than a few big experiments. The research suggests volume matters when most ideas fail, because the rare large effect is what moves the result. It doesn't say any particular number of tests will turn one up.

Optimizely, reporting on 127,000 experiments, offers the sanity check from the other side: "If nine out of ten of your tests are winning, something is off." A very high hit rate is more often a loose definition than a great account.

How many ads does it take to find at least one winner?

Once you've fixed a definition and have a rough hit rate, the planning question has an answer. The binomial distribution, as set out in the NIST/SEMATECH e-Handbook of Statistical Methods, gives it in one line:

Chance of at least one winner = 1 - (1 - hit rate) raised to the power of the number of ads.

At a 10% hit rate and 10 ads, that's 1 - 0.9^10, about 65% (illustrative). Both tables below are CreatStrat's own arithmetic from that formula, rounded. The rows in both tables are assumed hit rates, not sourced ones. The 5% row, for example, is an assumption and doesn't come from any study.

Chance of at least one winner. Illustrative arithmetic under assumed hit rates. Not a forecast, not a benchmark.
Assumed hit rateAds launched in the test window
510204080
2%9.6%18.3%33.2%55.4%80.1%
5%22.6%40.1%64.2%87.1%98.3%
10%41.0%65.1%87.8%98.5%>99.9%
15%55.6%80.3%96.1%99.8%>99.9%
30%83.2%97.2%99.9%>99.9%>99.9%
Illustrative arithmetic
More ads raise the chance of one winner, by less each time.
Chance of at least one winner on a 0 to 100% scale, for each count of ads launched in the test window, at three assumed hit rates.
Solid bar: 2% hit rateTick: 5% hit rateOutlined band ends at: 10% hit rate
5 adsIn the test window
9.6% / 22.6% / 41.0%
10 adsIn the test window
18.3% / 40.1% / 65.1%
20 adsIn the test window
33.2% / 64.2% / 87.8%
40 adsIn the test window
55.4% / 87.1% / 98.5%
80 adsIn the test window
80.1% / 98.3% / >99.9%
0%25%50%75%100%
Chance of at least one winner. Values read 2% / 5% / 10% hit rate

Illustrative arithmetic under assumed hit rates. Not a forecast, not a benchmark.

Illustrative example. CreatStrat's arithmetic: chance = 1 − (1 − hit rate)ads, the binomial model in the NIST/SEMATECH e-Handbook of Statistical Methods, section 1.3.6.6.18. The hit rates are assumptions, not sourced figures, and each ad is treated as an independent trial at a fixed rate.

Read it the other way round and you get the number of ads for a given chance.

Ads launched in the test window for a given chance of at least one winner, rounded up. Illustrative arithmetic under assumed hit rates. Not a forecast, not a benchmark.
Assumed hit rate50% chance80% chance90% chance95% chance
2%3580114149
5%14324559
10%7162229
15%5101519
30%2579

At an assumed 5% hit rate, 14 independent ads give about an even chance of at least one winner and 45 give about 90% (illustrative binomial arithmetic, not a forecast).

Two assumptions sit under every cell, and the NIST handbook names both: each ad is an independent trial, and the hit rate is fixed for every ad. Real accounts bend both. Ads built on the same concept tend to win or lose together, and a tired audience lowers the rate as the window runs on.

Under a relative definition like the top 2% of spend, winners aren't independent at all. Launching more ads changes the count of winners, not the rate. Use these tables with a CPA or significance definition you wrote down before launch. They give the chance of at least one winner under these assumptions, and they say nothing about how any particular ad will perform. Results vary.

Both tables assume you already have ads worth launching, and the Gap Scan shows where the first ones could point.

Free Creative Gap Scan →

Free. 3 researched opportunities, 1 customer-avatar gap, and 1 example concept within 24 hours.

Is a winner you found a winner you can trust?

Not automatically. According to Berman and Van den Bulte (2022), the false discovery rate in the experiments they studied was 18 to 25% at 5% significance. Roughly one in five results that cleared the bar may have had no real effect. A winner found isn't a winner proven. For the interval math on one result, see how to tell whether one winner is real.

What should count as a winner in your account?

Write the definition down before you launch, and leave it alone until the window closes. It takes three decisions.

First, the metric: CPA, ROAS, or a share of spend. Pick the one your business actually runs on. Second, the bar, which should be your own target CPA or ROAS rather than a benchmark from someone else's account. Third, the window: how many days after the learning phase ends the ad has to hold that bar, and the minimum spend it needs before it counts.

You set the target CPA, the spend floor and the number of days before launch. A definition you can change after seeing results will always produce a good hit rate, so fix it first and count second.

Once winners are defined, the next question is why the others lost, and which metric points at which part of the ad: hook rate, hold rate, CTR and CVR.

Here is how one customer described their start:

In our first 60 days, we launched 47 CreatStrat concepts. Nine became winning ads and our blended creative CPA dropped 23%.

Naomi C. · CreatStrat customer

That's one customer's account. Results vary, and it isn't a promise or an average.

Frequently asked questions

What is a good hit rate for Meta ads? No published figure covers DTC Meta ads. Sources range from 2% of ad variations taking most spend (AppsFlyer, 2024) to 12% of web experiments with a significant lift (Optimizely). A good rate is one measured against a definition you wrote before launch.

Why do hit rate statistics disagree so much? Each source counts a different win: a share of spend, a CPA at target, or a significant lift. Kohavi, Deng and Vermeer (2022) note that companies count success differently, so they treat published success rates as ballpark figures only.

How many ads do I need to test to find one winner? It depends on your hit rate. Under illustrative binomial arithmetic with independent ads, a 10% hit rate needs about 7 ads for a 50% chance of at least one winner and 22 for 90%. Not a forecast.

What counts as a winning ad? Whatever you define before launch. A common version is a CPA at or under your target, held for a set number of days after the learning phase, above a minimum spend. Fix the definition before results come in.

If I found a winner, how sure am I that it's real? Not fully. Berman and Van den Bulte (2022) found a false discovery rate of 18 to 25% at 5% significance, so about one in five significant results may have no real effect.

The hit rate worth tracking is yours

A published hit rate tells you how someone else counted. The one worth tracking is the share of your ads that clear a win condition you wrote down before launch, measured the same way every week.

CreatStrat is a software subscription that delivers finished ads every week; CreatStrat researches, plans the tests, writes the brief, and produces the finished statics, carousels, mashups, AI video, and AI UGC, each with its Meta copy; your creative strategist is built in. Your team launches the ads and sets the win condition.

If you want to see which researched opportunities your account isn't testing yet, start with a product page URL and an email.

Free Creative Gap Scan →

Free. 3 researched opportunities, 1 customer-avatar gap, and 1 example concept within 24 hours.

Want to see the app first? Book a demo →

Sources

  1. Kohavi, R., Deng, A., & Vermeer, L. 2022. A/B testing intuition busters: Common misunderstandings in online controlled experiments. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3168-3177. doi:10.1145/3534678.3539160.
  2. Berman, R., & Van den Bulte, C. 2022. False discovery in A/B testing. Management Science. 68(9):6762-6782. doi:10.1287/mnsc.2021.4207.
  3. Azevedo, E. M., Deng, A., Montiel Olea, J. L., Rao, J. M., & Weyl, E. G. 2020. A/B testing with fat tails. Journal of Political Economy. 128(12):4614. doi:10.1086/710607.
  4. NIST/SEMATECH. e-Handbook of Statistical Methods, section 1.3.6.6.18: Binomial distribution. National Institute of Standards and Technology. https://www.itl.nist.gov/div898/handbook/eda/section3/eda366i.htm
  5. AppsFlyer. 15 May 2024. Exploring the ad creative landscape in the era of AI: AppsFlyer's comprehensive report reveals key industry patterns (press release on the State of Ad Creatives in App Marketing report). https://www.appsflyer.com/company/newsroom/pr/creative-optimization-data-report/
  6. AppsFlyer. 29 Apr 2025. AppsFlyer's 2025 creative report reveals how AI and emotion drive creative performance (press release on the 2025 State of Creative Optimization report). https://www.appsflyer.com/company/newsroom/pr/ai-emotion-creative-trends/
  7. Optimizely. Top 10 takeaways from running 127,000 experiments. Optimizely Field Notes. https://www.optimizely.com/insights/top-10-takeaways-from-running-127000-experiments/
Free Creative Gap Scan

Start with something useful.

Share a product page URL and email. Within 24 hours we’ll send 3 researched opportunities, 1 customer-avatar gap, and 1 example concept.