Search for a creative hit rate and you get answers that refuse to agree. In AppsFlyer's 2024 report, 2% of ad variations took 68% of spend. Optimizely found 12% of web experiments produced a significant lift, and Microsoft put its experiment success rate at 33% (Kohavi, Deng and Vermeer, 2022).
If your own number sits somewhere else, the account probably isn't broken. These figures disagree because each one counts a different thing as a win, and only the first is about ads. The more useful question is how many ads you need to launch for a fair chance at one winner.
What is a creative hit rate?
A creative hit rate is the share of ads launched that meet a stated win condition, and published hit rates disagree because each source counts a different thing as a win: a share of spend, a CPA held at target, or a statistically significant lift.
The arithmetic is winners divided by ads launched. The hard part is the top of that fraction, because Ads Manager doesn't decide for you what a winner is.
Each win condition gives a different number for the same account. A spend-share definition that counts the top slice of ads as winners reports a hit rate close to the size of that slice, whatever you launch. A CPA definition moves with your target: tighten it and the rate falls. A significance definition depends on traffic, so small budgets produce fewer winners even when the ads are good.
So the hit rate is a choice before it's a result. Whether one winner is real is a separate check: see what statistical significance means for one test.
What counts as a win, and what rate does each definition give?
Here are the published figures people quote, lined up by what each one counted as a win. Only the AppsFlyer rows are about ads, and those are mobile app installs, not DTC brands. The other rows come from website and product experiments. They're included because ad testers quote them as if they were ad benchmarks.
| What counts as a win | Published rate | Source | What it measures |
|---|---|---|---|
| Top slice of spend | 2% of ad variations ("one in fifty") | AppsFlyer, 15 May 2024 | Where budget went across 220,000 creative variations from over 2,000 apps |
| Top slice of spend, by category | Top 2% of creatives took 53% of spend in Gaming and 43% in Non-Gaming | AppsFlyer, 29 Apr 2025 | Spend concentration across 1.1 million video creatives from 1,300 apps |
| Significant lift on the primary metric | 12% of experiments | Optimizely, 127,000 experiments | Website and feature tests, not ads |
| Company-reported experiment success | 8% (Airbnb Search) to 33% (Microsoft) | Kohavi, Deng and Vermeer, KDD 2022, Table 2 | Each company's own accounting of product and website experiments, not ads |
How to read it: a published rate only means something next to its definition. AppsFlyer measured spend concentration in mobile app marketing. Optimizely counted experiments that beat control on their primary metric. Kohavi and his co-authors collected what each company reported as a success. None of these is a DTC Meta benchmark, and we didn't find a published figure that is.
According to Kohavi, Deng and Vermeer (KDD 2022), published experiment success rates range from 8% at Airbnb Search to 33% at Microsoft, and the authors call them ballpark figures because companies count success differently.
Their own caveat: "These numbers may involve different accounting schemes, and we never know the true rates, but they suffice as ballpark estimates." A single figure lifted from that table can't tell you whether your account is on track.
Why a top-slice-of-spend definition sets its own hit rate
When a winner is defined as the top slice of spend, the hit rate is fixed by the definition: AppsFlyer's 2024 report found that 2% of ad variations took 68% of spend, which it described as one in fifty ads being a winner.
AppsFlyer's press release describes the dataset as "an anonymous aggregate of proprietary global data from over 2,000 apps, looking at 220,000 creative variations, and 720 million marketing-driven installs." The full report is gated, so the definition here comes from the release. Ad platforms sent budget to the ads that engaged best, and the ads at the top of that spend curve counted as winners.
Here's the catch. Under a relative cut-off, winners aren't independent draws. If the top 2% by spend are the winners, 100 ads produce about two and 1,000 ads produce about twenty (illustrative). The count of winners moves with volume. The rate stays pinned near the cut-off, because the cut-off is the rate.
So 2% describes where budget went across 220,000 app-install variations. It isn't a prediction for your next ten ads, and it says nothing about whether the ads that got the budget were profitable. How many new ads to launch each week at your spend level is a different question with its own answer.
One honest note about this post's title. The 33% end of the range comes from Microsoft's product experiments, not ads, and the 2% end comes from app-install ads. "2% to 33%" is the spread of published definitions of a win. It isn't the spread of ad results you should expect.
What the experiment research says about winners
The careful work on win rates comes from website and product experiments, but it answers a question every ad tester has: how often does a tested change do anything at all?
Rarely. According to a 2022 study by Berman and Van den Bulte in Management Science, which analysed 4,964 effects from 2,766 experiments, about 70% of tested changes had no true effect. Most ideas don't move the metric.
That sounds bleak until you read a second paper. Azevedo and colleagues, in a 2020 study in the Journal of Political Economy built on Microsoft Bing's experimentation platform, found that idea quality is fat-tailed. Most ideas do little and a rare few do a lot. Under that distribution, their analysis favours a "lean" approach: test more ideas with smaller samples rather than a few big experiments. The research suggests volume matters when most ideas fail, because the rare large effect is what moves the result. It doesn't say any particular number of tests will turn one up.
Optimizely, reporting on 127,000 experiments, offers the sanity check from the other side: "If nine out of ten of your tests are winning, something is off." A very high hit rate is more often a loose definition than a great account.
How many ads does it take to find at least one winner?
Once you've fixed a definition and have a rough hit rate, the planning question has an answer. The binomial distribution, as set out in the NIST/SEMATECH e-Handbook of Statistical Methods, gives it in one line:
Chance of at least one winner = 1 - (1 - hit rate) raised to the power of the number of ads.
At a 10% hit rate and 10 ads, that's 1 - 0.9^10, about 65% (illustrative). Both tables below are CreatStrat's own arithmetic from that formula, rounded. The rows in both tables are assumed hit rates, not sourced ones. The 5% row, for example, is an assumption and doesn't come from any study.
| Assumed hit rate | Ads launched in the test window | ||||
|---|---|---|---|---|---|
| 5 | 10 | 20 | 40 | 80 | |
| 2% | 9.6% | 18.3% | 33.2% | 55.4% | 80.1% |
| 5% | 22.6% | 40.1% | 64.2% | 87.1% | 98.3% |
| 10% | 41.0% | 65.1% | 87.8% | 98.5% | >99.9% |
| 15% | 55.6% | 80.3% | 96.1% | 99.8% | >99.9% |
| 30% | 83.2% | 97.2% | 99.9% | >99.9% | >99.9% |
Illustrative arithmetic under assumed hit rates. Not a forecast, not a benchmark.
Read it the other way round and you get the number of ads for a given chance.
| Assumed hit rate | 50% chance | 80% chance | 90% chance | 95% chance |
|---|---|---|---|---|
| 2% | 35 | 80 | 114 | 149 |
| 5% | 14 | 32 | 45 | 59 |
| 10% | 7 | 16 | 22 | 29 |
| 15% | 5 | 10 | 15 | 19 |
| 30% | 2 | 5 | 7 | 9 |
At an assumed 5% hit rate, 14 independent ads give about an even chance of at least one winner and 45 give about 90% (illustrative binomial arithmetic, not a forecast).
Two assumptions sit under every cell, and the NIST handbook names both: each ad is an independent trial, and the hit rate is fixed for every ad. Real accounts bend both. Ads built on the same concept tend to win or lose together, and a tired audience lowers the rate as the window runs on.
Under a relative definition like the top 2% of spend, winners aren't independent at all. Launching more ads changes the count of winners, not the rate. Use these tables with a CPA or significance definition you wrote down before launch. They give the chance of at least one winner under these assumptions, and they say nothing about how any particular ad will perform. Results vary.
Both tables assume you already have ads worth launching, and the Gap Scan shows where the first ones could point.
Free. 3 researched opportunities, 1 customer-avatar gap, and 1 example concept within 24 hours.
Is a winner you found a winner you can trust?
Not automatically. According to Berman and Van den Bulte (2022), the false discovery rate in the experiments they studied was 18 to 25% at 5% significance. Roughly one in five results that cleared the bar may have had no real effect. A winner found isn't a winner proven. For the interval math on one result, see how to tell whether one winner is real.
What should count as a winner in your account?
Write the definition down before you launch, and leave it alone until the window closes. It takes three decisions.
First, the metric: CPA, ROAS, or a share of spend. Pick the one your business actually runs on. Second, the bar, which should be your own target CPA or ROAS rather than a benchmark from someone else's account. Third, the window: how many days after the learning phase ends the ad has to hold that bar, and the minimum spend it needs before it counts.
You set the target CPA, the spend floor and the number of days before launch. A definition you can change after seeing results will always produce a good hit rate, so fix it first and count second.
Once winners are defined, the next question is why the others lost, and which metric points at which part of the ad: hook rate, hold rate, CTR and CVR.
Here is how one customer described their start:
In our first 60 days, we launched 47 CreatStrat concepts. Nine became winning ads and our blended creative CPA dropped 23%.
Naomi C. · CreatStrat customerThat's one customer's account. Results vary, and it isn't a promise or an average.
Frequently asked questions
What is a good hit rate for Meta ads? No published figure covers DTC Meta ads. Sources range from 2% of ad variations taking most spend (AppsFlyer, 2024) to 12% of web experiments with a significant lift (Optimizely). A good rate is one measured against a definition you wrote before launch.
Why do hit rate statistics disagree so much? Each source counts a different win: a share of spend, a CPA at target, or a significant lift. Kohavi, Deng and Vermeer (2022) note that companies count success differently, so they treat published success rates as ballpark figures only.
How many ads do I need to test to find one winner? It depends on your hit rate. Under illustrative binomial arithmetic with independent ads, a 10% hit rate needs about 7 ads for a 50% chance of at least one winner and 22 for 90%. Not a forecast.
What counts as a winning ad? Whatever you define before launch. A common version is a CPA at or under your target, held for a set number of days after the learning phase, above a minimum spend. Fix the definition before results come in.
If I found a winner, how sure am I that it's real? Not fully. Berman and Van den Bulte (2022) found a false discovery rate of 18 to 25% at 5% significance, so about one in five significant results may have no real effect.
The hit rate worth tracking is yours
A published hit rate tells you how someone else counted. The one worth tracking is the share of your ads that clear a win condition you wrote down before launch, measured the same way every week.
CreatStrat is a software subscription that delivers finished ads every week; CreatStrat researches, plans the tests, writes the brief, and produces the finished statics, carousels, mashups, AI video, and AI UGC, each with its Meta copy; your creative strategist is built in. Your team launches the ads and sets the win condition.
If you want to see which researched opportunities your account isn't testing yet, start with a product page URL and an email.
Free. 3 researched opportunities, 1 customer-avatar gap, and 1 example concept within 24 hours.
Want to see the app first? Book a demo →
Sources
- Kohavi, R., Deng, A., & Vermeer, L. 2022. A/B testing intuition busters: Common misunderstandings in online controlled experiments. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3168-3177. doi:10.1145/3534678.3539160.
- Berman, R., & Van den Bulte, C. 2022. False discovery in A/B testing. Management Science. 68(9):6762-6782. doi:10.1287/mnsc.2021.4207.
- Azevedo, E. M., Deng, A., Montiel Olea, J. L., Rao, J. M., & Weyl, E. G. 2020. A/B testing with fat tails. Journal of Political Economy. 128(12):4614. doi:10.1086/710607.
- NIST/SEMATECH. e-Handbook of Statistical Methods, section 1.3.6.6.18: Binomial distribution. National Institute of Standards and Technology. https://www.itl.nist.gov/div898/handbook/eda/section3/eda366i.htm
- AppsFlyer. 15 May 2024. Exploring the ad creative landscape in the era of AI: AppsFlyer's comprehensive report reveals key industry patterns (press release on the State of Ad Creatives in App Marketing report). https://www.appsflyer.com/company/newsroom/pr/creative-optimization-data-report/
- AppsFlyer. 29 Apr 2025. AppsFlyer's 2025 creative report reveals how AI and emotion drive creative performance (press release on the 2025 State of Creative Optimization report). https://www.appsflyer.com/company/newsroom/pr/ai-emotion-creative-trends/
- Optimizely. Top 10 takeaways from running 127,000 experiments. Optimizely Field Notes. https://www.optimizely.com/insights/top-10-takeaways-from-running-127000-experiments/
