Issue #185 | The Volume vs. Quality Debate

I hope you’re all enjoying the first weekend of (unofficial) fall! If your week was anything like mine, it was jam-packed with 2027 planning meetings, BFCM and a few last minute, we-have-a-problem-that-needs-solved-in-2026 client requests. Plus some back-to-school drama (because why not?) and the (glorious) return of football.
One of those many meetings was with a senior living operator spending $180k-$200k/yr on creative production, yet wildly frustrated with ad performance. When we finally got access to their account, we found that their team published a raw total of ~240 unique ad units over the last 12 months. But, when we decomposed and analyzed those ~240 individual assets, we found that there were only 6 core concepts, each with ~40 executions (different headlines, different image perspectives, different headlines, different first frames, different primary text, different ratios).
Honestly, this isn’t uncommon in Meta accounts, especially as the “just make more creative” side has taken a lead in the “how to optimize Meta” wars. More brands/agencies are simply modifying existing assets and pushing them into the account, under the mistaken notion that they are adding creative diversity.
But what’s really going on under the surface?
The brand/agency sees 240+ ad units launched in the last 12 months. That’s ~20 a month. They think they are CRUSHING it, especially relative to competitors who struggle to launch one new ad a quarter. Meta sees 6 ads with different presentations.
The distance between those two numbers is worth ~$1.5M/yr to that operator and closing it costs a 1/3rd of what they’re already spending.
So, let’s talk about what’s really happening here and how you can apply it to your account.
Two strategies, one equation
But first: a quick detour through my childhood. I promise this is relevant later, so just bear with me.
As a kid, I LOVED playing StarCraft. I’d do school/sports during the day, then stay up all night playing video games, get 3-4 hours of sleep from ~4am to ~8am, then do it all over again. I played competitively for years (I didn’t have much of a social life then). I was world-ranked. As my wife now reminds me, what really happened was that I used my lifetime allotment of video game time in the span of high school.
For the unfamiliar, StarCraft is a real-time strategy game, with 3 playable races that have different economies + different strategies. The two that matter here are Protoss & Zerg:
Protoss units are the American military of video games: wildly expensive, technologically powerful, and any unit lost takes forever to replace. They even come with rechargeable shields (which 12 year old me thought was really damn cool). 12 of them beat 40+ Zerg units in a straight fight. All of that means the strategy when you play Protoss is simple: win each individual engagement so fast that the opponent can’t destroy any of your units. The problem is that if you do lose a bunch of those units, you (almost always) lose the game, simply because it takes too long and costs too much to replace them.
Zerg is the polar opposite: units are cheap, individually weak and replaced almost immediately from a production mechanism that shares a resource pool with the economy itself. When you play Zerg, the objective shifts completely. You don’t try to win individual engagements; you try to lose in a resource-positive manner (i.e. trade 3 exceedingly cheap units for 1 very expensive unit), then rebuild your units faster than your opponent can replace theirs. Repeat that process over a longer game, and the Zerg player wins either through overwhelming numbers or resource exhaustion.
It took the StarCraft players a decade to formalize what was actually going on here: the Zerg advantage is not quantity. Quantity is just the visible part. The advantage is cheap units + fast replacement = a change to the optimization equation.
Protoss optimizes the outcome of a given engagement. Zerg optimizes the ratio of resources lost to resources destroyed, aggregated across many fights, while accepting that most individual encounters go badly.
Now, strip out all of the video game talk – and suddenly that pattern looks very familiar in creative production. In fact, I’d go as far as to say that virtually every creative program in paid media is one of these two, though almost no one chooses which side they’re playing intentionally. Agencies with production capability sell Protoss, because their cost structure requires it (maintaining a production studio is EXPENSIVE). Performance shops sell Zerg, because they don’t have the infrastructure to produce the shiny, expensive stuff + volume is what they can staff. Operators pick whichever their last CMO/Head of Growth/agency trained them to expect.
There’s nothing inherently wrong with either strategy. Both can be wildly effective, but only if you (1) acknowledge which game you’re playing and (2) measure the relevant parameters accordingly.
So, without further ado, let’s return to our initial example:
What testing actually buys you
The most common justification for launching hundreds of ads (and, in fact, the one this operator’s agency gave them) is that it “tests” different ads to identify winners.
But that’s not exactly true. At its most fundamental level, a creative test does NOT return the median performance of the creatives tested; it yields the performance of the best one, because that is the one you keep running. In creative testing, you’re not buying a mean, you’re buying a maximum.
Let’s define creative quality as a multiplier: a given creative has some factor, call it m, by which it multiplies conversion efficiency relative to the median creative in your account. A creative with m = 1.4 produces 40% more conversions per dollar than your median asset. Across an account, m is (approximately) lognormal: bounded below at zero, unbounded above, most assets clustered near the middle with a thin right tail.
One number describes that spread: σ, which is the standard deviation of log m. A σ of 0.15 means your creatives are tightly clustered and your best is not far from your median. A σ of 0.45 means the tail is long and your best asset is several times your median.
Run N creatives, keep the best, and what you get is E[max of N draws]. That expectation grows with N, sublinearly, and how fast it grows is dependent on σ.
That means we can model the impact of unique concepts on any account:
Focus on the columns, not the rows. At σ = 0.15, going from 4 concepts to 64 yields 21% improvement. Going from 64 to 256 yields another 8%.
But, at σ = 0.45, that same move from 4 to 64 yields 75%. The volume of unique concepts pays in exact proportion to how dispersed your outcomes are.
If your creative outcomes are tightly clustered (i.e. σ = ~0.15), there is no achievable amount of creative volume that will produce a 1.75x median creative. Mathematically, if you were to test 256 concepts, you should expect to land at 1.53x median – which is still very good, but maybe less good when you consider you had to develop, execute and monitor 256 unique concepts to get there. In an account with a low σ, the lever simply isn’t that long. The solution there is to raise the median instead, which is audience research, product development, user experience + offer work.
On the other extreme, at σ = 0.30+, craft can’t get you there either, because a production process that reliably lifts median performance 75% simply doesn’t exist. The only way to achieve that reliably is to sample the tail.
Which means, if we return to my video game roots, Protoss and Zerg are not differing philosophies; they are individually correct answers to a single question asked at different values of σ.
The Craft Premium
Let me illustrate.
Imagine if we compared these two approaches – the polished, high-end “Protoss-style” creative shop vs. the down-and-dirty meat-grinder “Zerg” performance shop – using the same budget.
Program A spends $180k to produce 6 concepts at $30k/each: proper scripting, real production, professionally-done testimonials, true creative direction, exceptional editing.
Program B spends that same $180k to produce 240 creative concepts at $750 each using an asset library + an editor working off a template (basically, the CreativeOS model).
Now, the comparison that gets made is whether the $30k work of art beats the $750 production-line ad. Candidly, that is the wrong comparison and making it is exactly what the high-end shop wants you to do. The reality is that a researched, scripted, expertly-produced $30k spot SHOULD beat a $750 template-driven spot head to head. To quote the immortal Chris Rock: “That’s what it’s SUPPOSED to do!”
The right comparison to make is whether Program A can beat Program B’s best of N. To do that, the expensive process must lift the entire distribution rather than produce a single good asset. At σ = 0.30, best-of-6 returns 1.49 and best-of-40 returns 1.93. The ratio is 1.296. Translated: the expensive “Protoss” style creative shop must raise the median performance of the entire account/program by 29.6% for 6 expensive concepts to match 40 cheap ones.
You might read that and think a real production process obviously clears 30%. Maybe it does in certain situations, but no brand I’ve worked with can demonstrate it. In the analyses we’ve done (and when we can isolate creative impact relatively cleanly), the actual incremental lift from exceptionally high creative tends to fall between 10% and 20%. So yes, the fancy agencies are right when they say high production value is the rising tide that lifts all boats. They just leave out the part that it’s a little 2’ storm surge, not a Roland Emmerich film.
Why?
Because most of what separates a winning ad from a losing one is the level of audience understanding, the angle and the offer, and those cost almost nothing to change. That’s why (a few months back) I wrote the article that 80% of your ad account’s success isn’t in your ad account.
None Of This Matters If You Can’t Survive Being Wrong
If we shift gears slightly, one of my favorite modern essayists + financial thinkers is Nassim Nicholas Taleb. One of the central premises of his work (which I highly recommend reading, just not before bed – it’s dense) is that high expected value outcomes only matter IF you can survive the cost of losing repeatedly.
Number of draws = budget divided by cost per attempt: N = B / c. Cost per attempt is the fully loaded cost of finding out you were wrong, not just the production bill (i.e. production, the media required to resolve the asset as a loser, the costs of any tools required to make the determination AND the fraction of the team’s time this entire process required).
Using this math, reaching 40 draws at $750 per attempt requires ~$30,000.
Reaching 40 draws at $30,000 per attempt requires $1.2M.
99% of all brands do not have a $1.2M creative budget, which means that strategy was never truly open to you. The cost per attempt was prohibitively high before anyone even had a chance to run the numbers.
This is the variable almost no one thinks about – let alone optimizes – and it decides which strategies you are permitted to consider at all. Cost per hit (or cost per winner, depending on your preferred vernacular) is discussed constantly. But Cost per failure is what sets the size of your option set, because any creative program will yield far more failures than winners and you pay for them repeatedly.
All of that means that – for 99% of brands – cheap production is not a quality compromise. It is an eligibility purchase.
The Ceiling That Everyone Ignores
Now, selection can only act on variance that differs between the things being selected. So, let’s divide creative quality discussed above into a shared component (e.g. everything your ads have in common, the same offer, the same landing page, the same warm-light-and-smiling-customer aesthetic) and a distinct component.
Let’s term the shared fraction ρ.
Selection works on the distinct component only, so the dispersion available to it is:
σselectable = σ × √(1 − ρ)
The shared component moves every variant identically. It cannot be tested away, because no variant in the account lacks it. You carry it whether it helps or hurts.
At the operator level this resolves into something more blunt: if 240 executions = 6 concepts * 40 executions, then replicates are near-identical draws and selection across them yields (effectively) nothing beyond noise.
Your effective draw count is the number of distinct concepts you produced, not the number of unique files in the creative library and certainly not the number of rows in the “ads” tab Meta Ads manager (or whatever reporting tool you use).
Putting this together with the math from above:
240 executions of 6 ideas yields ~1.49x median, while 40 distinct ideas produces ~1.93x median.
The distribution is the same, the difference between them is the same σ (29.6%) apart, and the second costs $150k less. Math is cool when it saves you a new Porsche Taycan.
You might read this and think, “There’s NO WAY I’d ever make that mistake! Our account clearly has more than 6 ideas in it!” Maybe. But have you actually checked?
Pull the last 12 months of creative and sort by the psychological claim being made rather than by asset name, format, style or launch date. The data from thousands of accounts is that the overwhelming majority resolve to between 4 – 8: VSL or branded message, features/benefits, a core emotional trigger (relief, peace of mind, happiness), a statistic/proof point, a location or pricing message & a seasonal promotion. Everything else is a variant on one of those.
So, the question is: well, what else is there?
The honest answer is that it depends on your product/market, but it’s design-able with proper planning. If a market contains K distinct viable psychological territories and you generate concepts without a plan, expected coverage of all K requires K·ln(K) concepts. If you plug in real numbers – say K = 8 – then you need 22 concepts to cover all of them if you’re designing without a plan. But, if you generate them deliberately (i.e. 1 per territory), then you only need 8 (maybe 10 if a few are on the fringes). Sampling design beats sample size by a factor of 3.
What You Keep vs. What You Measure
Everything above describes what you observe in your account/creative program – but what you actually deploy is – by definition – worse. And the best part? You can calculate exactly how much worse.
Whenever you select a “winning” ad, you’re doing so under noisy conditions, which means you select partly on noise. The concept that looked best got some help from stuff unrelated to its core properties – maybe the concept was especially resonant because of a trend, maybe many of your competitors decided to try out a weird ad format, maybe something in the Zeitgeist just looked favorably upon your creative team, maybe the macroeconomic conditions were just swell, whatever. Whatever component of the initial determinant was noise does NOT persist once the ad goes into the “evergreen” or “always on” rotation.
The shrinkage is the reliability ratio:
r = σ2true / ( σ2true + σ2noise )
Deployed lift ~= r * observed lift.
σ²_noise is set by how many conversions each concept accumulated during the “experimental” or “test” phase, which makes the penalty worse as concept count increases. An account generating 1,116 conversions across 6 concepts generally produces a pretty reliable read; but if you split those same 1,116 conversions across 40 concepts, the measurement gets far more noisy and unreliable:
If you’ve made it this far, you’re probably thinking, “Wait – this is directly opposed to the volume thesis. More ads = less reliable winners?”
Yes. That’s EXACTLY right. You always pay to learn; the question is how you pay.
The best solution I’ve found is to price it into the forecast vs. arguing with it. It also argues for observing performance across a longer window than feels comfortable, simply because if you decide too fast, the probability that the “winner” you’ve just promoted is part (or mostly) noise is substantially higher.
The Half Life Of Winners
The final mistake I see – and the most nefarious of all of them – is the presumption that winners perform at a constant level throughout their life. Nothing could be further from the truth.
The reality is that living organisms. And like all living organisms, they decay over time. The reasons why this happens are numerous, but the most common: audience saturation, higher frequency, decline in novelty, continued innovation, shifting macroeconomic environments (i.e. a premium angle is unlikely to perform as well in a declining economy vs. a booming one). All of them result in the same thing: the performance of a previously-exceptional concept regressing back toward the account median.
Mathematically + financially, that means that what you *really* yield from all that creative testing is the time-average of a sawtooth, not the height of its peak.
Guess what? We have math for that, too!
Model a winner decaying exponentially toward 1.0 with half-life h, replaced every T months, and the average multiplier you run is:
1 + (M − 1) × (1 − e−δT) / (δT), where δ = ln 2 / h
At a 4-month half-life and a best-of-40 winner, refreshing every 6 weeks maintains 1.74x median performance. Every 3 months holds 1.64. Every 6 months holds 1.54. Every 12 months holds 1.38. The net/net = a 40-concept creative program that only refreshes annually performs like an 8-concept program sampled at its peak. In economics, we’d consider production cadence and concept count to function as non-equivalent substitutes, which means you can price the exchange between them.
If you evaluate the program faster, then each tested concept accumulates fewer conversions, reliability falls, and you promote noise. Conversely, if you evaluate slower, decay eats the winner you correctly identified. The two errors move in opposite directions as you change production velocity, which means there is – necessarily – an interior optimum. For this portfolio at 40 concepts, that optimum sits at ~8 weeks: 1.657x at 2 months vs. 1.649x at 6 weeks and 1.642x at 3 months. The curve is flat across that whole range, which means you don’t need to hit 8 weeks precisely, but you do need to stop refreshing 2x a year.
3 Approaches To This Problem
For this particular brand, they had 18 communities at $9,400/community/month on Meta. That’s $2.03M annually. At 1,116 inquiries/mo, a 28% lead-to-tour rate and 23% tour-to-move-in rate, that yields 862 move-ins/yr at $2,354/move-in. At $6,800/mo in gross revenue and a 26-month average length of stay, each move-in is $176,800 in revenue and (using a ~38% contribution margin), ~$67,200 in contribution margin.
The honest reality is that creative does not control all (or even most) of that. Offer, audience, landing page, market conditions and sales execution carry the lion’s share. If we assume creative is responsible for 20% (an assumption reasonable people can argue about), φ = 0.20, then we can price lift accordingly:
Program C produces 47 more move-ins a year than the program it replaces, on a creative budget ~33% the size: roughly $3.2 million in additional lifetime contribution, about $1.5 million annualized. Program A, the expensive and well-made version, gets you 18 move-ins for 3x C’s creative cost.
The winner’s curse cut C’s advantage, because 40 concepts produce noisier data than 6. Decay handed it back and then some, because $750 concepts can be replaced every 8 weeks and $30,000 concepts simply can’t. Most of C’s edge comes from that 2nd effect vs. the raw creative count, which is not what anyone talks about.
What Happens After You Win?
The instinct once you identify winning creatives is to consolidate as much budget behind those single (or handful) of top performers. That’s what every “Meta Ads” expert on X tells you to do.
But…once again, the consensus is wrong.
The covariance math argues against it. A 50/50 blend of 2 uncorrelated concepts carries 0.71x the outcome volatility of running one alone. At ρ = 0.5 it’s 0.87x, and at ρ = 0.8 it’s 0.95x and gets you next-to-nothing.
If you’ve followed this entire issue through, you’re probably realizing the same ρ that capped your effective draws determines whether diversifying your live mix does anything at all. 2 winners from different psychological territories stabilize the account. 2 winners that are different executions of the same underlying idea/psychological territory do not. When one fails, the odds are the other will, too – and then you’re stuck desperately trying to find another winner before the cost of sub-median performance ruins your week/month/quarter/whatever.
So, What Should You Do?
First: please, for the love of all things holy, stop counting assets as unique concepts. Instead, start counting distinct psychological claims by sorting your library by what it argues rather than what it looks like. That number is your N.
Second: measure σ before choosing a strategy. Take 90 days of creative-level conversion efficiency, take logs, calculate the standard deviation. Under ~0.20 you are in a craft regime, so volume will not save you. At 0.30+, you are in a sampling regime, which means you need volume. Know which game you’re playing BEFORE you start making agency or allocation decisions.
Third: attack cost per failure vs. cost per winner, because that number decides whether the sampling regime is open to you at all.
Fourth: set your creative fresh rate intentionally and deliberately based on the above data, not by an agency schedule, vibes or X. In most cases, somewhere around 8 weeks is optimal for an account like the size of the one I’ve mentioned (~$2M in annual media), then accept that the winner you identify is actually worth ~15-25% less than what you think at the time you pick it. Decay is real, so just plan for it. .
And above all, remember that cheap units are bad units, but the Zerg player wins anyway (about 52% of the time vs Protoss), because being wrong cheaply, in different ways, and replacing the answer before it goes stale beats being right expensively once.
The best part? You can calculate all of it.

