A fractional CMO running roughly $300,000 a year on Meta said this to us in July, and it is the most accurate description of the problem we have heard:
"Meta is not clear what they mean when they say creative diversity. It is very challenging for the advertiser."
He is right, and the reason is structural. Creative diversity is not a creative-department concept that Meta borrowed. It is a description of what happens inside Meta Andromeda, the retrieval engine that decides which of your ads is even eligible to be considered. Once you look at it from the engine's side, the definition stops being fuzzy and becomes almost arithmetic.
What Andromeda actually is

Meta published the technical description on its own engineering blog in December 2024, and the numbers in it are the whole story.
Meta's ad system is multi-stage. Retrieval comes first, and its job is to take tens of millions of ad candidates and cut them down to a few thousand relevant ones. Only then do the larger ranking models take over, predict value, and decide the final set shown to a person.
Andromeda is the machine that does that first cut. Meta describes it as processing three orders of magnitude more ads than every subsequent stage, running on custom deep neural networks with a stated 10,000-fold increase in model capacity, deployed across Facebook and Instagram for a 6% recall improvement to the retrieval system and an 8% ads quality improvement on selected segments.
The part built specifically for you is the indexing. Meta says the hierarchical index exists "to support exponential ad creatives growth", designed to absorb the flood of assets coming from Advantage+ and generative tooling. At the time of writing, over a million advertisers were producing more than 15 million ads a month with Meta's generative tools.
And then the roadmap line, which is the one worth pinning to the wall. Meta writes that the architecture is expected to move to an autoregressive loss function "leading to a more efficient and faster inferencing solution that delivers a more diverse set of ad candidates", and that "increased ad diversity can improve people's experience with ads and drive better advertiser outcomes."
Diversity is not a hint. It is the stated direction of the retrieval engine, and Meta has since made the same argument to advertisers directly.
This article is about the mechanism and how to audit against it. If you want to see what acting on it looks like at scale over a full year, we documented that separately in how creative diversity scaled one account's Meta spend.
The reframe: you are not uploading ads, you are contributing candidates

Here is the sentence that resolves the confusion.
Retrieval selects distinct candidates, not files. If you upload forty assets that occupy nearly the same position in the model's representation of ad content, you have not contributed forty candidates to a pool of tens of millions. You have contributed something much closer to one candidate wearing forty outfits.
This is why "we tested a lot and learned nothing" has become such a common complaint. The volume was real. The distinctness was not.
It also explains why audience saturation and the classic ad fatigue story changed shape. Fatigue used to be a frequency problem: the same people saw the same ad too many times. Increasingly it is a retrieval problem: your account stops being able to reach new pockets of people because it has nothing structurally new to offer the first cut.
Weak axes and strong axes

When we ask advertisers what they varied, the answer is usually the format. Static, then carousel, then video. That is the weakest axis available, and it is the one most teams spend their budget on.
Here is a working hierarchy, ordered by how much each actually moves your position in the candidate pool.
Weak axes, worth very little on their own
- Vehicle. Static versus carousel versus video, holding everything else constant. It changes the container, not the content.
- Hook rotation on an identical body. Five openings, one script, one visual world. This was the standard iteration unit for years. It is now closer to a rounding error.
- Cosmetic variation. New colourway, new crop, new caption phrasing, same everything else.
- Aspect ratio and placement variants. Necessary hygiene. Not diversity.
Strong axes, where the distance actually comes from
- Persona. Not demographic buckets, but a genuinely different person with a different reason to buy. The offshore fisherman and the pond fisherman are not two segments of one ad, they are two ads.
- Problem framing. The same product sold as a fix for a symptom, as an upgrade to a routine, as a gift, as an identity signal. Four different products, functionally.
- Format world. Comic panels, screen recording, unpolished handheld, editorial still life, illustrated explainer, text-on-plain-background. These occupy genuinely different regions of the space.
- Register. Funny, clinical, urgent, calm, confessional. Register travels with the audience, not the product.
- Source of authority. Founder, customer, expert, data, demonstration, comparison.
The practical test is blunt and it works: put your last twenty ads on one screen. If a stranger scrolling past could tell they came from the same brand and the same campaign, that is not a brand-consistency win. That is your retrieval footprint, and it is narrow.
The strongest accounts we work in look almost incoherent laid out flat, and perfectly coherent to each individual person who sees only the two or three ads meant for them.
Why the hook-testing framework degraded

The same CMO described his historical method precisely: build a batch of statics, test a set of hooks, find the winners, then build other vehicles around the winning hook. It worked for years. It is now much harder to run.
The reason is that the method assumes the platform is holding targeting constant while you vary creative. That is no longer the arrangement. With broad targeting and automated audience construction, the creative is a substantial part of the targeting signal. Vary the hook alone and you have varied almost nothing that retrieval can act on, so all five variants get delivered into roughly the same pocket of people, and the differences you measure are mostly noise.
This produces a specific and expensive failure: a clean-looking test, a declared winner, and a scaling attempt that does not reproduce. The winner was a winner within one narrow audience the system had already decided to send everything to.
The replacement is not more hooks. It is testing at the axis level first, then iterating within the axis that showed signal.
An audit you can run this week
This takes an afternoon with an export from Ads Manager and does not require any tooling you do not already have.
Step 1. Export the last 90 days at the ad level. Spend, impressions, and your primary conversion metric per ad.
Step 2. Tag every ad on the five strong axes. Persona, problem framing, format world, register, source of authority. One value each. Do it by hand. It is tedious and it is the entire point: the tagging itself usually ends the argument.
Step 3. Count distinct combinations. Not ads. Combinations. Most accounts spending five figures a month discover they have between three and six, spread across sixty assets.
Step 4. Put spend against each combination. This is where it gets uncomfortable. The usual finding is that seventy to eighty percent of spend sits on one or two combinations, and the tail combinations were never given enough budget to say anything.
Step 5. Find the empty cells with spend nearby. A persona that consumes meaningful impressions but only ever sees one format world is your clearest opportunity. You already know the audience exists, because Meta is spending money reaching it. You have simply never offered retrieval a second way in.
Step 6. Decide production from the gaps, not from the calendar. Most creative briefs are generated by a schedule. This audit generates them from a map.
The volume question, answered properly
"How many creatives per month" is the wrong question and it comes up in every call.
The right question is how many distinct combinations you can support at your spend level, then how many assets each combination needs to reach a readable result. A combination that never accumulates enough impressions to be evaluated is not a test, it is a donation. We work through that arithmetic in detail in creative volume without intent.
The practical consequence for smaller accounts is liberating. If your budget only supports four combinations properly, run four. Four genuinely distinct concepts beat forty variations of one, both in what you learn and in what retrieval can do with them.
Where AI belongs in this

Meta's own numbers on generative tooling are worth reading precisely because they are Meta's: advertisers who turned on Advantage+ creative AI-driven targeting features saw a 22% increase in ROAS, and Meta estimates businesses using image generation see a 7% increase in conversions.
But the strategic value of generative tooling here is not cost per asset. It is that producing a genuinely different format world used to require a different production pipeline. An illustrated explainer, a comic sequence, and an editorial still life were three separate briefs, three vendors, three timelines. That production friction is precisely what pushed teams toward the weak axes: hook rotation was the only variation anyone could afford at volume.
Removing the friction on the strong axes is the actual unlock. Using AI to produce forty near-identical statics faster is the old mistake, executed more efficiently. It also now comes with an AI label attached, which is a good reason to make the generated work count.
Key takeaways
- Andromeda is Meta's retrieval engine. It reduces tens of millions of ad candidates to a few thousand before ranking ever runs, and Meta has publicly stated its roadmap is toward delivering a more diverse set of ad candidates.
- Creative diversity means the number of distinct candidates you contribute, not the number of files you upload. Forty near-identical assets are close to one candidate.
- Vehicle, hook rotation, and cosmetic variation are weak axes. Persona, problem framing, format world, register, and source of authority are where the distance actually comes from.
- The hook-testing framework degraded because creative is now part of the targeting signal. Varying the hook alone delivers all variants into the same pocket of people, which is why winners often fail to scale.
- Audit by tagging your last 90 days on the five strong axes and counting distinct combinations. Most five-figure accounts have between three and six. Brief from the empty cells.
We map creative diversity as a spend-weighted matrix, not a vibe, and we will show you where your account is concentrated. See our creative process



