The short answer
GEM (Generative Ads Recommendation Model) is Meta's largest ads foundation model. It does not rank ads directly: it trains the ranking models that predict whether a person will act on your ad. You can't tune it, but you can feed it better predictions. On our board, ads with strong proof reach the top tier at 16% vs about 10% overall.
What GEM is
GEM stands for Generative Ads Recommendation Model. Meta's engineering team calls it its "most advanced ads foundation model," built on an LLM-inspired design and trained across thousands of GPUs. It is the largest foundation model for recommendation that Meta has disclosed.
The important detail for advertisers: GEM is a teacher. It does not score each ad impression itself. Meta transfers what GEM learns into the smaller production models that do the ranking, using knowledge distillation, representation learning and parameter sharing. Meta says this transfer is about twice as effective as standard distillation.
The reported effect, in Meta's own numbers: a 5% increase in ad conversions on Instagram and 3% on Facebook Feed in Q2 2025. An August 2026 follow-up on the user-sequence models at GEM's core reports a cumulative 6% conversion lift on Instagram and 3% on Facebook. These are platform-wide averages, not a promise for your account.
Where GEM fits next to Andromeda
Andromeda decides which few thousand ads are even considered for a person. The ranking models decide the order of that shortlist by predicting what the person will do. GEM makes those ranking models smarter. Then the auction combines your bid with the predicted action rate and ad quality.
So when a media buyer says "GEM likes this ad," what it really means is that the ranking models predict a high chance this person converts on this ad. That prediction is what you are competing on.
What GEM looks at
Meta's post lists two kinds of inputs:
- Sequence features: a person's activity history, in order. What they viewed, clicked and bought, across Meta's apps and from the advertiser signals Meta receives.
- Non-sequence features: attributes of the person and of the ad, including age, location, ad format and the "creative representation."
The second list is where your work shows up. The model has a representation of your creative, it knows the format, and it learns which kinds of creative lead which kinds of people to act. The first list is where your conversion data shows up. Purchases and other events you send are what the model learns from, which is why Conversions API signal quality matters as much as the ads.
What you can and cannot influence
| You cannot touch | You control |
|---|---|
| The model architecture and training | The creative: concept, opener, format, length |
| How GEM's learning reaches ranking | The variety of creatives you give retrieval and ranking |
| A person's activity history | The conversion event you optimize for and how cleanly you send it |
| Auction competition | Bid strategy and budget |
| Any "GEM setting" (there is none) | The landing page that turns a click into the event the model learns from |
Most advice about "optimizing for GEM" is advice about the right column. That is fine, as long as nobody pretends there is a hidden switch.
Which creative traits over-index among winners
We can't see Meta's predictions. We can see which ads survive. Our October 9, 2026 snapshot has 2,443 live ads with $50k+ in estimated spend, and about 10% of them sit in Ad Radar's top tier (top 10% by winner score, live 21+ days). The table shows which tagged traits reach the top tier more or less often than that baseline.
| Creative trait | Ads | Share reaching top tier |
|---|---|---|
| Strong proof (demonstrations, data, named experts) | 104 | 16.3% |
| Image ads | 475 | 13.1% |
| Specific number opener | 411 | 12.9% |
| Video, 60 to 120 seconds | 483 | 10.6% |
| Board average | 2,443 | 9.6% |
| Question opener | 366 | 7.1% |
| Video, 300 seconds or longer | 105 | 5.7% |
| Video, under 15 seconds | 87 | 3.4% |
Two readings. First, proof travels well. Ads tagged with strong evidence are only 4% of the board, but their median estimated spend is $381K against $223K overall. A ranking model that predicts purchases has every reason to favor ads that resolve doubt. Second, very short and very long videos are rarer in the top tier. Under-15-second videos reach it 3.4% of the time, though with 87 ads the sample is small.
None of this proves what GEM rewards. Top tier status uses public signals (days live, variants launched, engagement growth), not Meta's internal scores. It does show what keeps getting funded once the ranking system has had its say.
Three ads that give the ranking models a lot to work with
I'm the founder of Mudwater, and a lot of people ask, what is the difference between Mudwater and mushroom coffee?
- Est. spend
- $7.4M
- Days live
- 69
- Format
- Video, 118s
- Variants
- 5
A founder answering the comparison question buyers already have. Strong proof, a clear reader, five variants in 69 days.
Forty bags were stolen from the carousel at Denver International in a single year.
- Est. spend
- $4.1M
- Days live
- 126
- Format
- Video, 85s
- Variants
- 6
A concrete number and place in the first line. It targets frequent flyers without any interest targeting, because only they care.
So you've started baking your own bread.
- Est. spend
- $988K
- Days live
- 334
- Format
- Video, 52s
The opener names a specific behavior, so the people who stop are already the buyers. Live almost a year with a single version.
What these share is clarity about who the ad is for. A model predicting conversions does better when the creative attracts a narrow, consistent type of person, because the engagement pattern it learns from is cleaner.
What to do with this
- Stop looking for GEM hacks. There is no toggle. Spend that energy on the inputs in the right-hand column above.
- Make each ad's reader obvious in the first line. A named behavior ("you've started baking your own bread") or a specific number helps both the person and the model sort quickly. Our hook pattern vs hook mechanism guide breaks down opener types.
- Put proof inside the ad. Demonstrations, comparisons and named experts are a small share of the board but over-index in the top tier. See social proof ads for formats that work.
- Fix your signal before you judge creative. If purchases arrive late, duplicated or unmatched, the ranking models learn from noise and your test results will be noise too.
- Give ranking real options. Different concepts, not ten cuts of one. Our creative diversification playbook shows how to plan the mix.
If you want to see which traits over-index in your own niche, Ad Radar's winning_patterns tool in the MCP server summarizes the tags of winners for the filter you give it, such as a niche or a format.
Figures marked as estimated spend come from Ad Radar's model of engagement on public Meta Ad Library ads. They are estimates, labeled as such, and are best used to rank ads against each other.