The short answer
There is no reliable public benchmark that shows ChatGPT or Claude writes better ad copy, and both change every few months. What decides the output is the input: real examples, hard constraints and a clear job. Winning Meta ads give you those constraints. On our board the median video opener is 11 words and the median headline is 6. Put numbers like those in the prompt and either model gets usable.
The honest answer first
We looked for a trustworthy benchmark of ChatGPT against Claude on ad copy and didn't find one. The public comparisons are mostly one person, one prompt, one afternoon, judged on taste. Both products ship new models several times a year, so even a careful test goes stale fast.
What we can say with confidence comes from the ads themselves. The 2,443 winning Meta ads on our board share patterns in length, voice and structure. A model that's given those patterns writes closer to a winner. A model that isn't writes the average of the internet. That gap is far larger than the gap between the two tools.
So this article does two things: lists the differences between the tools that are real and checkable, then shows how to prompt either one so the output is usable.
Differences that actually affect ad work
These are workflow differences, not quality claims.
| What matters | ChatGPT | Claude |
|---|---|---|
| Saved instructions and files per client | Projects and custom GPTs | Projects |
| Connecting live data (ad platforms, ad databases) | Supports MCP | Supports MCP (Anthropic created it) |
| Meta's official ads connector | Works with it | Works with it |
| Coding agent for bulk jobs (CSV in, 200 variants out) | Codex | Claude Code |
| Image generation in the same chat | Yes | No native image generation |
The MCP row is the one that changed in the last year. Anthropic created the Model Context Protocol and in December 2025 donated it to the Linux Foundation's Agentic AI Foundation, noting it had been adopted by ChatGPT, Gemini and Microsoft Copilot among others. In April 2026 Meta opened its own ads connector to both assistants (PPC Land). For a copywriter this means the choice of model no longer locks you out of data. More on that in MCP for marketers.
If your team already pays for one, start there. Switching tools is rarely the fix for weak copy.
What winning copy looks like, in numbers
Here are constraints you can put straight into a prompt. They describe the ads on our board on October 9, 2026, every one live for 30+ days with $50k+ in estimated spend.
11 words
median video opener (first sentence)
6 words
median headline
67 words
median primary text
19%
headlines that use an emoji
And the voice:
| Opener feature | Share of video openers |
|---|---|
| Uses "you" or "your" | 36% |
| Uses "I", "my" or "me" | 32% |
Two thirds of openers speak directly to the viewer or from a first-person narrator. Neither model defaults to that. Left alone, both drift toward third-person, brand-voice lines like "Introducing a revolutionary way to..." which barely appear on the board.
One more number worth knowing: opener length is not a magic dial. Openers of 11 to 15 words reach the top tier 11.4% of the time, but 21+ word openers do 10.3% and 6 to 10 words do 8.5%. Length is a constraint to keep drafts in the normal range, not a lever. (About 10% of the board sits in the top tier: the top 10% by winner score, live 21+ days.)
Why examples beat instructions
You can describe a voice in a paragraph, or you can show it. Showing works much better with both models.
The guys on the crew call me snacks, I laugh along with the joke, but every damn time I hear it lands like a punch to the throat.
- Est. spend
- $249K
- Days live
- 51
- Format
- Video, 126s
A top-tier ad written as a rhyming first-person monologue from a tradesman. No instruction like "authentic blue-collar voice" gets a model here. Pasting this script as the example does.
The same holds for long-form text ads. Describe "a story-led ad" and you get a story template. Paste a real one and the model picks up the rhythm: short declarative lines, a specific number early, a turn in the third paragraph.
Read this if you're a grandparent☝️
- Est. spend
- $499K
- Days live
- 142
- Format
- Image
- Variants
- 2
The primary text opens on a teacher, a $15,000 tuition figure and one classroom question. That's three concrete details in the first two sentences. Give a model this as the reference and ask it to match the density of specifics, not the topic.
A prompt that works in either model
Use the same prompt in both. If you want to compare, this is also your test prompt.
ROLE: You write Meta ad copy for [brand]. You are not the brand's
voice; you write as the person in the ad.
PRODUCT FACTS (use only these): [facts, price, offer]
CLAIMS YOU MAY NOT MAKE: [list]
AUDIENCE: [who], awareness level: [unaware/problem/solution/product]
REFERENCE ADS (structure and density to match, never copy wording):
1. [full script or primary text, with est. spend and days live]
2. [second example]
TASK: 5 primary texts and 5 headlines.
CONSTRAINTS:
- Openers 8 to 15 words, speaking to "you" or from "I".
- Headlines 4 to 8 words.
- Primary text 50 to 90 words.
- A specific number or concrete detail in the first two sentences.
- Each primary text uses a different hook pattern: problem/benefit,
number, first-person story, question, warning.
- No: "introducing", "revolutionary" or other launch-speak, no
rhetorical questions after the opener, no exclamation marks.
OUTPUT: a table with hook pattern, opener, full text, headline.
Three parts do the heavy lifting: the facts list (stops invented claims), the reference ads (sets the voice), and the "different hook pattern" rule (stops five rewrites of one idea). For a full set of single-purpose prompts, see Claude prompts for Facebook ads. The hook patterns themselves are covered in video ad hook formulas.
Where each tool tends to need correcting
We are not going to claim one is better. Drafts from either tool commonly need the same four corrections, and each can be written into the prompt:
- Over-polish. Both default to tidy, complete sentences. Winning scripts are full of fragments ("False." "Wrong." "Nope."). Ask for spoken rhythm and paste an example.
- Hedging. Both soften claims you never asked to soften, or harden claims you can't support. The facts list fixes both directions.
- Sameness across variants. Ask for ten hooks and you often get one hook in ten outfits. Name the patterns.
- Brand voice creep. Both slide back into "we" and "our". Tell the model who is speaking.
I'm about to destroy the biggest myths in the hair loss industry, starting with the one you've most probably read about all over socials.
- Est. spend
- $686K
- Days live
- 145
- Format
- Video, 366s
- Variants
- 2
The script runs as myth, one-word rebuttal, myth, one-word rebuttal. Fragment rhythm like this is what both models smooth away unless the example is in the prompt.
Run your own test in an afternoon
If you want to know which model suits your team, test it on your work, not on a blog post's.
- Pick 10 real briefs from the last two months, including ones that became winners.
- Write one prompt (the template above) and run it unchanged in both tools, same reference ads, same day.
- Strip the labels. Paste outputs into a sheet as A and B, shuffled per brief.
- Rank blind. Two people, separately: which draft would you launch? Note why.
- Launch the top picks as normal creative tests, with the same budget and kill rule as any other ad. See how many creatives to test for sizing.
- Judge on results. Preference in the blind round tells you what your team likes. Spend and CPA tell you what works.
You can also use both: one drafts, the other critiques with a "you are the media buyer who will reject this" prompt. A second model reading cold catches things the first one smoothed over.
Where the examples come from
Everything above depends on having real, recent winning ads to paste as references. Pull them from the Meta Ad Library by hand, or from a database that already filters for spend and runtime. If you use Claude, the Ad Radar MCP server can fetch the reference scripts directly with search_ads and get_ad; any MCP-capable client can connect the same way.
Figures marked as estimated spend come from Ad Radar's model of engagement on public Meta Ad Library ads. They are estimates, labeled as such, and are best used to rank ads against each other.