We Ran 500 Prompts Across 5 AI Engines: Here's What Actually Gets Cited
The results from the 500-prompt study answer the same kinds of questions brands and marketers ask when they are trying to figure out how to get mentioned by AI in the first place. Queries such as: which tools track AI visibility, how to monitor brand mentions across ChatGPT and Perplexity, how to improve AI search rankings, and similar questions about AI SEO, AEO, and GEO topics.
From those 500 prompts, we ran through the queries across ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews to find out what actually works for AI search optimization. What we’ve found out is that the citation behavior varied by more than a factor of two across engines: one engine backed up barely a third of its answers with a source, another backed up every single one; additionally, the type of question asked changed which sources got pulled in. This is what every brand and marketer should remember: which engine you’re targeting changes your odds more than you expect.
Methodology
The test covered 500 prompts across 12 different topic clusters and 6 intent types, spanning the full path someone takes from “Can I get alerts when ChatGPT mentions my brand?” to the fundamentals “Can you explain what AEO means?”. Some prompts named Yogoo AI directly, some named other AI visibility tracking tools, while the majority were unnamed, phrased the way an actual buyer would type them.
We ran each prompt at least three times on every engine, spaced hours apart over a roughly 20-hour window, with no variation in wording between runs. This then produced a total of 8,695 answers to work with. Citations were pulled from every answer at the moment it was captured, then normalized to the domain, with subdomains counted separately (blog.hubspot.com and hubspot.com are treated as two different sources), so a source only counts once per answer.
Five engines were tested: ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews. Three of these were run as API-based approximations with live web search enabled, not the consumer apps themselves: ChatGPT’s engine is GPT-4o-mini with web search, Claude’s is Claude Haiku 4.5 with web search, and Gemini’s is Gemini 2.5 Flash with web search. Perplexity used Sonar, and Google AI Overviews came from a real SERP artifact, not an API approximation.
This whole test ran through Yogoo’s own generative engine optimization tracking pipeline, the same infrastructure the product uses to monitor live citations, just pointed at a fixed prompt set instead of a customer’s ongoing tracking list.
What Gets Cited? Patterns Across Engines
Each engine behaves differently in how often they back their answers with a source, and how many sources it typically uses:
| Engine | % of answers citing a source | Avg. sources per answer |
|---|---|---|
| Perplexity (Sonar) | 100% | 17.8 |
| Google AI Overviews | 97.5% | 7.6 |
| Gemini (2.5 Flash + web search) | 96.8% | 10.4 |
| Claude (Haiku 4.5 + web search) | 71.2% | 3.3 |
| ChatGPT engine (GPT-4o-mini + web search, API) | 35.5% | 2.3 |
Perplexity cited a source in every answer. The ChatGPT engine (GPT-4o-mini with web search enabled via API), by contrast, backed up its answer with a source in barely a third of cases; it used the fewest sources of any engine.
Topic is also a big factor in how the engine answers the queries. Comparison-style prompts, like choosing between citation tracking platforms, pulled a different mix of sources than fundamentals-style “what is AEO” prompts, and how-to prompts skewed toward different domains again. If you want the mechanics of how Yogoo AI specifically divides and measures AI visibility across the five engines, you can check our methodology to see what that looks like.
Going back to what actually got cited: no domain showed up in every engine’s top 10. Google AI Overviews and Perplexity leaned heavily on user-generated content, YouTube and Reddit especially. Claude and Gemini favored niche vendor and blog content. GPT-4o-mini with web search via API leaned toward established tech media. To compare the five engines side by side, each one is pulling from a different pool of sources and not one shared list. This goes back to what we said earlier: if you’re building content to become an AI answer engine’s go-to source, the strategy varies depending on which engine you’re aiming at.
Another piece of data we’ve gathered is that we found the most cited domains that all five AI engines mentioned. Overall, YouTube and Reddit compete at the top with around 20% of the answers, while siftly.ai and therankmasters.com sit at the bottom with less than 10% of the answers.
Most-cited domains across all five engines:
| Rank | Domain | % of answers |
|---|---|---|
| 1 | youtube.com | 22.9% |
| 2 | reddit.com | 21.5% |
| 3 | tryprofound.com | 15.8% |
| 4 | linkedin.com | 12.3% |
| 5 | rankability.com | 9.4% |
| 6 | semrush.com | 9.1% |
| 7 | frase.io | 8.1% |
| 8 | hubspot.com | 7.9% |
| 9 | siftly.ai | 7.7% |
| 10 | therankmasters.com | 7.4% |
To see it closely side by side for each engine, we gathered data that showed the top five cited domains across five engines.
Top 5 cited domains, by engine:
| GPT-4o-mini (API) | Claude | Gemini | Google AI Overviews | Perplexity |
|---|---|---|---|---|
| techradar.com — 8.7% | rankability.com — 10.3% | tryprofound.com — 20.1% | youtube.com — 53.6% | reddit.com — 64.0% |
| semrush.com — 4.2% | tryprofound.com — 8.2% | semrush.com — 15.2% | reddit.com — 30.7% | youtube.com — 52.1% |
| blog.hubspot.com — 3.7% | stackmatix.com — 7.9% | hubspot.com — 15.2% | linkedin.com — 20.9% | linkedin.com — 37.3% |
| linkedin.com — 3.2% | frase.io — 7.3% | medium.com — 14.8% | tryprofound.com — 13.4% | tryprofound.com — 36.0% |
| surferseo.com — 3.1% | therankmasters.com — 7.1% | llmpulse.ai — 14.6% | otterly.ai — 9.2% | blog.hubspot.com — 25.7% |
We can see that each AI engine has different sources of its cited domains; one site might rank as a top source of one engine while it might rank lower for the other. Tryprofound, for example, is the top 1-rated domain by Gemini, but only the fourth most-cited source for Google AI Overviews. Meanwhile the same two platforms — YouTube and Reddit — occupy the top two spots in both Google AI Overviews and Perplexity, just in flipped order.
On another note: the gap between the ChatGPT engine’s behavior in this study and the consumer app is itself a finding worth sitting with. A same-prompt test run separately on the consumer ChatGPT app, which runs a frontier model, returned a completely different citation set than the API-driven engine in this study. The most likely explanation: smaller, retrieval-driven models tend to reward precise, exact-match on-page copy, pulling directly from pages that answer the question cleanly. Frontier models lean more on brand authority signals, things like roundup listicles, review coverage, and backlinks, rather than matching a page’s wording directly. Citation sets shift with model tier. There is no single “AI search” to write for, even within one company’s product line.
Where AI Search Optimization Tools Fit
Running a test like this at any real scale isn’t something you do by hand. Tracking 8,695 answers across five engines, extracting every cited domain, and normalizing it all down to something you can actually compare needs dedicated AI search optimization tools built for exactly this job, not a spreadsheet and a lot of patience.
That’s also the practical takeaway for anyone trying to apply this outside a one-off study: citation behavior changes often enough, and differs enough by engine, that a single manual check tells you very little. Ongoing monitoring is what turns a snapshot like this into something you can act on. If you want to go deeper on how a visibility score is actually calculated from citation data like this, we captured this data in one single score for each site that can be tracked. If you’re curious about it, you can check how Yogoo Score captures these different elements and factors to see where you rank.
Key Takeaways for Brands
1. Citation behavior varies in each AI engine. One engine cited a source in every answer; another did it only about a third of the time. If you want to target being mentioned by a specific AI engine, strategize content based on the engines your buyers or consumers actually use.
2. Match content to the engine’s source. Engines leaning on UGC reward different content than engines leaning on blogs or established media. This matters to your content by checking what a given engine actually cites in your category before optimizing for it.
3. Exact-match or direct-answer copy helps with retrieval-driven engines. Smaller models pulling from live web search tend to reward pages and content that answer the question precisely and concisely right from the start.
4. Authority signals matter more for frontier models. Roundups, reviews, comparison content, and backlinks carry more weight with larger models that draw on broader training knowledge.
5. The prompt changes the outcome. A comparison prompt and a how-to prompt on the same topic can pull from different sources entirely. Content built for one intent won’t automatically perform for another.
What’s Next?
This was a single 20-hour snapshot, not a trend line. Citation behavior on retrieval-driven engines especially will keep shifting as models update and web search behavior changes. If you want to track how this looks for your own brand over time rather than as a one-time study, learning about AI visibility is a good next stop.