← All posts

Why AI Answer Engines Change Their Mind About Your Brand Overnight

AI brand mentions don't hold steady. An engine can cite your brand in one answer and drop it from the next, even when nothing about the prompt changes. We tested this across five engines by running identical prompts multiple times each, then measured how much the cited sources actually held steady between runs. The results varied more by engine than we expected.

The Same Question, Five AI Answer Engines, Five Different Answers

Ask ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews the same question, and you're not just getting five different writing styles. You're often getting five different source lists, and the gap between engines is bigger than the gap between prompts.

EngineOverlap between repeated runs
Perplexity~85%
Claude~66%
GPT-4o-mini + web search (API)~62%
Google AI Overviews~59%
Gemini~42%

Overlap here means the share of a run's cited domains that also appeared in the smaller of the two runs being compared, not raw Jaccard similarity. Jaccard overlap for these same engines runs lower (see methodology below).

Perplexity held up best. Its citations overlapped about 85% between runs, meaning most of what it cited the first time, it cited the second and third time again. Gemini went the other way, with only about 42% overlap between runs, which means asking it the same thing twice gets you a mostly different source list both times.

Gemini's citations moved the most of any engine we tested. Ask it the same question twice, and the source it leans on first is different more than 9 times out of 10, the lowest consistency of any engine in the study. If a brand shows up once in a Gemini answer, that's no guarantee it shows up again.

What Causes Citations to Change Between Runs

Nothing here is cached. Every run pulls a fresh retrieval pass off the live web, and retrieval itself carries some randomness, so a word-for-word repeat of the same prompt can still surface a different mix of pages. Nothing about the wording changed. Only the timing did.

That's the mechanism behind the numbers above. A single check only tells you what an engine said once, not what it usually says.

What This Means If You're Only Checking Once

Most brand-monitoring tools run a prompt one time and call that the result. Given what we found on Gemini, a single check has decent odds of showing you a citation set that won't repeat the next day.

Other platforms in this space typically publish that kind of one-time snapshot. What's different here is that we ran the same prompts several times per engine and reported the overlap between runs as an actual number. A one-off check can mislead you. Watching ChatGPT brand mentions and citations on other engines over several runs is what tells you whether something's a real pattern or just a fluke. That includes how we track Google AI Overviews mentions, which behaves differently from the chat-based engines above.

If you want this kind of run-over-run tracking applied to your own brand, Yogoo AI's ChatGPT and Perplexity monitoring checks run on a repeating schedule instead of a single pass.

How We Tested This

Every prompt ran at least three times per engine, spaced hours apart, worded the same each time. We pulled the cited domains from every run and compared the sets to get an overlap percentage per engine.

ChatGPT and Claude also flicker between citing something and citing nothing at all. Roughly a quarter of the prompts where either model cited anything didn't get a citation in every run. The overlap percentages above are calculated only on runs that actually cited a source.

GPT-4o-mini + web search (API) is an API version of ChatGPT with live retrieval turned on. It isn't the ChatGPT app, so these numbers describe that specific setup, not necessarily what the consumer product would return on the same prompt.

The full breakdown of how we built the original prompt set is in How to Measure AI Visibility. For a deeper look at what gets cited across all five engines, see our 500-prompt citation study.