Stay ahead of AI search

The changes that move your rankings and AI citations — and the exact move to make — before they cost you. Free, three times a week.

GEO Beat · DATA

ChatGPT, Copilot, Google and Perplexity cite almost entirely different pages for the same question — check each engine before concluding a brand is missing from AI answers

2026-09-19Credibility: Medium 中文版 →

Ask ChatGPT, Microsoft Copilot, Google and Perplexity the same buying question, then compare any two of their answers: most of the time, the two will not share a single cited URL. That is the measured result of an observational audit that independent researcher Benjamin Tannenbaum posted to arXiv (cs.IR) on 19 September 2026. On 6 June 2026 he ran 15 fixed commercial prompts through all four engines and logged 589 citation observations: 528 unique URLs across 356 domains. The overlap figures sit close to zero at every level the paper measured. Mean same-prompt URL overlap between engine pairs, scored as Jaccard similarity, was 0.0079, with a median of zero. In 84.9% of engine pairs, not one cited URL was shared. And 96.4% of URLs appeared on only one engine, against the 12.7% overlap a matched random baseline would predict. A single engine captured between 11.4% and 42.6% of the combined four-engine URL set. At the top of the answer the split is complete: top-five exact-URL overlap was zero in all 60 pairwise comparisons. A second check found one engine unstable against itself as well. Comparing 5 and 6 June 2026, that engine's URL set turned over by 67.0% on average from one day to the next. The paper's stated conclusion is aimed at a specific kind of tool: engine-free "page scores" measure query-page fit only; they leave out the separate exposure and citation-selection steps that decide whether a page is actually cited on any one engine.

Sources: arXiv (cs.IR), Tannenbaum, B., "Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores" (full citation below).

Credibility: MEDIUM. This is a single-author arXiv preprint that has not been through peer review, and the sample is a one-day, 15-prompt snapshot plus a two-day same-engine check. No other outlet or researcher has corroborated it yet. It clears the artifact bar because the method, the prompt count and the full statistics are published in the open and can in principle be reproduced.

Fact bullets:

  • 6 June 2026: 15 fixed commercial prompts run across ChatGPT, Copilot, Google and Perplexity produced 589 citation observations (528 unique URLs, 356 domains).
  • Cross-engine overlap: mean pairwise Jaccard similarity 0.0079; 84.9% of engine pairs shared zero cited URLs; top-five URL overlap was zero in all 60 pairwise comparisons.
  • A single engine captured only 11.4%–42.6% of the four-engine combined URL set; 96.4% of observed URLs appeared on just one engine.
  • A separate same-engine comparison (5–6 June 2026) found 67.0% mean URL-set turnover day-to-day on one engine alone.
OPERATOR TAKE

For marketing leaders who report an AI-visibility number to a CEO or a client, the finding is narrower than the paper's title suggests, and more useful. What it tests is whether a page score predicts citation on a given engine. A grade computed without asking the engines cannot say which engine will cite a page, and four engines given the same question return four largely separate citation lists. A tracker that queries each engine and reports each one on its own line is doing the measurement this data calls for. A blended figure still works as a leading indicator of direction. It cannot carry the sentence "we are cited" or "we are not cited" by itself, because 96.4% of the citations Tannenbaum logged lived on exactly one engine.

The more expensive mistake runs the other way. A team that asks ChatGPT its top buying question once, finds no mention, and declares an AI-visibility problem has sampled a single engine on a single day. In this dataset 84.9% of engine pairs shared nothing, and one engine's own citations turned over 67.0% between two consecutive days. That "no" is a single draw from four moving lists. (The caveat cuts against this brief too: the whole dataset is 15 prompts run on a June day by a researcher whose paper nobody has reviewed yet. Trust the direction more than the decimals.)

Recommended actions:

  • Open the last AI-visibility report that went to your CEO or a client and check whether citations are broken out by engine (ChatGPT, Copilot, Google AI Overviews, Perplexity); if it shows one blended number, ask the tracker or vendor for the per-engine split before the next reporting cycle.
  • This week, have the team re-run the top 10–15 buyer-intent prompts on all four engines, on more than one day, before anyone calls a missing citation a site-wide problem.
  • When an aggregate visibility score goes to a client or CEO this month, record the decision rationale next to it: the number is directional, and it travels with the per-engine citation list.
Sources

arXiv (cs.IR): Tannenbaum, B. "Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores." Submitted 19 Sep 2026. https://arxiv.org/abs/2609.22655