Stay ahead of AI search

The changes that move your rankings and AI citations — and the exact move to make — before they cost you. Free, three times a week.

GEO Beat · DATA

Your website has no OpenAI licensing deal and ChatGPT quotes it like Reuters anyway — and what ChatGPT quotes is a page's opening headline plus the next ~150 characters.

2026-08-11Credibility: Medium 中文版 →

The theory this data retires belonged to one of the researchers now disproving it. In June 2026, SEO researcher Suganthan Mohanadasan reverse-engineered ChatGPT's network traffic and concluded that OpenAI's own search index, tagged "labrador" in ChatGPT's server responses, was a licensed tier reserved for major publishers. He published a correction on 2026-07-14: an Italian reader's free-account captures showed small, non-partner sites coming back through the identical pipeline as Reuters, The Guardian, The Wall Street Journal and Wikipedia. French SEO consultancy Resoneo arrived at the same place from a different direction, capturing 1,249 ChatGPT answers through a browser extension in July 2026 and finding that labrador served sites with no OpenAI content-licensing deal in the same format, length and freshness as pages from licensed partners. Search Engine Journal reported Resoneo's findings on 2026-08-11. What the same dataset does discriminate on is the page itself: of 534 ChatGPT-cited pages Resoneo reviewed, 463 carried an H1 tag, and 387 of those, 83.6%, had that H1 text reused word for word in the citation snippet, which caps out near 200 characters. The evidence also has a date on it. OpenAI stopped tagging search results with the pipeline name around 2026-07-21, closing the exact signal both investigations used to identify labrador, so the finding is verified through mid-July 2026 and cannot currently be re-tested the same way.

Credibility: MEDIUM. Two independent researchers, Resoneo and Mohanadasan, reverse-engineered ChatGPT's network traffic using separate, published methodologies and landed on the same conclusion. That is real corroboration, and it is also where the evidence stops. Neither is an official OpenAI artifact, Resoneo is a commercial SEO consultancy publicising proprietary data partly to promote its Chrome extension and its consulting services, and the pipeline-name signal both investigations relied on has since been removed by OpenAI.

  • Resoneo captured 1,249 ChatGPT answers in July 2026; in the free-account sample, the in-house labrador index handled most of the search results shown.
  • In a separate paid-account, thinking-mode sample of 16,407 search results, Google-scraping supplied about 75% and the in-house labrador index about 24% (Resoneo, published at think.resoneo.com/chatgpt-retrieval/).
  • Of 534 ChatGPT-cited pages Resoneo reviewed, 463 had an H1 tag, and 387 of those (83.6%) had the H1 text reused verbatim in the citation snippet, which is capped near 200 characters.
  • Mohanadasan retracted his June 2026 "licensed tier" reading on 2026-07-14, after an Italian reader's free-account captures showed small, non-partner sites cited through the identical pipeline as Reuters, The Guardian, The Wall Street Journal, and Wikipedia.
OPERATOR TAKE

For anyone weighing an OpenAI content-licensing deal as the way into ChatGPT's free-tier answers: on this evidence, that is not the lever. We flagged the multi-pipeline architecture underneath ChatGPT's search on 2026-07-08, when the labrador pipeline supplied 88.1% of primary search sources across roughly 10,000 runs and switching sources changed which URLs got cited. Resoneo's July 2026 capture settles what labrador is: OpenAI's own in-house index, not a third-party retrieval vendor. It also adds the part an operator can use. Through mid-July 2026, it served sites with no licensing deal in the same format, length and freshness as licensed partners like Reuters and The Guardian.

What is confirmed instead is a mechanism, and the mechanism is unglamorous. The citation snippet is built from a page's H1 tag plus roughly the next 150 characters, capped near 200 characters in total. So if an H1 states a topic rather than a specific claim or number, the topic is what ChatGPT quotes. If a template prints a kicker, a byline and a date block above the first paragraph, that boilerplate fills the characters after the H1, and the boilerplate is what ChatGPT quotes. Licensing status never enters the calculation.

A licensing deal can still be the right call for reasons this dataset never tested: paid-tier or thinking-mode treatment, where Google-scraping supplied about 75% of results in Resoneo's sample, or contractual and IP protection. Do not use this finding to close a conversation that is happening for those reasons. And keep the expiry attached to the claim. OpenAI removed the tag both investigations read, so the honest position is that this held through mid-July 2026 and nobody can currently re-run the test.

Recommended actions:

  • Audit the H1 tags on your most commercially important and most-cited pages, and rewrite any H1 that states a topic instead of a specific claim or number. That is the first ~50 characters ChatGPT's citation snippet reuses.
  • Ask your web team to check what renders immediately after the H1 in your page template, and trim any kicker, byline or date block eating into the ~150 characters after it that the snippet also captures.
  • If an OpenAI content-licensing conversation is live for reasons other than free-tier ChatGPT citation eligibility, record that rationale separately. Resoneo's dataset does not test paid-tier or IP-protection scenarios.
Sources

Search Engine Journal: "ChatGPT's Search Index Serves Small Sites Too, Data Shows" by Matt G. Southern, August 11, 2026. https://www.searchenginejournal.com/chatgpts-search-index-serves-small-sites-too-data-shows/585388/

Resoneo: "ChatGPT Retrieval" report (primary data — browser-extension capture of 1,249 ChatGPT answers, July 2026). https://think.resoneo.com/chatgpt-retrieval/

Suganthan Mohanadasan: "How ChatGPT Picks Sources, Part 2" (correction published July 14, 2026). https://suganthan.com/blog/how-chatgpt-picks-sources-part-2/

Prior Novastacks GEO Beat coverage (arc origin): "ChatGPT citations shift when its hidden search pipeline switches — 11.6% of prompts changed primary source across repeated runs," GEO Beat run 2026-07-10.