Insights / Blog

ChatGPT’s Hidden Shortlist: It Picks Brands Before the Search Even Runs

ChatGPT builds a brand shortlist from training data before any web search. New data shows shortlisted brands are cited 33x more often than retrieved-only…
L
Lam Nguyen - Founder
Share
ChatGPT's Hidden Shortlist: It Picks Brands Before the Search Even Runs
ON THIS PAGE

Table of Contents

ChatGPT’s shortlist is built from training data, not from search results. Research published in August 2026 by Search Engine Journal found that in 21 of 27 tested conversations, ChatGPT embedded brand names into its own search queries before fetching a single page. Brands that made the shortlist were cited in the final answer at a rate of 68.9%, compared to 2.1% for brands only encountered during retrieval, roughly 33 times higher.

68.9%Citation rate for brands named in ChatGPT’s own search query vs. 2.1% for brands only retrievedSearch Engine Journal, August 2026
33xHow much more likely a shortlisted brand is to be cited vs. a retrieved-only brandSearch Engine Journal, August 2026
21 of 27Conversations where ChatGPT’s first search query contained brands the user never typedSearch Engine Journal, August 2026

How Does ChatGPT Build Its Shortlist Before Any Search Runs?

When a user asks for “the best AI note-taking app,” ChatGPT rewrites that prompt into search queries of its own. Those queries already contain specific brand names, Granola, Notion AI, Otter, Fireflies, Fathom, Mem, Limitless, that the user never typed and that no web result had yet supplied. The queries sit inside the browser’s response JSON under a field called search_queries (renamed from search_model_queries in early August 2026, according to SEJ’s reporting). Any account holder can read them in browser DevTools in about two minutes.

The implication is structural. ChatGPT is not searching for candidates; it is going down a list it already holds, one name at a time, running a separate targeted probe at each brand’s own website. If your brand is absent from that initial list, your website never gets contacted in that conversation, however well-built it is.

What Does the Data Show About Which Brands Make the Cut?

The SEJ researcher analyzed first search queries across 27 conversations, comparing each user’s opening message against the model’s first query before any retrieval took place. In 21 of those 27 conversations, the query contained brands the user never mentioned. The shortlist is not fixed: the same question phrased slightly differently expanded the list from three brand names to seven in one documented case.

Testing across 12 unrelated product categories showed the behavior holds beyond software. For robot vacuums, ChatGPT recalled specific current model numbers, Roborock Saros, Dreame X50, Eufy S1 Pro, in the first query, unprompted. In the electric SUV category, the model defaulted to automotive publications (Car and Driver, Edmunds, Top Gear) rather than manufacturer sites, while software categories drew it toward vendor pricing pages. That instinct flipped on repeated runs, which suggests it is a tendency rather than a fixed property of any given category.

The researcher also ran 24 queries that deliberately avoided the word “best.” The shortlist injection persisted. What matters, per the analysis, is whether ChatGPT must supply the product names itself. Displacement queries followed the same pattern: “Alternatives to Zendesk” produced Help Scout; “what can I use instead of QuickBooks” produced Zoho Books; “something like Duolingo but better for grammar” produced Kwiziq and Babbel, all in the query, before any page was fetched.

Once You’re on the Shortlist, What Decides Who Gets Cited?

Being named in the query is the entry ticket, not the win. The SEJ researcher built a labeled dataset from 57 conversations and 3,554 retrieved pages. ChatGPT reads around 600 pages per answer and credits roughly 30. Three factors separated cited pages from ignored ones, according to the analysis:

  • Position within the domain’s result group: citation rates collapsed sharply below the top two results. Below that, citation was, in the researcher’s words, “a rounding error.”
  • Number of pages per domain: two tightly matched pages was the sweet spot. Beyond six pages from one domain, per-page citation rates dropped as the domain competed against itself, what practitioners call cannibalization.
  • Relevance threshold, not ranking: cited pages sat in the top 5% for claim relevance but were the single best match only 20% of the time. Relevance qualifies a page for consideration; something else picks the final winner.

The research also documented 86 cases where a brand was recommended without its website being fetched at all in that conversation. Training data alone drove those recommendations.

What This Means for AI-Search Visibility

The most important takeaway from this research is simple: most of what determines whether ChatGPT recommends a brand happens before any website is contacted.

Think of it like a restaurant critic who already has a mental shortlist before walking in the door. She will evaluate the shortlisted restaurants carefully. If your restaurant was never on her list, she will not visit, no matter how good the food is. That is the mechanism described here: a brand absent from ChatGPT’s internal shortlist faces roughly a 2% chance of appearing in a final answer, versus 68.9% for a brand on the list, per SEJ’s August 2026 findings.

In Hingewise’s assessment, this separates AI visibility into two distinct problems that are currently being treated as one.

Problem one is getting into the category vocabulary, becoming the brand ChatGPT associates with a given problem space. Based on the SEJ findings, this appears to be shaped by training data over time: third-party reviews, comparisons, analyst coverage, and consistent category-defining content across the open web. Technical changes, schema markup, page speed, llms.txt files, have no direct path to influencing this layer, because the decision happens before any server is contacted.

Problem two is winning the citation once you are on the shortlist. This is where page-level work applies: one tightly matched page per search intent, the direct answer placed near the top in plain HTML text, and internal cannibalization eliminated by consolidating pages that compete over the same question.

Hingewise’s view: most investment currently labeled GEO is being directed at problem two. The data suggests the larger multiplier sits in problem one. That is not a technical fix. It looks more like a coverage and authority problem, which has different budget implications than an on-page audit.

One methodological note worth flagging: the SEJ findings come from one account across a limited number of conversations. The researcher explicitly states the percentages indicate direction rather than statistical measurement. The underlying mechanism, that shortlists exist and are readable in the response JSON, is independently reproducible. The specific citation rates should be read as directional signals, not benchmarks.

What to Check Before Running an AI Visibility Audit

  • Run your primary “best [your category]” question in ChatGPT five times and read the search queries each time, not just the final answer.
  • Note which brand names appear in the query itself. Those are the brands ChatGPT already associates with your category.
  • Check whether your brand appears in the query consistently, occasionally, or not at all across five runs.
  • Run “alternatives to [top competitor]” and check whether your brand appears in the resulting query.
  • Audit your site for pages competing over the same question intent; identify which to consolidate into a single authoritative page.
  • Confirm your primary answer page places its direct response near the top in plain HTML, not locked inside JavaScript or buried below navigation.
  • Review your third-party footprint: are you being reviewed, compared, and listed in the roundups and analyst pieces your category buyers actually read?

What to Watch Next

The search_queries field in ChatGPT’s response JSON is currently readable by any account holder. OpenAI renamed the field in early August 2026, which signals the structure may continue to change. How long this diagnostic window stays open, and whether the shortlist behavior shifts as OpenAI updates its underlying models, are the variables most worth tracking in the months ahead.

Source: Search Engine Journal, August 2026

By Lam Nguyen, Hingewise

Keep reading

All articles →
Free · no strings

See what AI says about you, today.

Get the report showing how ChatGPT, Gemini & Perplexity answer about your brand.

Get free report →
Reply within 48 hours.