info.link logo
Go back
Citations & Retrieval

ChatGPT API vs UI: do they return the same citations?

Across 52,170 chats on 555 prompts, average brand visibility differed 41% between the ChatGPT interface and the API. The sampling is unusually strong at 47 runs per prompt per surface. The configuration is undisclosed, and setting confounds do not average out the way random error does.

Across 555 prompts run 47 times on each surface, average brand visibility differed by 41% between the ChatGPT interface and the API Landwehr · Peec AI · 2026. The most visible brand in the interface did not reach the API top three. The figure comes from Peec AI's Malte Landwehr. It is a single vendor study with no disclosed configuration; treat it as provisional.

The divergence itself is not in doubt, and the sample is an unusually large one for this question. What the study cannot establish is how much of the gap belongs to the surface and how much to the settings. That distinction decides whether the number means anything to a brand team.

The 41% gap is real, and the sampling behind it is unusually strong

The design runs 555 prompts 47 times on each surface, which is 52,170 chats Landwehr · Peec AI · 2026. That is well past the point where sampling noise explains a gap this size. Reliable citation-share estimates need roughly 40 to 50 queries on Gemini and 150 or more on SearchGPT Sielinski · IQRush · 2026. Typical 95% confidence intervals in that work spanned 3 to 6 percentage points.

The reported differences are large rather than marginal. Competitor domains rose from 38% to 52% of citations between the two surfaces Landwehr · Peec AI · 2026. Listicles fell from 25% to 4%, and product pages and how-to guides grew from 34% to 63%. Movements of that size do not come from a sampling artefact at 47 repeats.

Forty-seven runs per prompt already answers the instability objection

The most common objection to the study is that AI answers are unstable against themselves. As a general matter the objection is well founded. Within-prompt resampling explained 34.8% of outcome variance across 12,933 responses, while brand identity explained just 0.7% Żatuchin · Estonian Entrepreneurship University of Applied Sciences (EUAS); Rankfor.AI · 2026. A single run on either surface is one draw.

The objection does not land on this particular design. The same decomposition found that repeats past the fifth added almost nothing Żatuchin · Estonian Entrepreneurship University of Applied Sciences (EUAS); Rankfor.AI · 2026. A separate framework recommends a minimum of 5 to 10 repeated runs per evaluation Mustahsan · arXiv · 2025. At 47 runs per prompt per surface, the study clears both thresholds several times over. Repetition is the one thing it did thoroughly, and the argument around it has largely gone unmade.

Repetition removes random error, not a setting

A setting confound is systematic, so averaging more runs does not remove it. Suppose the two surfaces ran different model variants, or one had web search enabled and the other did not. Every one of the 47 repeats then inherits that difference identically. The study discloses no temperature, no web-search setting, no interface account state and no model variant Landwehr · Peec AI · 2026.

This is the distinction the surrounding argument keeps collapsing. Run-to-run variance is a sampling problem that repetition fixes. Configuration divergence is an identification problem that repetition cannot touch. Treating generative engine optimisation as a stochastic, partially observable pipeline captures the first problem and not the second Martinez · arXiv (cs.IR) · 2026. A partially observable system still has to be observed under stated settings.

GPT-5.6 is three models, so "both are GPT-5.6" does not name the comparison

The study states that both surfaces run GPT-5.6. OpenAI documents three models under that name, with separate API identifiers and three price tiers OpenAI · 2026. Sol runs at five and thirty dollars per million input and output tokens, Terra at 2.50 and fifteen, and Luna at one and six. Sol is the flagship, Terra a lower-cost option, and Luna the fastest and most cost-efficient.

Which one the interface serves depends on the account. GPT-5.6 Luna became the default model for Free and Go users OpenAI · 2026. OpenAI states separately that the Chat build of Sol is distinct from the build powering Work and Codex. A comparison that does not name the variant on each side may be measuring model against model in part. Naming the variant is a documented requirement rather than a pedantic one.

A web-search setting can manufacture a citation collapse

Several of the study's reported differences are citation-composition findings Landwehr · Peec AI · 2026. Those are vulnerable to a single unstated parameter. An ungrounded API call retrieves nothing, so it returns no citations at all. A near-total fall in user-generated and editorial sources on the API side is consistent with a genuine retrieval difference. It is equally consistent with web search being switched off on one side.

The study does not say which. Rankfor.AI's Dmitrij Żatuchin reported a small check in the comment thread, running six prompts six times on gpt-5.6-terra with and without the web-search tool. That is an unverified test by a competitor of the study's author, so it belongs in the question rather than the answer. The question stands on its own. A citation-composition finding is interpretable only once the grounding configuration is stated for both arms.

No published API fan-out count should anchor a decision yet

The study reports a much lower fan-out ceiling in the API than in the interface Landwehr · Peec AI · 2026. A separate analysis ran about 4,000 prompts through the GPT-5.6 Sol API Long · Nectiv · 2026. It reports an average of 7.61 fan-out queries per prompt, with the longest single chain reaching 29 searches. That study is a single vendor analysis; treat its figures as provisional. The two accounts of API fan-out cannot both describe the same instrument, and neither is reconciled.

OpenAI's own documentation supplies the likely reason. The web_search_call output item will usually, but not always, include the search queries that were searched OpenAI · 2025. Fan-out observed through the API is therefore partial by the platform's own account. An API-side count may be a floor rather than a full count.

On that reading, a low API fan-out figure measures what the API reveals rather than what the model did. The disagreement is a counting problem before it is a behavioural finding.

The model names brands before it searches, so neither surface isolates retrieval

The argument treats the API as the model and the interface as the model plus retrieval. Network-traffic analysis complicates that split. In 21 of 27 conversations, the model's first search query already contained brand names the user never typed Mohanadasan · Snippet Digital; Keyword Insights · 2026. No page had been fetched at that point, and 11 of 13 unrelated product categories showed the same pattern. That analysis is a single practitioner study; treat it as provisional.

What the model already holds therefore shapes what it goes looking for. Brands named in ChatGPT's own query reached the final answer 68.9% of the time, against 2.1% for brands fetched but never named Mohanadasan · Snippet Digital; Keyword Insights · 2026. The author states those percentages are directional rather than measurements. He narrowed the headline in August 2026, after further testing found unnamed brands can still reach the answer. The direction still holds, and a two-surface comparison has more than two variables in it.

Neither surface is ground truth, and the divergence is engine-wide

Interface results are not a stable reference point either. ChatGPT's switch to GPT-5.3 Instant on 4 March 2026 cut cited URLs per response by about 20% de Segonzac · Resoneo · 2026. That study tracked 27,000 responses to 400 prompts over 14 weeks, and unique domains per response fell from roughly 19.6 to 15.5. Those figures belong to that switch and do not transfer to later models. The surface a tracker calls real changes underneath the tracker.

Cross-engine divergence sets the outer bound on what any single-surface number describes. Of 11,647 domains cited across five engines, 69.6% were cited by only one engine and 2.7% by all five Khallad · SurfacedBy · 2026. A critical survey of 45 GEO studies reaches the general form of the same conclusion Martinez · arXiv (cs.IR) · 2026. It describes generative engine optimisation as a stochastic, partially observable pipeline rather than one ranking task.

What a defensible AI visibility number carries

The practical output of this argument is a disclosure list, not a verdict on either surface. A visibility figure becomes comparable when it states five things. Access path, model variant, grounding configuration, run count, and spread. Without those, two numbers that look like the same metric may describe different systems.

Single-snapshot measurement understates true brand presence, and stochasticity is a measurement parameter rather than noise to discard Schulte et al. · University of St. Gallen · 2026. The addition this episode makes is narrow. The access path belongs in the same disclosure block as the run count. The study that started the argument is large enough to show the gap exists, and incomplete enough to show which fields the category leaves out.


Related reading:

FAQ
Frequently Asked Questions

Sources

Sources are tiered per our methodology & sources page.

Tier A — Strongest evidenceRead source

A preview of GPT-5.6 Sol, Terra, and Luna

OpenAI · 2026

Key finding

GPT-5.6 is not one model. OpenAI documents three, with separate API identifiers and three price tiers: gpt-5.6-sol at five and thirty dollars per million input and output tokens, gpt-5.6-terra at 2.50 and fifteen, gpt-5.6-luna at one and six. Sol is the flagship, Terra a lower-cost option, Luna the fastest and most cost-efficient. During the preview the family ran through the API and Codex only, and not in ChatGPT.

Methodology note

First-party OpenAI Help Center documentation, fetched directly on 31 August 2026; the page carries a relative update stamp rather than a publication date. Normative rather than empirical: it names the model identifiers, the price tiers and the preview access rules. Read alongside R280, OpenAI's general-availability announcement, which supersedes the preview-era ChatGPT statement.

OpenAI Help Center·Accessed
Key finding

In an internal OpenAI evaluation of financial, medical and legal prompts requiring factual detail, responses containing at least one factual error were about 68% less common with GPT-5.6 Sol and about 62% less common with GPT-5.6 Luna than with GPT-5.5 Instant, which OpenAI attributes to the model making better use of the sources it finds. GPT-5.6 Luna became the default model for Free and Go users; the Chat build of GPT-5.6 Sol is separate from the build powering Work and Codex. OpenAI states around 1 billion people use ChatGPT weekly.

Methodology note

First-party OpenAI product announcement, published 6 August 2026 and fetched directly on 2026-08-24. Not an independent study: the accuracy comparison is an internal evaluation with no disclosed sample size, prompt count, grader design or confidence intervals, covering three verticals only. Cited as evidence of OpenAI's own stated behaviour and intent, not as measurement.

OpenAI·Accessed
Key finding

A critical survey of 45 GEO studies (Nov 2023 to Jul 2026) argues GEO is not one ranking task but a stochastic, partially observable pipeline. The foundational Princeton gains are valid only for content already present in a fixed context, establishing neither organic discoverability nor durable traffic; topical relevance and context position are the most reproducible levers, generic heuristics transfer poorly, and citation-oriented rewrites can impair retrieval.

Methodology note

Single-author academic critical survey (Olivier Martinez), arXiv cs.IR, 18 pages, 8 tables, covering 45 GEO studies plus RAG and evaluation work, published 2026-07-15, not yet peer-reviewed. Ancillary literature matrix and search protocol included. Verified by direct fetch of the arXiv abstract page and metadata on the 2026-07-30 run.

arXiv·Accessed
Tier A — Strongest evidenceRead source

Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (arXiv 2607.13304)

Estonian Entrepreneurship University of Applied Sciences (EUAS); Rankfor.AI · Dmitrij Żatuchin · 2026

Key finding

Across 12,933 responses covering 20 brands, 8 languages and 3 engines at temperature 0.3, within-prompt resampling explained 34.8% of outcome variance and query language 31.6%, while brand identity explained 0.7%. Repeats past the fifth added almost nothing; adding languages cut relative-error variance roughly fifteen times more than five extra repeats.

Methodology note

Preprint by Dmitrij Żatuchin (Estonian Entrepreneurship University of Applied Sciences and Rankfor.AI), submitted 2026-07-14. A crossed random-effects model partitioned variance in brand sentiment across resampling, paraphrase, model identity and query language, using GPT-5.2, Gemini 3 Flash and Perplexity. Fetched and read directly.

arXiv·Accessed
Key finding

Citation visibility behaves as a sample estimate, not a fixed value. Reliable citation-share estimates required roughly 40 to 50 queries on Gemini, about 100 on Perplexity, and 150 or more on SearchGPT. Typical 95% confidence intervals spanned 3 to 6 percentage points, and apparent week-over-week movements often fell inside sampling noise.

Methodology note

Preprint by Ronald Sielinski (IQRush), submitted 2026-06-09. Bootstrap resampling built 95% confidence intervals around citation share and prevalence across Perplexity Search, OpenAI SearchGPT and Google Gemini, using 200 queries per topic across three consumer topics, sampled daily over nine days plus ten-minute-interval runs. Fetched and read directly.

arXiv·Accessed
Tier A — Strongest evidenceRead source

Don't Measure Once: Measuring Visibility in AI Search (GEO)

University of St. Gallen · Schulte et al. · 2026

Key finding

Argues that single-snapshot AI visibility measurement understates true brand presence in generative search. Proposes a longitudinal measurement framework that captures variation across runs, prompts, and platforms, demonstrating that any one-time snapshot of citation rate or mention rate can swing materially across repeated queries. Stochasticity itself is a measurement parameter, not noise to discard.

Methodology note

arXiv preprint 2604.07585 (April 2026). Position paper proposing a multi-run, multi-prompt evaluation protocol for GEO. Direct fetch on arxiv.org returned the canonical abstract page; PDF body was inaccessible but methodology summary was confirmed through the abstract and the linked DOI.

arXiv·Accessed
Key finding

Quantifies stochasticity in agentic LLM evaluations using intraclass correlation coefficients (ICC). Shows that single-run evaluations of agentic systems are unreliable because run-to-run variance is large relative to the gap between system variants. Recommends a minimum of 5 to 10 repeated runs per evaluation and reports the ICCs for several common agentic benchmarks.

Methodology note

arXiv preprint 2512.06710 (December 2025). Direct fetch on arxiv.org returned the abstract page. The paper applies the intraclass correlation coefficient framework from psychometrics to LLM agent evaluation and reports ICC values across multiple published benchmarks.

arXiv·Accessed
Tier A — Strongest evidenceRead source

Web search (OpenAI API documentation)

OpenAI · 2025

Key finding

OpenAI's web-search API documentation states that the web_search_call output item will usually (but not always) include the search queries that were searched, and that the sources field can reveal all URLs consulted during the search run. This is first-party proof that some query rewrites can be observed for requests under the caller's control — but the 'usually but not always' caveat means observed fan-out is partial rather than exhaustive.

Methodology note

Official OpenAI developer documentation for the web search tool exposed via the Responses API. Describes the schema of the web_search_call output item, including which fields are populated and the explicit caveat that searched queries are returned 'usually (but not always).' Content verified by fetch on 2026-05-27. No aggregate usage data is disclosed.

OpenAI Developer Platform·Accessed
Key finding

SurfacedBy analyzed 127,198 source citations from ChatGPT, Claude, Gemini, Perplexity, and Google AI Mode across roughly 16,400 commercial-intent answers between March and June 2026. Of 11,647 cited domains, 69.6% were cited by only one engine and just 2.7% by all five. Vendor, product, and long-tail pages drew 90.6% of citations; Reddit 1.8% and Wikipedia 0.6%. Gemini averaged 11.0 sources per answer, ChatGPT 3.7.

Methodology note

First-party experiment by SurfacedBy, an AI-visibility tracking vendor with commercial interest, published 27 June 2026 and updated 29 June. About 16,400 answers to real buyer and category questions across five engines; citations counted at the domain level. Authors disclose limits: commercial-query skew, citations are not clicks, engine behavior shifts. Verified by direct fetch.

SurfacedBy Blog·Accessed
Tier B — Citable with caveatsRead source

The Bigfoot Effect: How ChatGPT's Search Visibility Collapsed Overnight (Resoneo)

Resoneo · Olivier de Segonzac · 2026

Key finding

Across 27,000 responses to 400 prompts over 14 weeks, ChatGPT's switch to GPT-5.3 Instant on 4 March 2026 cut the number of cited URLs shown per response by about 20%, with unique domains per response falling from roughly 19.6 to 15.5. The ratio of URLs to domains stayed constant at 1.26 throughout, so results are effectively deduplicated by domain and fewer domains occupy the same visibility surface. With over 90% of ChatGPT weekly users on the free tier, the default experience triggers fewer web searches, uses fewer queries and produces fewer citations than paid tiers.

Methodology note

Longitudinal tracking study by Resoneo, published March 2026 and read from PDF on 2026-08-24. Disclosed sample: 27,000 responses across 400 prompts over 14 weeks, spanning the 4 March 2026 default-model switch. Prompt selection criteria, vertical mix and confidence intervals are not disclosed. Resoneo is an agency with a commercial interest in AI-search consulting and distributes a ChatGPT scraping plugin used by other researchers in this strand. Scope is explicitly GPT-5.3 Instant and earlier.

Resoneo·Accessed
Tier C — Tactical signals onlyRead source

API vs UI - How to track LLM visibility?

Peec AI · Malte Landwehr · 2026

Key finding

Across 555 prompts run 47 times each on both the ChatGPT UI and the API, 52,170 chats, the average brand's visibility differed by 41% between surfaces and the most visible UI brand missed the API top three. Competitor domains rose from 38% to 52% of citations, listicles fell from 25% to 4%, product pages and how-to guides grew from 34% to 63%, and API answers ran 31% longer.

Methodology note

Self-published LinkedIn article by Peec AI's CPO, 29 August 2026, read in full from a PDF capture. The design is large and repeated, and 47 runs per prompt per surface clears the sampling thresholds in R242 and R243, but the article discloses no temperature, no web-search setting, no UI account state and no GPT-5.6 variant, and reports no confidence intervals.

LinkedIn (article)·Accessed
Tier C — Tactical signals onlyRead source

What We Learned From Analyzing 28K+ ChatGPT And Gemini Fan-Out Queries (Nectiv)

Nectiv · Chris Long · 2026

Key finding

Running about 4,000 prompts from its 2025 baseline set through the GPT-5.6 Sol API, Nectiv found average fan-out queries per prompt rising from 2.17 to 7.61 and the longest single chain from 4 searches to 29. The site: operator appeared in 64% of all fan-out queries, with 'official' the second most common unigram and 'gov' also in the top five. Freshness unigrams dropped out of the top five and ChatGPT began running multi-year searches covering both the current and previous year. Software was the most searched vertical at 10.7 fan-outs per prompt.

Methodology note

Single-agency data study by Chris Long, co-founder of Nectiv, published 13 August 2026 and fetched directly on 2026-08-24. Roughly 4,000 prompts, a subset of the agency's 2025 study set with stated equal representation across verticals, re-run through the GPT-5.6 Sol API with fan-out queries extracted programmatically. No absolute sample sizes per vertical, no confidence intervals, no control for personalisation. API extraction is clean and repeatable but may diverge from what consumer-interface users receive.

Nectiv Blog·Accessed
Tier C — Tactical signals onlyRead source

ChatGPT Already Knows Who's In The Running Before It Searches

Snippet Digital; Keyword Insights · Suganthan Mohanadasan · 2026

Key finding

Reading ChatGPT network traffic, the author found that in 21 of 27 conversations the model's first search query already contained brand names the user never typed, before any page was fetched, and 11 of 13 unrelated product categories showed the same pattern. Brands named in ChatGPT's own query reached the final answer 68.9% of the time versus 2.1% for brands whose pages were fetched but never named, a gap of roughly 33 times. Across 57 conversations and 3,554 retrieved pages only 110 were cited, a rate of 3.1%. Citation rate fell sharply with position inside a domain group, from 5.2% at first position to 0.3% at sixth or later.

Methodology note

Single-researcher study by Suganthan Mohanadasan, published 10 August 2026, updated 17 August 2026, read from PDF on 2026-08-24. Read from one logged-in ChatGPT Plus account in Dubai between 24 and 25 July 2026: 57 conversations for the citation figures, 27 for the first-query test, 12 fresh category queries and 3 repeats, weighted toward software and AI tools. The author states every percentage is directional rather than a measurement, discloses personalisation effects visible in his own data, and flags brand-name tokenisation errors in his counts. The mechanism is independently reproducible; the magnitudes are not generalisable.

suganthan.com·Accessed

About the author Max Ackermann

Max Ackermann is founder and Managing Director of info.link, the product data platform that makes brands visible in AI search and connects every physical product to the web through GS1 Digital Link. He writes about AI search and generative engine optimization (GEO), AI-powered commerce, and how brands can structure product data for ChatGPT, Gemini, Perplexity, and retailer AI assistants like Amazon Rufus. For the past two years he has built the pipelines that put structured product data into AI answers, and run the experiments that test what actually moves AI citations.

Max has 20+ years of experience building digital products and businesses. He previously led McKinsey's Corporate Venture and Design teams across Europe, and as Managing Director of a leading US digital agency he built platforms with Nike, Google, Meta, and Airbnb. He founded the UX Design program at Central Saint Martins College, University of the Arts London, and is a Fellow of the UK's Higher Education Academy. Based in Hamburg, he works closely with GS1 on Digital Link adoption; info.link is headquartered in Hamburg and Berlin and counts GS1 Germany among its investors.

Follow Max on LinkedIn.

Interested?

From compliant digital labels to AI-verified product answers, we help leading brands ensure their products are visible and accurately represented everywhere consumers look. Book your free consultation and demo.

digital label preview
digital label preview
digital label preview
ChatGPT API vs UI: do they return the same citations? | info.link