You can't measure AI search with one number: four instruments and what each one can't see
Reliable AI visibility estimates need roughly 40 to 150 queries per platform depending on the engine, and typical 95% confidence intervals span 3 to 6 percentage points, so week-over-week movement is usually sampling noise. Four instruments measure AI search, and each one is blind to what the others catch.
One company's web analytics attributed 0% of new leads to AI assistants, while roughly 25% of its lead-form respondents named one Landwehr · Peec AI · 2026. Both figures come from Peec AI's Malte Landwehr and rest on no disclosed method. Both can still be correct, because clicks and influence are different quantities. Four instruments measure AI search: prompt tracking, server logfile analysis, web analytics, and self-reported attribution. Each has a blind spot, usually documented by the vendor selling the instrument, and the sections below take each in turn.
The four-instrument framework is emerging consensus, not one vendor's idea
The framework circulating as a new insight was published independently, twice, eleven weeks apart. Search Engine Land set out five layers on 18 May 2026: GA4 attribution, crawl-log diagnostics, share of voice, self-reported attribution, and incrementality testing DeMott · Search Engine Land · 2026. Landwehr's four-pillar version followed in August 2026 and drops the fifth Landwehr · Peec AI · 2026.
Independent convergence is the reason to take the framework seriously. It also means the framework is not the contribution. Neither version supplies the statistical basis: how many measurements a number needs, and which measurements are wasted.
Prompt tracking is bounded by the prompt set, and the vendors document this themselves
Prompt tracking is synthetic data. It measures how often a brand appears in answers to prompts the customer chose. It does not measure answers to prompts real buyers typed.
The vendors say so plainly. Otterly's help documentation states there is "no way to learn which prompts are most asked at ChatGPT or Perplexity" Otterly AI · 2024. It recommends building prompts from brand terms, domains, industries, URLs, and SEO keywords. Its companion page confirms prompts are customer-defined and proxy-derived Otterly AI · 2024. Peec's documentation describes the same shape, running customer-supplied prompts daily across platforms Peec AI · 2024.
This makes prompt tracking the only instrument that produces a competitor benchmark. It is also the only one whose output moves when the prompt list changes, with nothing changing in the market.
Logfiles prove a URL was requested, not that it was used
Server logs are real, server-side, unsampled records of which AI crawlers fetched which pages. They are silent on everything that happens next. A log entry cannot show whether the page shaped the answer, whether the brand was named, or how it was described.
The volume gap makes the limit concrete. Anthropic's ClaudeBot crawled around 70,000 pages for every visitor it referred Cloudflare · 2025. OpenAI's GPTBot crawled around 1,700 per visitor and Perplexity around 5 Cloudflare · 2025. Crawl activity at that ratio confirms attention and settles nothing about outcome. Logfiles also carry no competitor benchmark, and their reliability declines as crawler caching improves.
The analytics number is miscounted, not undercounted
Web analytics sees only visits that produced a click, and most AI-referred visits arrive unlabelled. Across a 446,405-visit dataset, 20,428 visits came from AI sources Di Cesare · Loamly · 2025. Of those, 14,413, or 70.6%, had no referrer header and landed in GA4 as Direct Di Cesare · Loamly · 2025. The named mechanisms are users pasting copied URLs, and ChatGPT's mobile app and Atlas browser stripping referrers. That is misattribution with a fixable cause, not absence.
The hidden visits are unusually valuable. Those unlabelled AI visits converted at 10.21% against 2.46% for non-AI traffic Di Cesare · Loamly · 2025. A separate study put AI-assistant traffic at roughly 3 times the conversion rate of other channels Microsoft Clarity · 2026.
Volume is where analytics is weakest. Across 81,947 sites, average AI traffic rose about 9.7 times in a year Ahrefs · 2025. It still reached only 0.25% of total traffic. Clicks are scarce by design. Users click a traditional link on about 8% of searches showing an AI summary, against 15% without one Pew Research Center · 2025.
Self-reported attribution sees the decision, and it runs on memory
Asking buyers where they heard about a brand is the only method that captures dark chat. That is a conversation with an assistant, followed by a direct visit or a branded search. The signal sits closer to revenue than anything the other three instruments return.
The behaviour it measures is well evidenced. Among 1,076 B2B software buyers, 51% now start research with an AI chatbot more often than Google G2 · 2026. On the assistant's guidance, 69% chose a different vendor than planned.
The instrument's weakness is the respondent. Self-report captures what a buyer remembers and will write down, which favours the most recent and most nameable touch. It carries the recall limits every stated-preference measure carries. The G2 figures also come from a single vendor with a commercial interest in the answer.
A visibility percentage without a sample size is not a measurement
The sample sizes are published, and they are larger than most reporting assumes. Reliable citation-share estimates needed roughly 40 to 50 queries on Gemini and about 100 on Perplexity Sielinski · IQRush · 2026. SearchGPT needed 150 or more, with typical 95% confidence intervals spanning 3 to 6 percentage points Sielinski · IQRush · 2026. The same work found apparent week-over-week changes often fall inside sampling noise.
A second study points the same way. Single-snapshot measurement understates brand presence, and run-to-run variation is a measurement parameter rather than noise to discard Schulte et al. · University of St. Gallen · 2026.
Drift compounds the problem across time. Across roughly 80,000 prompts per platform in June and July 2025, Google AI Overviews drifted 59.3% and Perplexity 40.5% Blyskal · Profound · 2025. Over a January-to-July window, drift rises to 70% to 90%. A number reported without its interval describes the instrument as much as the market.
Running the same prompt every day is the least productive use of query budget
The standard advice is to run a fixed prompt set daily and aggregate the results. A variance decomposition tested that. Across 12,933 responses covering 20 brands, 8 languages, and 3 engines, within-prompt resampling explained 34.8% of outcome variance Żatuchin · Estonian Entrepreneurship University of Applied Sciences (EUAS); Rankfor.AI · 2026. Query language explained 31.6%, while brand identity explained 0.7%. A repeat past the fifth changed the estimate by 0.0003, and adding languages cut relative-error variance about fifteen times as much as five more repeats.
The caveat matters as much as the finding. That study measured multilingual sentiment polarity across 20 Central and Eastern European brands. Its language component partly reflects its own design, so it should not be carried over to a single-market English tracker. The allocation conclusion survives: repeats are the least efficient of the four axes available.
Repeated runs are still necessary, and a separate estimate lands in the same place. Work using intraclass correlation recommends a minimum of 5 to 10 repeated runs Mustahsan · arXiv · 2025. Five runs spread across more prompts, languages, and engines buy more precision than thirty runs of one prompt.
Only incrementality supports a causal claim, and the popular framework omits it
The four instruments describe correlation. Prompt tracking shows a brand appearing more often. Logs show more crawling, analytics shows more sessions, and self-report shows more mentions. None of the four isolates the effect of the work that was done.
Portfolio-level difference-in-differences analysis does. It compares properties with different levels of investment, and it is the fifth layer the four-pillar version drops DeMott · Search Engine Land · 2026. It is also the layer fewest teams can run, because it needs comparable properties, a clean intervention date, and a control group. Agencies with portfolios can attempt it. Single brands usually cannot.
One claim in the circulating framework has no published test behind it
The most consequential claim in the popular version is the least supported. Landwehr states that ChatGPT's and Perplexity's APIs return different brands and sources than their user interfaces Landwehr · Peec AI · 2026. He says Peec AI tested this across thousands of prompts, and that tools relying only on the chat API therefore show false data. If it holds, that claim disqualifies a category of tracking tools.
No published methodology supports it, from Peec AI or from any independent party. The vendor making the claim also sells a tool built on the interface-simulation approach the claim favours. The reasoning is plausible, because an API endpoint and a consumer product can run different retrieval configurations. It remains an assertion, and a buyer told that a rival tool shows false data has no public evidence available to check it.
FAQFrequently Asked Questions
Sources
Sources are tiered per our methodology & sources page.
Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (arXiv 2607.13304)
Estonian Entrepreneurship University of Applied Sciences (EUAS); Rankfor.AI · Dmitrij Żatuchin · 2026
Across 12,933 responses covering 20 brands, 8 languages and 3 engines at temperature 0.3, within-prompt resampling explained 34.8% of outcome variance and query language 31.6%, while brand identity explained 0.7%. Repeats past the fifth added almost nothing; adding languages cut relative-error variance roughly fifteen times more than five extra repeats.
Methodology note
Preprint by Dmitrij Żatuchin (Estonian Entrepreneurship University of Applied Sciences and Rankfor.AI), submitted 2026-07-14. A crossed random-effects model partitioned variance in brand sentiment across resampling, paraphrase, model identity and query language, using GPT-5.2, Gemini 3 Flash and Perplexity. Fetched and read directly.
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement (arXiv 2603.08924)
IQRush · Ronald Sielinski · 2026
Citation visibility behaves as a sample estimate, not a fixed value. Reliable citation-share estimates required roughly 40 to 50 queries on Gemini, about 100 on Perplexity, and 150 or more on SearchGPT. Typical 95% confidence intervals spanned 3 to 6 percentage points, and apparent week-over-week movements often fell inside sampling noise.
Methodology note
Preprint by Ronald Sielinski (IQRush), submitted 2026-06-09. Bootstrap resampling built 95% confidence intervals around citation share and prevalence across Perplexity Search, OpenAI SearchGPT and Google Gemini, using 200 queries per topic across three consumer topics, sampled daily over nine days plus ten-minute-interval runs. Fetched and read directly.
Don't Measure Once: Measuring Visibility in AI Search (GEO)
University of St. Gallen · Schulte et al. · 2026
Argues that single-snapshot AI visibility measurement understates true brand presence in generative search. Proposes a longitudinal measurement framework that captures variation across runs, prompts, and platforms, demonstrating that any one-time snapshot of citation rate or mention rate can swing materially across repeated queries. Stochasticity itself is a measurement parameter, not noise to discard.
Methodology note
arXiv preprint 2604.07585 (April 2026). Position paper proposing a multi-run, multi-prompt evaluation protocol for GEO. Direct fetch on arxiv.org returned the canonical abstract page; PDF body was inaccessible but methodology summary was confirmed through the abstract and the linked DOI.
Stochasticity in Agentic Evaluations: Quantifying Inconsistency with Intraclass Correlation
arXiv · Mustahsan · 2025
Quantifies stochasticity in agentic LLM evaluations using intraclass correlation coefficients (ICC). Shows that single-run evaluations of agentic systems are unreliable because run-to-run variance is large relative to the gap between system variants. Recommends a minimum of 5 to 10 repeated runs per evaluation and reports the ICCs for several common agentic benchmarks.
Methodology note
arXiv preprint 2512.06710 (December 2025). Direct fetch on arxiv.org returned the abstract page. The paper applies the intraclass correlation coefficient framework from psychometrics to LLM agent evaluation and reports ICC values across multiple published benchmarks.
The Crawl-to-Click Gap: Cloudflare Data on AI Bots, Training, and Referrals
Cloudflare · 2025
AI crawlers read content far more than they send referrals back. Anthropic's ClaudeBot crawled around 70,000 pages for every visitor it referred; OpenAI's GPTBot crawled around 1,700 for every visitor; Perplexity around 5 for every visitor. Mistral was the only major AI engine where referrals outweighed crawl volume.
Methodology note
Aggregate analysis of crawl requests and referral traffic across the Cloudflare network. For each major AI crawler, the team divided pages crawled by visits sent to the same destinations during the same window, producing a crawl-to-refer ratio. Published August 2025.
Google Users Are Less Likely to Click on Links When an AI Summary Appears in Search Results
Pew Research Center · 2025
When a Google search result page includes an AI summary, users click on a traditional link in roughly 8% of visits. On result pages without an AI summary, they click in roughly 15% of visits. Users rarely click on the citations inside the AI summary itself, doing so on about 1% of visits.
Methodology note
Pew Research panel study covering 900 US adults and 68,879 Google searches conducted between March and May 2025. Sessions were tracked through opt-in browser participation; click behaviour was observed directly rather than self-reported. Published July 2025.
The 5-Layer Framework for Measuring GEO Performance (Search Engine Land)
Search Engine Land · Paul DeMott · 2026
An independently published five-layer GEO measurement framework: GA4 direct attribution, crawl-log diagnostics, share of voice plus structured AI interrogation, self-reported attribution form fields, and portfolio-level incrementality testing. It cites Cloudflare's June 2025 crawl-to-referral ratios of 1,700:1 for OpenAI and 73,000:1 for Anthropic, and the Loamly 70.6% Direct-misattribution figure.
Methodology note
Trade-press contributor article by Paul DeMott, published in Search Engine Land on 2026-05-18. A practitioner framework rather than a study: it presents no original dataset and sources its figures from Cloudflare and Loamly. Value here is independent convergence on the same measurement pillars. Fetched and read directly.
The Answer Economy: How AI Search Is Rewiring B2B Software Buying
G2 · 2026
G2's survey of 1,076 B2B software buyers, fielded March 2026, found 51% now start research with an AI chatbot more often than Google, up from 29% a year earlier, and 71% use AI chatbots for vendor research. 69% chose a different vendor than planned on AI guidance, and a third bought from a vendor new to them. Single-vendor survey with commercial interest.
Methodology note
First-party survey by G2 of 1,076 B2B software buyers and decision-makers, fielded March 2026 and published 15 April 2026 as 'The Answer Economy'. G2 sells answer-engine-optimization products, so treat as directional, not independent. Verified against G2's own release and PR Newswire. An unverified '85% think more highly' figure circulating in aggregators does not trace to G2's release.
AI Traffic Converts at 3× the Rate of Other Channels (Study)
Microsoft Clarity · 2026
Visitors arriving from AI assistants convert at roughly 3 times the rate of visitors from other channels, and at up to 11 times the rate in certain publisher segments. AI traffic still represents a small share of total visits, but its per-visitor commercial value is materially higher than traditional search or social.
Methodology note
Analysis of Microsoft Clarity user-session data across a multi-publisher dataset. The study compared conversion rates of sessions originating from AI assistants against sessions from other referral channels. Published January 2026. Single-vendor study with disclosed methodology, downgraded to Tier B in v1.1.
The AI Traffic Attribution Crisis: Why Your Analytics Are Wrong (Loamly)
Loamly · Marco Di Cesare · 2025
Across 446,405 visits in Loamly's database, 20,428 came from AI sources, and 14,413 of those, or 70.6%, arrived without a referrer header and landed in GA4 as Direct. Named causes are users pasting copied URLs, and ChatGPT mobile and Atlas stripping referrers. Those dark AI visits converted at 10.21% versus 2.46% for non-AI traffic.
Methodology note
Vendor analysis by Marco Di Cesare of Loamly, published 2025-11-07 and last updated 2026-02-16, drawing on the company's own visit database. Visit count and AI-visit split are disclosed; the number of distinct websites and the collection time window are not. Fetched and read directly.
AI Search Volatility: Citation Drift Across ChatGPT, Google AI Overviews, Microsoft Copilot, and Perplexity
Profound · Josh Blyskal, Sartaj Rajpal · 2025
Across roughly 80,000 prompts per platform tested in June 2025 and again in July 2025, Profound measured citation drift — the share of domains appearing in the later window but not the earlier one. Google AI Overviews drifted 59.3%, ChatGPT 54.1%, Microsoft Copilot 53.4%, Perplexity 40.5%. Over a January-to-July comparison, drift rises to 70–90%, making single-snapshot AI visibility measurements unreliable.
Methodology note
Vendor research study by Profound, published 17 July 2025. Compared domain-level citations on identical open-ended prompts across two three-day windows: 11–13 June and 11–13 July 2025. Sample roughly 80,000 prompts per platform. Drift defined as the percentage of domains cited in the later window but absent in the earlier window. Source verified by direct fetch.
AI Traffic Has Increased 9.7× in the Past Year (81,947 Websites Study)
Ahrefs · 2025
Across 81,947 websites, average AI traffic grew about 9.7 times in a year. The average site's search traffic dropped about 21% over the same period. AI traffic now represents 0.25% of a site's total traffic on average. ChatGPT grew 85% since January 2025 and now sends more traffic than Reddit or LinkedIn. Google still sends about 210 times more traffic than the big three AI platforms combined.
Methodology note
Ahrefs analysed referral traffic patterns across 81,947 websites between mid-2024 and mid-2025, comparing AI referrals (ChatGPT, Perplexity, Gemini, Copilot) against traditional search, social platforms, and direct traffic. The dataset more than doubled the size of the earlier March 2025 study.
How to find relevant prompts for your brand? (Otterly Help)
Otterly AI · 2024
Otterly's own help documentation explicitly states there is 'no way to learn which prompts are most asked at ChatGPT or Perplexity' and 'no way to know what exactly people are searching for in the AI engines.' Otterly recommends constructing prompts from available external inputs such as brand terms, domains, industries, URLs, and SEO keywords. This is a vendor admission that aligns with the public-proxy thesis.
Methodology note
Otterly AI Help Center article (last updated April 2026) describing the vendor's own recommended methodology for building a brand's prompt list. Self-reported vendor documentation; the page explicitly states that AI search engines do not publish query data and lists three substitute methods Otterly supports (Prompt Research tool, Google Search Console import, AI-assisted brainstorming). Content verified by direct fetch.
Otterly supports prompt construction from external proxies including SEO keywords, brand names, industry terms, and URLs. The page reinforces that prompts are customer-defined and proxy-derived, not drawn from a privileged platform-wide feed of real chatbot user prompts.
Methodology note
Otterly AI Help Center article (December 2025) describing the three ways customers can add prompts inside the Otterly platform: individual entry, CSV import, or the AI Prompt Research tool. Self-reported vendor documentation. Useful as evidence of the kinds of inputs Otterly accepts; not a controlled study or independent benchmark. Content verified by direct fetch.
Peec's documentation says the platform runs customer prompts daily across AI platforms. This supports the interpretation that vendors like Peec observe outcomes from prompts they execute rather than drawing from a secret platform-wide prompt firehose.
Methodology note
Peec AI's official Quickstart Guide, published on its Mintlify-hosted documentation site. Describes the four-step onboarding workflow (set up prompts, identify competitors, read the dashboard, analyse sources) and confirms that Peec runs customer-defined prompts daily across ChatGPT, Perplexity, Gemini and Copilot. Content verified by direct fetch on 2026-05-27.
Can You Measure AI Search? Combining Four Data Sources (GEO 101, Peec AI)
Peec AI · Malte Landwehr · 2026
Landwehr reports one company where web analytics attributed 0% of new leads to LLMs while more than 20% of lead-form respondents named ChatGPT, and roughly 25% named any LLM. Peec AI customers Graphite and n8n showed 0.9% of leads from LLMs in Google Analytics versus 9% self-reported, a tenfold gap. Neither figure discloses a sample size.
Methodology note
Ten-minute GEO 101 explainer video by Malte Landwehr, CPO and CMO of prompt-tracking vendor Peec AI, with a companion LinkedIn summary post dated 2026-08-05. No sample sizes, time windows or survey instruments are disclosed. Content confirmed against the full verbatim transcript supplied by Max; YouTube blocked direct fetch.
About the author Max Ackermann
Max Ackermann is founder and Managing Director of info.link, the product data platform that makes brands visible in AI search and connects every physical product to the web through GS1 Digital Link. He writes about AI search and generative engine optimization (GEO), AI-powered commerce, and how brands can structure product data for ChatGPT, Gemini, Perplexity, and retailer AI assistants like Amazon Rufus. For the past two years he has built the pipelines that put structured product data into AI answers, and run the experiments that test what actually moves AI citations.
Max has 20+ years of experience building digital products and businesses. He previously led McKinsey's Corporate Venture and Design teams across Europe, and as Managing Director of a leading US digital agency he built platforms with Nike, Google, Meta, and Airbnb. He founded the UX Design program at Central Saint Martins College, University of the Arts London, and is a Fellow of the UK's Higher Education Academy. Based in Hamburg, he works closely with GS1 on Digital Link adoption; info.link is headquartered in Hamburg and Berlin and counts GS1 Germany among its investors.
Follow Max on LinkedIn.


