AI visibility measurement
AI visibility measurement estimates how often a brand or page appears in the answers of an answer engine, usually by sampling a fixed set of prompts and counting citations or brand mentions. It is harder than rank tracking because the surface is unstable: resubmitting an identical prompt returns only a 0.50 to 0.61 Jaccard overlap in the brands recommended, and paraphrasing the same buying question drops agreement to between 0.135 and 0.288[1]. A number from one run is not comparable with a number from another.
Why one reading is not enough
The instability is partly mechanical. Production inference endpoints are non-deterministic even at temperature 0, because server load changes the batch size and batch-size-dependent numerical kernels send identical requests down different floating-point paths[2]. It is partly the search surface. An empirical study of 11,500 queries across Google Search, AI Overviews and Gemini Flash 2.5 found an AI Overview on 51.5% of queries, drawing on sources with under 0.2 Jaccard similarity to the organic results, and measurably less consistent across two runs of the same query[3]. The vendor Profound tracked 10,000 randomly sampled prompts over 14 days in spring 2026, against ChatGPT, Perplexity and Copilot, and reported stable results, with day-of-week variation small next to the differences between engines[4]. A 2026 measurement paper draws the opposite conclusion from the same behaviour: answers vary across runs, prompts and time, so a one-off observation is unreliable and visibility belongs in a report as a distribution rather than a single-point outcome[5].
The gap the numbers show
A study of more than 100,000 prompt responses covering over 100 brands between March and May 2026 found that on a brand's first visibility run, global household names appeared in 73% of relevant AI answers, established mid-market and regional brands in 44%, and niche and small brands in 11%[6]. The same work found a brand's measured sentiment flips between runs roughly 6.7 times more often than whether the brand is mentioned at all[7], so presence tracking is far steadier than tone tracking and should be reported separately from it.
An audit of roughly 37,000 commercial recommendation runs puts the failure at the bottom of the prominence ladder rather than spread evenly across it: among specialist and regional brands, 48% to 52% never surfaced in any run of the set[8]. A visibility programme for a brand in that band is measuring absence, and needs enough runs to distinguish it from sampling noise.
Fabricated citations rise with familiarity
Measurement also has to allow for citations that do not exist. A per-entity bias study of 100 Hungarian B2B entities, using 1,400 probe runs over 2,062 sources, found well-known tier 1 brands produced a fabricated-citation rate of 52.69% against 37.87% for obscure tier 3 entities (p = 1.67e-11)[9]. Model familiarity with a brand yields confident wrong citations rather than fewer of them, so a visibility score needs per-entity calibration before it can be set beside another brand's. Citation accuracy covers the failure mode itself.
Language and model both change the answer
A 12-language study of 66 brands across three models and 35,640 responses found that moving from an English query to a brand's home-market language raised recommendation share by 0.80 for local champions but only 0.15 for global multinationals. Response stability varied more with which model was asked than with which language was used, at an eta-squared of 0.32 against 0.01[10]. A programme that samples one model in one language has measured that combination and nothing wider, which is the practical argument behind multilingual GEO.
What the samples say drives citation
A controlled experiment of 252,000 trials across six models, testing 18 content factors with brand names anonymised so content effects could be separated from position bias, found topical relevance and position in the retrieved list were the strongest predictors of being cited first[11]. The founding GEO paper reported content changes could raise visibility by up to 40%, with the efficacy of the strategies varying across domains[12]. Both results argue for measuring per engine and per market rather than publishing one visibility figure; SEO and GEO covers how much of that overlaps with rank tracking.
Data
Show the numbers (3)
| Point | % |
|---|---|
| Global brands | 73 |
| Established mid-market brands | 44 |
| Niche and small brands | 11 |
References (12)
- Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline Archive
- Defeating Nondeterminism in LLM Inference Archive
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews Archive
- What AI Engines Actually Search For
- Don't Measure Once: Measuring Visibility in AI Search (GEO)
- Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines Archive
- Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines Archive
- Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit
- Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration Archive
- The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages Archive
- What Gets Cited: Competitive GEO in AI Answer Engines Archive
- GEO: Generative Engine Optimization Archive
Last updated 2026-09-04. Written and maintained by Baseline Labs.