Skip to content

AI visibility measurement

Concept in AI search and SEO
Volatile, last checked 2026-09-04 This page carries measured figures that move quickly.Reviewed 2026-09-04

AI visibility measurement estimates how often a brand or page appears in the answers of an answer engine, usually by sampling a fixed set of prompts and counting citations or brand mentions. It is harder than rank tracking because the surface is unstable: resubmitting an identical prompt returns only a 0.50 to 0.61 Jaccard overlap in the brands recommended, and paraphrasing the same buying question drops agreement to between 0.135 and 0.288[1]. A number from one run is not comparable with a number from another.

Why one reading is not enough

The instability is partly mechanical. Production inference endpoints are non-deterministic even at temperature 0, because server load changes the batch size and batch-size-dependent numerical kernels send identical requests down different floating-point paths[2]. It is partly the search surface. An empirical study of 11,500 queries across Google Search, AI Overviews and Gemini Flash 2.5 found an AI Overview on 51.5% of queries, drawing on sources with under 0.2 Jaccard similarity to the organic results, and measurably less consistent across two runs of the same query[3]. The vendor Profound tracked 10,000 randomly sampled prompts over 14 days in spring 2026, against ChatGPT, Perplexity and Copilot, and reported stable results, with day-of-week variation small next to the differences between engines[4]. A 2026 measurement paper draws the opposite conclusion from the same behaviour: answers vary across runs, prompts and time, so a one-off observation is unreliable and visibility belongs in a report as a distribution rather than a single-point outcome[5].

The gap the numbers show

A study of more than 100,000 prompt responses covering over 100 brands between March and May 2026 found that on a brand's first visibility run, global household names appeared in 73% of relevant AI answers, established mid-market and regional brands in 44%, and niche and small brands in 11%[6]. The same work found a brand's measured sentiment flips between runs roughly 6.7 times more often than whether the brand is mentioned at all[7], so presence tracking is far steadier than tone tracking and should be reported separately from it.

An audit of roughly 37,000 commercial recommendation runs puts the failure at the bottom of the prominence ladder rather than spread evenly across it: among specialist and regional brands, 48% to 52% never surfaced in any run of the set[8]. A visibility programme for a brand in that band is measuring absence, and needs enough runs to distinguish it from sampling noise.

Fabricated citations rise with familiarity

Measurement also has to allow for citations that do not exist. A per-entity bias study of 100 Hungarian B2B entities, using 1,400 probe runs over 2,062 sources, found well-known tier 1 brands produced a fabricated-citation rate of 52.69% against 37.87% for obscure tier 3 entities (p = 1.67e-11)[9]. Model familiarity with a brand yields confident wrong citations rather than fewer of them, so a visibility score needs per-entity calibration before it can be set beside another brand's. Citation accuracy covers the failure mode itself.

Language and model both change the answer

A 12-language study of 66 brands across three models and 35,640 responses found that moving from an English query to a brand's home-market language raised recommendation share by 0.80 for local champions but only 0.15 for global multinationals. Response stability varied more with which model was asked than with which language was used, at an eta-squared of 0.32 against 0.01[10]. A programme that samples one model in one language has measured that combination and nothing wider, which is the practical argument behind multilingual GEO.

What the samples say drives citation

A controlled experiment of 252,000 trials across six models, testing 18 content factors with brand names anonymised so content effects could be separated from position bias, found topical relevance and position in the retrieved list were the strongest predictors of being cited first[11]. The founding GEO paper reported content changes could raise visibility by up to 40%, with the efficacy of the strategies varying across domains[12]. Both results argue for measuring per engine and per market rather than publishing one visibility figure; SEO and GEO covers how much of that overlaps with rank tracking.

Data

First-run appearance rate in relevant AI answers, by brand tierGlobal brands73%Established mid-market brands44%Niche and small brands11%
First-run appearance rate in relevant AI answers, by brand tier. Figures in %.Source: arxiv.org
Show the numbers (3)
Point%
Global brands73
Established mid-market brands44
Niche and small brands11
References (12)
  1. Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline Archive
    Academic Published 2026-05-22 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  2. Defeating Nondeterminism in LLM Inference Archive
    Documentation Published 2025-09-10 Retrieved 2026-09-04
  3. How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews Archive
    Academic Published 2026-04-30 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  4. What AI Engines Actually Search For
    Vendor documentation Published 2026-04-30 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  5. Don't Measure Once: Measuring Visibility in AI Search (GEO)
    Academic Published 2026-04-08 Retrieved 2026-09-04 Single-source
  6. Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines Archive
    Academic Published 2026-06-18 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  7. Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines Archive
    Academic Published 2026-06-18 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  8. Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit
    Academic Published 2026-05-22 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  9. Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration Archive
    Academic Published 2026-06-19 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  10. The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages Archive
    Academic Published 2026-06-22 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  11. What Gets Cited: Competitive GEO in AI Answer Engines Archive
    Academic Published 2026-05-25 Retrieved 2026-09-04
  12. GEO: Generative Engine Optimization Archive
    Academic Published 2024-06-28 Retrieved 2026-09-04

Last updated 2026-09-04. Written and maintained by Baseline Labs.

George
Online
0%