Skip to content

How we measure AI visibility

Every engine you select is asked every keyword in your template. Each answer records whether your domain is cited, at what position, and whether your brand is named at all.

Depth: wordings per question

How many wordings of each question the run asks. Draw one is your question as written; the rest are meaning-preserving rewordings in your template's language, asked on every engine you selected and pooled under the question.

Draws per intentGradeCostWhat it buys
-T1Snapshot-Not offered. The first rung already reaches T2.
1 DefaultT2DirectionalShows the direction. The rate can move.×1Plus or minus 45 points, reads about 15 high.
2T2DirectionalShows the direction. The rate can move.×2Plus or minus 25 points.
3T2DirectionalShows the direction. The rate can move.×3Plus or minus 25 points, keyword bias about 5.
5T3ConsensusA rate within 30 points either way.×5Plus or minus 19 points.
-T4Established-Not offered.
-T5Verified-Not offered.

Why it stops at T3. Five wordings measured plus or minus 19 points; the next grade needs 15.

Wording carried 26.5% of total variance against brand identity's 1.5%, so depth buys new wordings rather than repeats.

Tested on 20 questions across 6 studies, in English and Swedish, on 5 engines. Each run was scored against a second run's fresh wordings, so the error is to the question, not to the keyword.

WordingsPlus or minus, mention ratePlus or minus, citation rateReads high by
142 points51 points13 to 16 points
225 points24 points6 to 8 points
325 points24 points5 to 7 points
519 points16 points3 points

Your question as written reads high because a template is written in the words a brand already ranks for: on 12 of 20 questions the original wording found the brand more often than its rewordings, and on 2 it found it less. Running the same wording again agrees with itself far more closely than a new wording does, so a stable number is not the same as an accurate one. Five wordings stop short of Established, which needs 15 points.

Google is resampled too, because a reworded question is a new query with its own result set. Within one wording the Google result set is fetched once and shared by Google, ChatGPT, Claude and DeepSeek, so their disagreement is about the models, not fetch timing. A rewording we cannot produce falls back to repeating the original, and the report's method says how many did. Reports run before sampling existed read n=1, because that is what they were.

Breadth: keywords per template

How many keywords the template asks.

IntentsGradeCostWhat it buys
-T1-Not on the scale. Narrow is the lowest breadth grade.
3 DefaultT2NarrowNew finds had not slowed yet.×1Finds about a third of a wider run.
10T3Partial*Still finding new things at the end.×3.3
25T4Mostly saturated*New finds had nearly stopped.×8.3
50T5Saturated*The last samples found nothing new.×16.7

* Our estimate, not yet validated by real runs.

Each run measures whether new ground was still appearing when it stopped, so a finished report's breadth grade is measured on that report, not promised from this table.

What the report claims

NumberTypeWhat it can support
Cited, position, brand namedCountedFacts about what the engine returned, never a claim that the engine was right.
AI summaryInterpretedA model's narrative read, marked as AI-generated.
Visibility scoreDerivedNever more certain than the counts underneath it.

Surfaces versus sources

Seven engines read four retrieval systems. The run screen shows that count before you pay.

EngineWhat it reads
Google organic, ChatGPT, Claude, DeepSeekThe one Google result set we fetch for the run. The three models re-rank it, which measures how AI weighs those results, not what corpus it searched.
PerplexityIts own index.
GeminiAn independent web index.
Google AI ModeGoogle's generated answer over its own index.

Engines sharing the fetched result set correlate at 0.54 across 6,830 stored cells, so where intervals are computed their draws count at effective weight rather than face count.

What we do not claim

  • That the engines are right about you.
  • A confidence percentage on any model judgment.
  • One number where the wordings disagree. The rate pools across them, and the report also shows how many wordings found you: a brand found on every phrasing and one found on a lucky phrasing are different facts. That spread is a property of your placement, not an error bar.
  • More than one wording where a number rests on one. It is labelled as one.

Other audits

How we measure, across every audit
George
Online
0%