How we measure AI visibility
Every engine you select is asked every keyword in your template. Each answer records whether your domain is cited, at what position, and whether your brand is named at all.
Depth: wordings per question
How many wordings of each question the run asks. Draw one is your question as written; the rest are meaning-preserving rewordings in your template's language, asked on every engine you selected and pooled under the question.
| Draws per intent | Grade | Cost | What it buys |
|---|---|---|---|
| - | T1Snapshot | - | Not offered. The first rung already reaches T2. |
| 1 Default | T2DirectionalShows the direction. The rate can move. | ×1 | Plus or minus 45 points, reads about 15 high. |
| 2 | T2DirectionalShows the direction. The rate can move. | ×2 | Plus or minus 25 points. |
| 3 | T2DirectionalShows the direction. The rate can move. | ×3 | Plus or minus 25 points, keyword bias about 5. |
| 5 | T3ConsensusA rate within 30 points either way. | ×5 | Plus or minus 19 points. |
| - | T4Established | - | Not offered. |
| - | T5Verified | - | Not offered. |
Why it stops at T3. Five wordings measured plus or minus 19 points; the next grade needs 15.
Wording carried 26.5% of total variance against brand identity's 1.5%, so depth buys new wordings rather than repeats.
Tested on 20 questions across 6 studies, in English and Swedish, on 5 engines. Each run was scored against a second run's fresh wordings, so the error is to the question, not to the keyword.
| Wordings | Plus or minus, mention rate | Plus or minus, citation rate | Reads high by |
|---|---|---|---|
| 1 | 42 points | 51 points | 13 to 16 points |
| 2 | 25 points | 24 points | 6 to 8 points |
| 3 | 25 points | 24 points | 5 to 7 points |
| 5 | 19 points | 16 points | 3 points |
Your question as written reads high because a template is written in the words a brand already ranks for: on 12 of 20 questions the original wording found the brand more often than its rewordings, and on 2 it found it less. Running the same wording again agrees with itself far more closely than a new wording does, so a stable number is not the same as an accurate one. Five wordings stop short of Established, which needs 15 points.
Google is resampled too, because a reworded question is a new query with its own result set. Within one wording the Google result set is fetched once and shared by Google, ChatGPT, Claude and DeepSeek, so their disagreement is about the models, not fetch timing. A rewording we cannot produce falls back to repeating the original, and the report's method says how many did. Reports run before sampling existed read n=1, because that is what they were.
Breadth: keywords per template
How many keywords the template asks.
| Intents | Grade | Cost | What it buys |
|---|---|---|---|
| - | T1 | - | Not on the scale. Narrow is the lowest breadth grade. |
| 3 Default | T2NarrowNew finds had not slowed yet. | ×1 | Finds about a third of a wider run. |
| 10 | T3Partial*Still finding new things at the end. | ×3.3 | |
| 25 | T4Mostly saturated*New finds had nearly stopped. | ×8.3 | |
| 50 | T5Saturated*The last samples found nothing new. | ×16.7 |
* Our estimate, not yet validated by real runs.
Each run measures whether new ground was still appearing when it stopped, so a finished report's breadth grade is measured on that report, not promised from this table.
What the report claims
| Number | Type | What it can support |
|---|---|---|
| Cited, position, brand named | Counted | Facts about what the engine returned, never a claim that the engine was right. |
| AI summary | Interpreted | A model's narrative read, marked as AI-generated. |
| Visibility score | Derived | Never more certain than the counts underneath it. |
Surfaces versus sources
Seven engines read four retrieval systems. The run screen shows that count before you pay.
| Engine | What it reads |
|---|---|
| Google organic, ChatGPT, Claude, DeepSeek | The one Google result set we fetch for the run. The three models re-rank it, which measures how AI weighs those results, not what corpus it searched. |
| Perplexity | Its own index. |
| Gemini | An independent web index. |
| Google AI Mode | Google's generated answer over its own index. |
Engines sharing the fetched result set correlate at 0.54 across 6,830 stored cells, so where intervals are computed their draws count at effective weight rather than face count.
What we do not claim
- That the engines are right about you.
- A confidence percentage on any model judgment.
- One number where the wordings disagree. The rate pools across them, and the report also shows how many wordings found you: a brand found on every phrasing and one found on a lucky phrasing are different facts. That spread is a property of your placement, not an error bar.
- More than one wording where a number rests on one. It is labelled as one.