How we measure AI Vision
An AI Vision scan crawls your site and runs ten checks. Seven read the bytes we fetched and nothing else; three ask AI models what they already know about you. The report keeps the two halves apart.
Depth: repeat judgments per model check
Only the three model checks repeat. The crawl happens once and the seven page checks read bytes already in hand, so neither is resampled or charged for.
| Repeat judgments per model-judged check | Grade | Cost | What it buys |
|---|---|---|---|
| 1 Default | T1SnapshotOne read. Rerun to confirm it. | ×1 | |
| - | T2Directional | - | Not offered. The ladder steps from T1 straight to T4. |
| - | T3Consensus | - | Not offered. The ladder steps from T1 straight to T4. |
| 2 | T4EstablishedRerun, and at least 4 in 5 reads matched. | ×1.38 | Of 19 rerun scans, 17 matched on every check. |
| 3 | T4EstablishedRerun, and at least 4 in 5 reads matched. | ×1.75 | Of 19 reruns, 17 Established and 2 Consensus. |
| 5 | T4EstablishedRerun, and at least 4 in 5 reads matched. | ×2.5 | Of 19 reruns, 18 Established and 1 Consensus. |
| - | T5Verified | - | Not offered. |
Why it stops at T4. Agreement between model reads is not proof they are right.
The grade does not climb with repeats: what they buy is the floor. At two, one odd answer halves the agreement and the check reads Directional; at three it reads Consensus. The evidence is 129 repeat reads across 18 subjects in stored scans, pooled agreement 0.97 and pairwise 0.94, measured between scans days or weeks apart across model versions, so repeats inside one run should agree at least that often.
Disagreement is recorded as a neutral finding, never against your score: agreement findings sit outside every score denominator. A template with no model-judged module or no engine selected runs once and is billed once whatever you asked for. Repeats of one run ask the same category questions, pinned from the first repeat, because rewording moves an answer far more than asking again does.
Breadth: pages crawled
What each page is counted for is an entity fact: a schema type, or a type and one of its properties. Never the property's value, because every product page names a different product and a value-keyed measure would make every catalogue look unexplored.
| Pages | Grade | Cost | What it buys |
|---|---|---|---|
| 1 | UnknownNot recorded, or too few samples to tell. | ×1 | Home page only. |
| - | T1 | - | Not on the scale. Narrow is the lowest breadth grade. |
| 5 | T2Narrow*New finds had not slowed yet. | ×1.6 | |
| 14 | T3Partial*Still finding new things at the end. | ×1.6 | The median stored scan. |
| 50 Default | T4Mostly saturated*New finds had nearly stopped. | ×1.6 | The standard cap. |
| 250 | T5Saturated*The last samples found nothing new. | ×3 |
* Our estimate, not yet validated by real runs.
The run grades its own curve: new facts per page, in crawl order, with a page that declared nothing kept as a real zero. The grade on your report is measured on your site and can land below or above the rung you paid for.
A crawl that finished is a census, not a sample. When the crawler followed every same-origin link and sitemap entry and stopped below its page cap, the run grades Saturated on having finished, says so, and keeps the curve's own grade beside it. A page nothing links to and the sitemap omits is invisible to every crawler, ours included.
Two refusals hold. A scan that found nothing anywhere reads Unknown, never Saturated, even when the crawl finished. A crawl stopped by its page cap or by the daily request budget we hold for your server is not a census, and the report says so beside the grade. Scans run before this existed read Unknown and always will: 1,596 of the 1,639 findings stored before the change carry no page, so there is no curve to rebuild, and inferring one would flatter us.
What the report claims
| Number | Type | What it can support |
|---|---|---|
| Bot and crawler access, llms.txt quality, Organization schema, structured data coverage, rich result and E-E-A-T signals, content clarity, AI indexability | Verified | Rules applied to markup we hold. A second run over the same bytes returns the same answer, so repeating them buys nothing. |
| Engines that recognised your domain, cited URLs pointing at your site | Counted | Facts about what came back, never that the engine was right. A failed engine, or one answered by a substitute model, is named and counted in neither side of any fraction. |
| What a model says your business is, its industry and audience, gaps, category mentions, names alongside yours | Interpreted | Agreement between repeated runs and nothing else. The category-name list is a lead list, not competitor discovery. |
| Module scores, AI Vision Score, AI verdict | Derived | Never firmer than the weakest input. A module that could not take its measurement scores nothing and drops out of the average, its weight redistributed. |
Fidelity: page checks and model checks
| Checks | What is read |
|---|---|
| Seven page checks | The real surface. Your pages fetched with a declared crawler identity at a paced rate, and the markup that arrives graded. |
| Three model checks | A proxy. A model asked through an API what it knows about your domain from training, temperature pinned to zero. A consumer chat app adds a session, a system prompt and often live retrieval we do not see. |
Temperature zero reduces how far an answer wanders between calls; it does not make it repeatable, and it is no evidence the model is right. Each run records which model actually answered: when a model family is exhausted our provider substitutes another, and the stand-in's answer is reported under its own name and left out of every count.
What we do not claim
- That a model is right about you.
- A confidence percentage on any model judgment. We do not ask a model how confident it is either: a model writing "87" about itself generates that number the way it generates the rest of the sentence.
- A rate from a count. Four engines asked once tells you what four engines said that day, not how often they say it.
- A full mark from an absence. A check that could not be taken shows nothing rather than a hundred.
- Our outages as your result. A failed engine is reported as failed and counted neither way.
- Competitor discovery. Names alongside yours are a lead list; competitor work lives in the Active Competitors audit.
- More than one judgment where a number rests on one. It is labelled as one.