How we measure
Every number in every Baseline audit answers three questions, and we grade our own answer to each one, on every report, in the open. This page is the general contract; each audit has its own page on what its numbers can support.
The three axes
Did we look everywhere? A brand can rank perfectly on the three questions we asked and be invisible on the forty we did not.
Is the number stable? AI engines answer differently between runs, so a number from one draw is a reading, not a rate.
Were we looking at the real thing? A precise measurement of a stand-in is still a measurement of a stand-in.
The axes are independent. Sampling a report ten times sharpens depth and moves the other two not at all. In 1936 the Literary Digest polled 2.4 million people and called the U.S. election for the wrong man: enormous depth, catastrophic fidelity.
What kind of claim is it
Before a number gets a grade it gets a type, and no amount of sampling changes what the type can support.
| Type | What it is | What it can support |
|---|---|---|
| Verified | A deterministic check of bytes we fetched. No model in the path. | The strongest claim we make. Sampling it is wasted money, so we do not. |
| Counted | What an engine returned: named or not, in which position, cited or not. | A real rate with a real interval. We assert what we counted, never that the engine was right. |
| Interpreted | A model's judgment, like a sentiment read. | Agreement between repeated runs, and nothing else. A model confidently wrong five times out of five would score 100% on a confidence percentage, so we never attach one. |
| Derived | Computed from the above. | No more than its weakest input. A score built on one draw is a one-draw score however elaborate the formula. |
Depth grades
One 1 to 5 scale, so three chips read as three comparable numbers. Counted claims are graded on the width of their 95% Wilson interval, which behaves at small samples where the textbook formula claims impossible certainty: Established is within 15 points either side, Consensus within 30. Interpreted claims are graded on agreement between runs.
| Grade | Meaning |
|---|---|
| T5Verified | Checked directly. Same answer every time. |
| T4Established | A rate within 15 points either way. |
| T4Established | Rerun, and at least 4 in 5 reads matched. |
| T3Consensus | A rate within 30 points either way. |
| T3Consensus | Rerun, and most reads matched. |
| T2Directional | Shows the direction. The rate can move. |
| T2Directional | Rerun, and the reads differed. |
| T1Snapshot | One read. Rerun to confirm it. |
These grades describe our measurement, never your site. Snapshot says we asked once, not that your brand scored badly. On a model judgment, Directional means the runs disagreed: the models have no settled view of the brand, which is a finding in its own right.
Breadth grades
"How many queries cover a sector" has no honest answer. "Are we still finding new things" does, and every run measures it: each extra probe either surfaces something new (a competitor, a citation domain, an issue) or it does not, and the shape of that curve is the coverage claim.
| Grade | Meaning |
|---|---|
| T5Saturated | The last samples found nothing new. |
| T4Mostly saturated | New finds had nearly stopped. |
| T3Partial | Still finding new things at the end. |
| T2Narrow | New finds had not slowed yet. |
| Unknown | Not recorded, or too few samples to tell. |
Unknown is the one grade with no tier. It is our gap, not yours, and the absence of a measurement is not the bottom of a scale.
Fidelity grades
| Grade | Meaning |
|---|---|
| T3Truth | Read from what a real person sees. |
| T2Mixed | Partly the real surface, partly a stand-in. |
| T1Proxy | A stand-in for the real surface. |
Some consumer surfaces cannot be read directly at all. Where we proxy, we say so.
The star
| T4Established | We validated the grade by running the curve. |
| T4Established* | Our current best estimate for this audit, assigned from evidence but not yet tested. |
Every ladder starts starred and earns its way out one audit at a time, as real multi-sample runs accumulate and we check the promise against them. An honest estimate marked as one beats hiding the ladder until every rung is proven.
How we sample
Intents, not wordings
The same prompt rerun overlaps its own results about half the time; a meaning-preserving reword overlaps under a third. A rate attached to one phrasing measures the wording as much as the brand. So we report on the intent, a question a buyer actually has: deeper samples draw fresh paraphrases and the published rate pools across them. When one wording finds you and another does not, that spread is reported too.
Why repeats stop at five
Past about the fifth draw, repeating one wording buys almost nothing: in our decomposition a sixth repeat reduced variance by 0.0003. Budget beyond that goes to fresh paraphrases and more intents.
Engines are not independent, and we count accordingly
In a visibility run, Google organic, ChatGPT, Claude and DeepSeek all read the one Google result set we fetch, so their agreement is one reading confirmed, not four. Over 6,830 stored report cells, engines sharing a fetched result set correlate at 0.54; engines on different corpora at 0.2 or less. Correlated draws count at their effective weight, and every run picker shows surfaces selected against retrieval systems actually read.
Temperature
We pin model temperature to 0 wherever we control it. That reduces variance but does not make runs repeatable, because providers route, batch and update models under fixed names, so we record which model and provider answered each call.
Per audit
Where each audit's ladders top out on the one scale. Each page shows the full ladder and what every rung costs.
| Audit | Top depth rung | Top breadth rung |
|---|---|---|
| AI Visibility | T3ConsensusA rate within 30 points either way. | T5Saturated*The last samples found nothing new. |
| PR & Brand | T2Directional*Rerun, and the reads differed. | T3Partial*Still finding new things at the end. |
| Site Pulse | T5VerifiedChecked directly. Same answer every time. | T5Saturated*The last samples found nothing new. |
| Schema Audit | T5VerifiedChecked directly. Same answer every time. | T4Mostly saturated*New finds had nearly stopped. |
| Reputation | T4Established*Rerun, and at least 4 in 5 reads matched. | T4Mostly saturated*New finds had nearly stopped. |
| AI Vision | T4EstablishedRerun, and at least 4 in 5 reads matched. | T5Saturated*The last samples found nothing new. |
| Active Competitors | T3Consensus*A rate within 30 points either way. | T3Partial*Still finding new things at the end. |
| Fan-out | T4Established*A rate within 15 points either way. | T3Partial*Still finding new things at the end. |