Skip to content

How we measure

Every number in every Baseline audit answers three questions, and we grade our own answer to each one, on every report, in the open. This page is the general contract; each audit has its own page on what its numbers can support.

The three axes

Breadth

Did we look everywhere? A brand can rank perfectly on the three questions we asked and be invisible on the forty we did not.

Depth

Is the number stable? AI engines answer differently between runs, so a number from one draw is a reading, not a rate.

Fidelity

Were we looking at the real thing? A precise measurement of a stand-in is still a measurement of a stand-in.

The axes are independent. Sampling a report ten times sharpens depth and moves the other two not at all. In 1936 the Literary Digest polled 2.4 million people and called the U.S. election for the wrong man: enormous depth, catastrophic fidelity.

What kind of claim is it

Before a number gets a grade it gets a type, and no amount of sampling changes what the type can support.

TypeWhat it isWhat it can support
VerifiedA deterministic check of bytes we fetched. No model in the path.The strongest claim we make. Sampling it is wasted money, so we do not.
CountedWhat an engine returned: named or not, in which position, cited or not.A real rate with a real interval. We assert what we counted, never that the engine was right.
InterpretedA model's judgment, like a sentiment read.Agreement between repeated runs, and nothing else. A model confidently wrong five times out of five would score 100% on a confidence percentage, so we never attach one.
DerivedComputed from the above.No more than its weakest input. A score built on one draw is a one-draw score however elaborate the formula.

Depth grades

One 1 to 5 scale, so three chips read as three comparable numbers. Counted claims are graded on the width of their 95% Wilson interval, which behaves at small samples where the textbook formula claims impossible certainty: Established is within 15 points either side, Consensus within 30. Interpreted claims are graded on agreement between runs.

GradeMeaning
T5VerifiedChecked directly. Same answer every time.
T4EstablishedA rate within 15 points either way.
T4EstablishedRerun, and at least 4 in 5 reads matched.
T3ConsensusA rate within 30 points either way.
T3ConsensusRerun, and most reads matched.
T2DirectionalShows the direction. The rate can move.
T2DirectionalRerun, and the reads differed.
T1SnapshotOne read. Rerun to confirm it.

These grades describe our measurement, never your site. Snapshot says we asked once, not that your brand scored badly. On a model judgment, Directional means the runs disagreed: the models have no settled view of the brand, which is a finding in its own right.

Breadth grades

"How many queries cover a sector" has no honest answer. "Are we still finding new things" does, and every run measures it: each extra probe either surfaces something new (a competitor, a citation domain, an issue) or it does not, and the shape of that curve is the coverage claim.

GradeMeaning
T5SaturatedThe last samples found nothing new.
T4Mostly saturatedNew finds had nearly stopped.
T3PartialStill finding new things at the end.
T2NarrowNew finds had not slowed yet.
UnknownNot recorded, or too few samples to tell.

Unknown is the one grade with no tier. It is our gap, not yours, and the absence of a measurement is not the bottom of a scale.

Fidelity grades

GradeMeaning
T3TruthRead from what a real person sees.
T2MixedPartly the real surface, partly a stand-in.
T1ProxyA stand-in for the real surface.

Some consumer surfaces cannot be read directly at all. Where we proxy, we say so.

The star

T4EstablishedWe validated the grade by running the curve.
T4Established*Our current best estimate for this audit, assigned from evidence but not yet tested.

Every ladder starts starred and earns its way out one audit at a time, as real multi-sample runs accumulate and we check the promise against them. An honest estimate marked as one beats hiding the ladder until every rung is proven.

How we sample

Intents, not wordings

The same prompt rerun overlaps its own results about half the time; a meaning-preserving reword overlaps under a third. A rate attached to one phrasing measures the wording as much as the brand. So we report on the intent, a question a buyer actually has: deeper samples draw fresh paraphrases and the published rate pools across them. When one wording finds you and another does not, that spread is reported too.

Why repeats stop at five

Past about the fifth draw, repeating one wording buys almost nothing: in our decomposition a sixth repeat reduced variance by 0.0003. Budget beyond that goes to fresh paraphrases and more intents.

Engines are not independent, and we count accordingly

In a visibility run, Google organic, ChatGPT, Claude and DeepSeek all read the one Google result set we fetch, so their agreement is one reading confirmed, not four. Over 6,830 stored report cells, engines sharing a fetched result set correlate at 0.54; engines on different corpora at 0.2 or less. Correlated draws count at their effective weight, and every run picker shows surfaces selected against retrieval systems actually read.

Temperature

We pin model temperature to 0 wherever we control it. That reduces variance but does not make runs repeatable, because providers route, batch and update models under fixed names, so we record which model and provider answered each call.

Per audit

Where each audit's ladders top out on the one scale. Each page shows the full ladder and what every rung costs.

AuditTop depth rungTop breadth rung
AI VisibilityT3ConsensusA rate within 30 points either way.T5Saturated*The last samples found nothing new.
PR & BrandT2Directional*Rerun, and the reads differed.T3Partial*Still finding new things at the end.
Site PulseT5VerifiedChecked directly. Same answer every time.T5Saturated*The last samples found nothing new.
Schema AuditT5VerifiedChecked directly. Same answer every time.T4Mostly saturated*New finds had nearly stopped.
ReputationT4Established*Rerun, and at least 4 in 5 reads matched.T4Mostly saturated*New finds had nearly stopped.
AI VisionT4EstablishedRerun, and at least 4 in 5 reads matched.T5Saturated*The last samples found nothing new.
Active CompetitorsT3Consensus*A rate within 30 points either way.T3Partial*Still finding new things at the end.
Fan-outT4Established*A rate within 15 points either way.T3Partial*Still finding new things at the end.
George
Online
0%