Skip to content

How we measure PR & Brand

A scan writes search questions from your business profile, asks Google each one, and has a model read every result. It counts where you are mentioned and interprets how you are described, and grades those two claims separately.

Depth: wordings per question

How many wordings of each question the scan searches. Draw one is the question on your template; the rest are rewordings in your study's language, pooled under the question. Depth buys the sentiment read, which cannot be graded at all at one draw - the state of every scan run so far.

Wordings per questionGradeCostWhat it buys
1 DefaultT1SnapshotOne read. Rerun to confirm it.×1
2T2Directional*Rerun, and the reads differed.×2
3T2Directional*Rerun, and the reads differed.×3
-T3Consensus-Not offered.
-T4Established-Not offered.
-T5Verified-Not offered.

Why it stops at T2. Draws past three add half as much for twice the price.

* Our estimate, not yet validated by real runs.

Wordings, not repeats: two runs of one phrasing overlap about half to three fifths of the time, a meaning-preserving reword under a third, and wording carried 26.5% of the variance against the brand's own identity at 1.5%. Rewordings are written in your market by a cheap model held to a rubric (same meaning, no new constraints, places or brands). Any that come back as the original, a shuffle, a duplicate or an answer are thrown away; the draw then repeats the original, and a run whose rewordings all fell through is labelled repeats.

The mention rate barely moves with depth. Draws of one question pool as one cluster, so three count for about 1.44 independent reads: on the median scan of 52 questions the rate lands at about plus or minus 13 points at one draw and 11 at three. That correction was measured on repeats, and distinct wordings are less alike, so the real gain is larger; we publish the smaller figure until the wording correlation is measured on this audit, which is why every rung above one carries a star.

Two draws can only score 1.00 or 0.50, so two and three draws promise the same floor; your report shows what it actually measured, which may be Established. A fourth and fifth draw add 0.139 effective reads per question between them, against 0.298 for the first extra draw. The cost multiple is the multiple charged: writing the questions happens once, a few per cent of a scan, and is not discounted.

Breadth: questions per scan

How many questions the scan asks - on this audit, the axis carrying the weight. Each rung is a search depth on your template at the credits it is charged, and the count is the measured median, not a cap.

QuestionsGradeCostWhat it buys
-T1-Not on the scale. Narrow is the lowest breadth grade.
25T2NarrowNew finds had not slowed yet.×1Quick, 110 credits. Mention rate plus or minus 18.
52 DefaultT2NarrowNew finds had not slowed yet.×1.36Standard, 150 credits. Mention rate plus or minus 13.
100T3Partial*Still finding new things at the end.×2.91Deep, 320 credits.
-T4Mostly saturated-Not offered.
-T5Saturated-Not offered.

Why it stops at T3. New sources kept arriving at every size run so far.

* Our estimate, not yet validated by real runs.

Questions per scanSources per question
22 to 241.53
50 to 541.39
55 to 591.38

Across 720 completed scans, sources found per question barely falls as the question set grows; the marginal rate between the ends is 1.21 new sources per extra question. The stars here are inferred from totals, because no scan recorded which question found which result. Each scan now records its own curve over the domains of admitted mentions and stamps its grade. Questions return in network order, so it is a rarefaction curve: how much we would have missed stopping earlier, not a planned sequence. Errored searches count as dry samples, nudging towards Saturated, and the report states how many errored.

What the report claims

NumberTypeWhat it can support
Mention rateCountedSearches that found you divided by searches that came back. It used to read 100% on every report, because the denominator had been overwritten with the mention count. Errored searches are kept out: an outage is not you being absent.
Sentiment, stance, overall toneInterpretedGraded by how often independent passes agree, never by a confidence score. At one draw, a single read.
Source breakdown, sentiment mix, positioning, summaryDerivedNo firmer than the weakest claim each rests on.

Fidelity: mixed

What we readHow close to the real surface
Search resultsGoogle's real index, fetched with your study's country and language, so what a searcher there sees. It is not an AI engine's answer: this is the material answers are built from. The AI Visibility report asks engines directly.
Top-ranked pagesFetched as GeorgeBot, honouring robots.txt, and the model reads the opening text.
Other resultsJudged from title and snippet, which is thinner evidence.
Every kept mentionBacked by a sentence quoted from the source. A quote not in the document discards the mention.
Publication datesFrom the search index or the page's markup, never the model: a wrong date on a trend is worse than a gap.

Why old trends are undefined

Until now every scan wrote fresh questions at a temperature chosen to vary them, and nothing recorded them, so no two scans of one brand asked the same things. A line between two of them compares two instruments. Our data shows it: ten pairs of scans of one template run within two days share only 55% of their sources, and two pairs run six hours apart share 56% and 61%.

Every scan now records its questions, the model that wrote them and the version of the admission rules, and can rerun the previous question set. That is off by default: fresh questions are what discovery needs. Freeze the set for a trend, leave it free for discovery. Older scans read as one draw with no question set recorded, and are not backfilled.

What we do not claim

  • That this is what AI engines say about you. It is what search returns.
  • That we found everything written about you. On this audit the breadth grade usually says more was still arriving.
  • A confidence percentage on anything a model judged. A model confidently wrong three times running agrees perfectly.
  • That two scans are comparable, unless one reran the other's questions. The report says which.
  • That a draw is an independent look. Draws of one question are pooled as one cluster and the range is widened accordingly.
  • That every draw was reworded. The report says how many fell back to the original.

Other audits

How we measure, across every audit
George
Online
0%