How we measure PR & Brand
A scan writes search questions from your business profile, asks Google each one, and has a model read every result. It counts where you are mentioned and interprets how you are described, and grades those two claims separately.
Depth: wordings per question
How many wordings of each question the scan searches. Draw one is the question on your template; the rest are rewordings in your study's language, pooled under the question. Depth buys the sentiment read, which cannot be graded at all at one draw - the state of every scan run so far.
| Wordings per question | Grade | Cost | What it buys |
|---|---|---|---|
| 1 Default | T1SnapshotOne read. Rerun to confirm it. | ×1 | |
| 2 | T2Directional*Rerun, and the reads differed. | ×2 | |
| 3 | T2Directional*Rerun, and the reads differed. | ×3 | |
| - | T3Consensus | - | Not offered. |
| - | T4Established | - | Not offered. |
| - | T5Verified | - | Not offered. |
Why it stops at T2. Draws past three add half as much for twice the price.
* Our estimate, not yet validated by real runs.
Wordings, not repeats: two runs of one phrasing overlap about half to three fifths of the time, a meaning-preserving reword under a third, and wording carried 26.5% of the variance against the brand's own identity at 1.5%. Rewordings are written in your market by a cheap model held to a rubric (same meaning, no new constraints, places or brands). Any that come back as the original, a shuffle, a duplicate or an answer are thrown away; the draw then repeats the original, and a run whose rewordings all fell through is labelled repeats.
The mention rate barely moves with depth. Draws of one question pool as one cluster, so three count for about 1.44 independent reads: on the median scan of 52 questions the rate lands at about plus or minus 13 points at one draw and 11 at three. That correction was measured on repeats, and distinct wordings are less alike, so the real gain is larger; we publish the smaller figure until the wording correlation is measured on this audit, which is why every rung above one carries a star.
Two draws can only score 1.00 or 0.50, so two and three draws promise the same floor; your report shows what it actually measured, which may be Established. A fourth and fifth draw add 0.139 effective reads per question between them, against 0.298 for the first extra draw. The cost multiple is the multiple charged: writing the questions happens once, a few per cent of a scan, and is not discounted.
Breadth: questions per scan
How many questions the scan asks - on this audit, the axis carrying the weight. Each rung is a search depth on your template at the credits it is charged, and the count is the measured median, not a cap.
| Questions | Grade | Cost | What it buys |
|---|---|---|---|
| - | T1 | - | Not on the scale. Narrow is the lowest breadth grade. |
| 25 | T2NarrowNew finds had not slowed yet. | ×1 | Quick, 110 credits. Mention rate plus or minus 18. |
| 52 Default | T2NarrowNew finds had not slowed yet. | ×1.36 | Standard, 150 credits. Mention rate plus or minus 13. |
| 100 | T3Partial*Still finding new things at the end. | ×2.91 | Deep, 320 credits. |
| - | T4Mostly saturated | - | Not offered. |
| - | T5Saturated | - | Not offered. |
Why it stops at T3. New sources kept arriving at every size run so far.
* Our estimate, not yet validated by real runs.
| Questions per scan | Sources per question |
|---|---|
| 22 to 24 | 1.53 |
| 50 to 54 | 1.39 |
| 55 to 59 | 1.38 |
Across 720 completed scans, sources found per question barely falls as the question set grows; the marginal rate between the ends is 1.21 new sources per extra question. The stars here are inferred from totals, because no scan recorded which question found which result. Each scan now records its own curve over the domains of admitted mentions and stamps its grade. Questions return in network order, so it is a rarefaction curve: how much we would have missed stopping earlier, not a planned sequence. Errored searches count as dry samples, nudging towards Saturated, and the report states how many errored.
What the report claims
| Number | Type | What it can support |
|---|---|---|
| Mention rate | Counted | Searches that found you divided by searches that came back. It used to read 100% on every report, because the denominator had been overwritten with the mention count. Errored searches are kept out: an outage is not you being absent. |
| Sentiment, stance, overall tone | Interpreted | Graded by how often independent passes agree, never by a confidence score. At one draw, a single read. |
| Source breakdown, sentiment mix, positioning, summary | Derived | No firmer than the weakest claim each rests on. |
Fidelity: mixed
| What we read | How close to the real surface |
|---|---|
| Search results | Google's real index, fetched with your study's country and language, so what a searcher there sees. It is not an AI engine's answer: this is the material answers are built from. The AI Visibility report asks engines directly. |
| Top-ranked pages | Fetched as GeorgeBot, honouring robots.txt, and the model reads the opening text. |
| Other results | Judged from title and snippet, which is thinner evidence. |
| Every kept mention | Backed by a sentence quoted from the source. A quote not in the document discards the mention. |
| Publication dates | From the search index or the page's markup, never the model: a wrong date on a trend is worse than a gap. |
Why old trends are undefined
Until now every scan wrote fresh questions at a temperature chosen to vary them, and nothing recorded them, so no two scans of one brand asked the same things. A line between two of them compares two instruments. Our data shows it: ten pairs of scans of one template run within two days share only 55% of their sources, and two pairs run six hours apart share 56% and 61%.
Every scan now records its questions, the model that wrote them and the version of the admission rules, and can rerun the previous question set. That is off by default: fresh questions are what discovery needs. Freeze the set for a trend, leave it free for discovery. Older scans read as one draw with no question set recorded, and are not backfilled.
What we do not claim
- That this is what AI engines say about you. It is what search returns.
- That we found everything written about you. On this audit the breadth grade usually says more was still arriving.
- A confidence percentage on anything a model judged. A model confidently wrong three times running agrees perfectly.
- That two scans are comparable, unless one reran the other's questions. The report says which.
- That a draw is an independent look. Draws of one question are pooled as one cluster and the range is widened accordingly.
- That every draw was reworded. The report says how many fell back to the original.