Skip to content

How we measure fan-out

Each seed question goes to each engine you select, and the report records the searches the engine ran while answering, or, where an engine hides them, the questions it surfaces to a reader.

Depth: draws per seed question

How many separate times each engine is asked the same thing. This is the tallest ladder in the product, because an engine asked twice does not run the same searches twice: here a repeat draw is a new observation, not a receipt.

Draws per seed questionGradeCostWhat it buys
-T1Snapshot-Not offered. The first rung already reaches T2.
1T2DirectionalShows the direction. The rate can move.×1No recurrence to read.
2 DefaultT2DirectionalShows the direction. The rate can move.×2The default. Plus or minus 40 points.
5T2Directional*Shows the direction. The rate can move.×5Plus or minus 33 points.
10T3Consensus*A rate within 30 points either way.×10Plus or minus 26 points.
25T3Consensus*A rate within 30 points either way.×25Plus or minus 18 points.
50T4Established*A rate within 15 points either way.×50Plus or minus 13 points.
100T4Established*A rate within 15 points either way.×100Plus or minus 10 points.
-T5Verified-Not offered.

Why it stops at T4. The top grade is for deterministic checks, not sampled answers.

* Our estimate, not yet validated by real runs.

Every extra draw is another paid provider call: one draw across all four engines is 470 credits per seed question, and ChatGPT alone is 280 a draw. The price is printed on the control before you spend it. One draw is unstarred because it cannot produce a recurrence rate by construction; two is unstarred because it is what the product has run since the ChatGPT upgrade in August, across fifteen stored seeds.

Why repeats do not stop at five here

A proportion settles after about five draws. Fan-out measures a set, and every draw can add a member. One seed, drawn one at a time on 22 September 2026:

EngineNew questions per drawDistinct
Gemini, 5 draws5, 4, 3, 5, 320
ChatGPT, 3 draws3, 3, 410

Neither curve has started to bend, and not one of ChatGPT's ten questions repeated word for word. On a surface that rewords itself, exact-match recurrence sits near one in n however deep you go: depth buys ChatGPT reach, not frequency.

Google is read once at every depth and priced once: three scans of one question two minutes apart returned the identical six questions, eighteen rows of eighteen. Perplexity is drawn and charged like the rest, because its Related block is written per answer and nobody has yet measured whether it repeats.

Breadth: seed questions

How many different things you ask. It is the least flattering ladder we publish, and that is the finding.

Seed questionsGradeCostWhat it buys
1UnknownNot recorded, or too few samples to tell.×1
-T1-Not on the scale. Narrow is the lowest breadth grade.
5T2Narrow*New finds had not slowed yet.×5
10T2Narrow*New finds had not slowed yet.×10
20 DefaultT2Narrow*New finds had not slowed yet.×20The template default.
50T3Partial*Still finding new things at the end.×50The cap.
-T4Mostly saturated-Not offered.
-T5Saturated-Not offered.

Why it stops at T3. No template has been seen to run out of sub-questions.

* Our estimate, not yet validated by real runs.

No stored scan has ever run out of new sub-questions. In one twenty-question scan the twentieth question added twelve unseen sub-questions, as many as the first; in a seven-question scan the sixth was the largest of the run.

Every scan records its own novelty curve over draws. A draw that found nothing new is kept as a real zero, the only evidence of saturation there is; a draw where no engine answered is dropped and disclosed, so an outage on the last draws cannot pass for having found everything.

How much of the question space a run has seen

Each finished scan estimates how many distinct sub-questions sit behind it, the ecologist's way: questions seen in exactly one draw carry the information about questions seen in none. It gives the smallest plausible population, so the share we report is a best case, and the report says so. The five-draw Gemini run above estimates about sixty-eight questions behind its seed, twenty seen: roughly twenty-nine per cent at best.

A September review first measured this across engine-and-question pairs and found about seven per cent coverage. Engines and questions barely overlap, so that was measuring how different the engines are. Coverage now comes from repeat draws of one question on one engine, the only probe that samples the same population twice. Google, drawn once, is left out of the estimate and the report says how many rows were excluded.

What the report claims

NumberTypeWhat it can support
Questions capturedCountedThe size of a deduplicated union: every distinct question any draw of any engine produced, once. A bigger union means we reached more, never that anything came back more often.
How often a question came backCountedA rate over the draws we made, failed draws included. Matched on exact wording, and a rate about the engine, not your brand.
Themes, summary, headlineInterpretedOne model pass over the captured list, a single read with no confidence figure. Nothing else depends on it.
Coverage estimateDerivedInherits every weakness of the counts it is built from.

Fidelity

EngineWhat we read
GeminiThe grounding metadata on a grounded answer: searches the model really ran.
ChatGPTThe web-search calls in the API response: searches the model really ran.
Perplexity, GoogleStood in for. The Related block, People Also Ask and related searches are written for a human, not run against an index: the question space around your topic, not retrieval.

The API is close to the consumer app but not the same surface, and a personalised chat session may fan out differently. Gemini never returns a citation's source URL, only a redirect wrapper Google expires, so we recover the domain from the citation title: the domain lasts, the link will rot.

What we do not claim

  • That these are the searches your customer's chat session would run. They are what the provider API ran for us, at a point in time.
  • That Perplexity or Google fan out this way internally. We say which rows are real searches and which are surfaced questions.
  • A count as a rate. Questions captured is a union size and pages a domain was cited on is a count; the one frequency is labelled with its own denominator.
  • That a scan has seen the whole question space. Where a run is still finding new things at its last draw, the report says so.
  • Depth or coverage for scans run before depth was recorded. Fan-out drew a different number of times in different months, per engine, so we leave the history blank.

Other audits

How we measure, across every audit
George
Online
0%