Citation accuracy
Citation accuracy is the share of AI-generated answers whose attributions name the right source and resolve to a working page. A 2025 Tow Center test of eight tools over 1,600 queries returned incorrect answers to more than 60% of them[1], with per-tool rates from 37% to 94%[2]. Two failure modes are counted separately: naming the wrong publisher, and supplying a URL that is fabricated or dead[3].
The Tow Center tests
The Tow Center for Digital Journalism ran two controlled tests. The first, in November 2024, fed ChatGPT 200 verbatim quotes drawn from 10 articles at each of 20 publishers and asked it to identify the source. It returned partially or entirely wrong attributions 153 times and admitted it could not answer on 7 occasions[4]. The second, in March 2025, repeated the design across eight chatbots, for 20 publishers times 10 articles times 8 tools, or 1,600 queries in total[5]. Perplexity was least wrong at 37%, Grok-3 most wrong at 94%[2]. Paid tiers answered more prompts correctly than free ones and still scored a higher error rate, because they replaced refusals with confident wrong answers[6].
Mention accuracy against link accuracy
An answer can attribute a fact to the correct publisher and still hand the reader a broken link, so the two are measured apart. Grok-3 produced 154 citations that led to error pages out of 200 responses[3]. A separate error names a real page that is the wrong copy: tools cited syndicated republications on aggregator sites and, in one case, a plagiarised copy of a New York Times article[7].
Replications
A 2025 study by 22 public service media organisations across 18 countries and 14 languages graded more than 3,000 assistant responses. It found a significant issue in 45%, serious sourcing problems in 31% and major accuracy problems in 20%[8]. Results held across languages and territories[8]. Gemini scored worst, with significant issues in 76% of its responses[9].
AI Overviews
Google's AI Overviews are audited on claim fidelity rather than article identification. An academic audit of 55,393 trending queries split the overviews into 98,020 atomic claims and found 11.0% unsupported by the cited pages, with omission the dominant failure[10]. Close to 30% of cited domains never appear in the first-page results beside the overview[11], so the citation pool is not the ranking set. A 2026 evaluation over about 4,000 factual questions found 91% of overviews carried the correct answer while only 39% were both correct and fully supported by their sources[12].
Data
- 91%Contains the correct answer
- 39%Correct and fully supported
Show the numbers (2)
| Point | % |
|---|---|
| Contains the correct answer | 91 |
| Correct and fully supported | 39 |
References (12)
- Jazwinska and Chandrasekar, AI Search Has A Citation Problem, Tow Center for Digital Journalism Archive
- Jazwinska and Chandrasekar, AI Search Has A Citation Problem, Tow Center for Digital Journalism Archive
- Jazwinska and Chandrasekar, AI Search Has A Citation Problem, Tow Center for Digital Journalism Archive
- Jazwinska and Chandrasekar, How ChatGPT Misrepresents Publisher Content, Tow Center for Digital Journalism
- AI Search Has a Citation Problem
- Jazwinska and Chandrasekar, AI Search Has A Citation Problem, Tow Center for Digital Journalism Archive
- Jazwinska and Chandrasekar, How ChatGPT Misrepresents Publisher Content, Tow Center for Digital Journalism Archive
- News Integrity in AI Assistants, European Broadcasting Union and BBC Archive
- News Integrity in AI Assistants, European Broadcasting Union and BBC Original
- Xu, Iqbal and Montgomery, Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact, arXiv:2605.14021 Archive
- Xu, Iqbal and Montgomery, Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact, arXiv:2605.14021 Archive
- Oumi, Study finds half of AI Overviews untrustworthy Archive
Last updated 2026-09-04. Written and maintained by Baseline Labs.