Content licensing deals
Content licensing deals are paid agreements under which a publisher or platform lets an AI company use its material for model training, for display inside generated answers, or both. Buyers include OpenAI, Google, Meta, Microsoft, Amazon and Perplexity, and reported values run from $10m spread across five newsrooms through the Lenfest Institute[1] to more than $250m over five years for News Corp[2]. Around a dozen claims against OpenAI have been filed since the end of 2023, so the market is being built by contract and by lawsuit at the same time[3]. Whether a licence buys citation is a separate question: Google ties eligibility to indexing rather than to contracts[4], and a published audit of licensed publishers found no attribution advantage[5].
What the money buys
Two distinct rights are on sale, often in the same contract. Training rights cover ingesting an archive into a model's weights. Display rights cover summarising live articles and linking to them inside an answer product, which is the half that can return traffic. Under the Associated Press agreement with Google, AP will "deliver a feed of real-time information to help further enhance the usefulness of results displayed in the Gemini app"[6]. Nine's agreement with Microsoft lets "Microsoft Copilot use its journalism in real time when creating AI answers"[7], and Meta's Reuters deal is framed the same way[8]. A third shape pays per use instead of per archive: signed publishers under Perplexity's programme "will be able to share the revenue generated by interactions where their content is referenced"[9].
Reported values
Most deals disclose nothing. Of those that leak a figure, News Corp's OpenAI agreement was reported at "more than $250m over five years"[2], its later Meta agreement at "up to $50m per year for at least three years"[10], and the New York Times deal with Amazon at "$20m to $25m per year"[11]. Firmer numbers come from filings. Reddit told the SEC in February 2024 that it had "entered into certain data licensing arrangements with an aggregate contract value of $203.0 million and terms ranging from two to three years"[12]. The revenue line those contracts sit in went from $15.2m in 2023 to $140.0m in 2025[13], and $143.7m of future revenue was still under contract at the end of 2025[14].
Why the deals exist
The commercial pressure to license comes from the open web closing. A longitudinal audit of the consent protocols behind the common training corpora found that, counting Terms of Service crawling restrictions rather than robots.txt alone, a full 45% of the C4 corpus is now restricted[15]. A crawler reading only the file understates what the site's own terms forbid, and a contract is what closes that gap.
Provenance in the existing corpora is weak in the other direction. An audit tracing more than 1,800 text datasets found license omission above 70% and license error rates above 50% on widely used dataset hosting sites[16], so a buyer cannot establish from the public record alone what an assembled corpus was ever permitted to contain. See AI crawlers for how the restrictions are expressed.

Does a licence buy citation
A contract is not a ranking signal. Google's guidance for publishers says that "to be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search"[4], which puts crawling and indexing ahead of any commercial arrangement. The Tow Center tested 200 quotes from 20 publications in ChatGPT and found it "returned partially or entirely incorrect responses on a hundred and fifty-three occasions"[17]. Licensed publishers were not spared: the New York Post and The Atlantic "have licensing deals with OpenAI and enabled access to all the crawlers, but their content was frequently cited inaccurately or misrepresented"[5]. A 2026 audit of citations for 44 Web3 enterprises found exposure skewed toward "already prominent voices"[18].
Access sold as plumbing
Some licensing sells delivery rather than permission. Wikimedia Enterprise sells APIs, snapshots and real-time streams of Wikipedia content, described as "Built for AI, Search, and Knowledge Graphs"[19]. The content is already free to reuse under a Creative Commons licence, so what buyers pay for is the pipe and the service level, not the right.
Licence or litigate
The same publishers appear on both lists. Around six cases have been brought against Perplexity, and a federal judge largely rejected a motion to dismiss Reddit's case, saying Reddit had plausibly pleaded that Perplexity "conspired with at least one of the three data scrapers to bypass access controls and obtain Reddit's material"[20].
Data
Show the numbers (3)
| Point | $m |
|---|---|
| 2023 | 15.2 |
| 2024 | 114.7 |
| 2025 | 140 |
References (20)
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI features and your website - Google Search Central Archive
- How ChatGPT Search (mis)represents publisher content - Columbia Journalism Review Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- AI deals with publishers: full list - Press Gazette Archive
- Reddit, Inc. Form S-1 registration statement Archive
- Reddit, Inc. Form 10-K, disaggregation of revenue by source Archive
- Reddit, Inc. Form 10-K, revenue note Archive
- Consent in Crisis: The Rapid Decline of the AI Data Commons
- The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing and Attribution in AI
- How ChatGPT Search (mis)represents publisher content - Columbia Journalism Review Archive
- When Attention Becomes Exposure in Generative Search Archive
- Wikimedia Enterprise Archive
- AI deals with publishers: full list - Press Gazette Archive
Last updated 2026-09-04. Written and maintained by Baseline Labs.