Skip to content
Search - regulation

Google must hand its European click data to rivals

On 16 July the European Commission ordered Google to license the anonymised record of what Europeans search for, which results they saw and which links they opened, to any rival search engine or AI chatbot that qualifies. Google published the operating rules at the start of September: agreements from 17 September, first data samples on 16 November, and sharing itself from January 2027. The Commission's summary of what the data is for is that it helps recipients "understand which websites to prioritise". Your pages are the websites in that sentence, so which engines can already find them?

George, the Baseline Labs mascot, holding two identical jars

Google already ran a voluntary version, and it took the queries out

Google was not starting from nothing. It already ran a voluntary data-sharing scheme, and the Commission's verdict on it is the most damning figure in the story: it "was removing between 90 and 100% of unique search queries from the dataset", with the result that "there has been no meaningful uptake by potential beneficiaries". A record of European search with the searches taken out has no buyers.

So on 27 January 2026 the Commission opened specification proceedings covering the scope of the data, the anonymisation method and "the eligibility of AI chatbot providers to access the data". On 16 July it adopted binding measures under Article 6(11) of the Digital Markets Act, on market-position grounds: "With a market share of more than 90% for decades, Google Search has a strong market position". Specification proceedings carry no fine. They tell a gatekeeper exactly how to comply.

That "more than 90%" is the Commission's characterisation. StatCounter's panel, which measures tracked page views on traditional search engines and counts no chatbot queries at all, puts Google at 89.69% of European search in August 2026, with Bing second on 4.74% and nothing else above 5%. Ireland is tighter still, at 94.01%.

90-100%
Of unique queries stripped from Google's earlier voluntary dataset, per the Commission
89.69%
Google's share of European search, StatCounter, August 2026
94.01%
Google's share of Irish search, StatCounter, August 2026

Queries, clicks, the URLs people opened, and where they ranked

The shared fields are specific, and they are exactly the ones a search engine needs. Alphabet has to hand over "queries entered by end users in Google Search on any access point, query metadata (for example, language and device type), URLs viewed by users, users' actions interacting with search engine results, and information on where a search result is positioned (i.e. ranked) in the search results page", across free and paid search alike. Google calls the licensable Search Dataset "ranking, query, click, and view data from Google Search in the European Economic Area". What is held back is equally specific.

In the datasetKept out of it
Queries typed into Google Search in the EEA, from any access pointUser account information and search histories
Query metadata, for example language and device typeThe precise timestamp of a search
The URLs users viewedVery long queries and queries containing rare words
The actions users took on the resultsLocation finer than a NUTS 3 region, dropping to country level where the anonymity threshold is not met
Where each result was ranked on the pagePrecise interaction durations, replaced by time intervals
Free and paid search resultsThe URLs of paid results, which are removed

Three further limits shape how useful it is. Data arrives with a latency of at least seven days, so nobody gets a live feed. Each beneficiary's access runs for five years at most. And the measures "do not require the sharing of Google's algorithms or technology". Rivals get the evidence of what worked. The machine that decided it stays inside Google.

Anonymisation is k-anonymity. Each user sits "in a group with at least 1,000 users with the same location, device type and query language", and the Commission says 95% of users will land in larger groups of at least 25,000.

Rivals may ground on this data. They may not train on it.

This is the part that gets summarised wrongly, and it changes what you should expect. The Commission names the permitted uses plainly: improve indexing, improve ranking, and for chatbots improve grounding, the retrieval step where "AI chatbots rely on search retrieval systems to fetch most recent information from the web and ensure accuracy". The annex to the decision adds "web crawling and index building" to that list. Then it draws the line in one sentence: "Training of the general purpose LLM underlying the AI chatbot is excluded from the permitted purpose."

Three uses are banned outright. Recipients "cannot use the data to train general-purpose AI models, improve services that are unrelated to online search engine services like consumer profiling and advertising, or use the data to systematically replicate Alphabet's search results". So the accurate sentence about your website is narrow. The clicks Europeans give Google become part of how rival engines decide what to crawl, what to index, and what to retrieve when someone asks them a question. They do not become model weights.

Enforcement runs through auditors. Before touching the 5% sample or the full dataset an applicant must "provide a Level 1 reasonable assurance report", and to keep access it must "submit to regular, continuous monitoring by an independent assurance practitioner" under a Level 2 report. That means an audit before any data arrives, another within six months of starting processing, and one every year after.

The privacy objection is not only Google's. Google's programme page concedes that "the data that will be provided will still contain personal data". Brave, one of the rivals the remedy is meant to help, opposes the scheme anyway: it "is not, in our view, provably anonymous, and its release would impose severe privacy risk on hundreds of millions of European citizens". Google's Kent Walker argued on the day the decision landed that "Europeans' private searches would be exposed to unfamiliar companies". A would-be beneficiary and the gatekeeper landing on the same objection is a sign that the anonymisation question is still open.

Three dates, and who gets through the gate

Google's programme page carries the operating detail and was last updated on 4 September 2026. Agreements go to relevant applicants "starting on September 17, 2026", and the data samples "will be available as of November 16, 2026". January 2027 is the Commission's milestone: by then Alphabet "must finalise the pricing offer and communicate this to the Commission and third-party online search engines", and Google "must start sharing search data with eligible search engine providers from January 2027".

There are three samples. Sample A is 1,000 rows, free. Sample B is synthetic, up to 10 million queries, paid for. Sample C is a 5% cut of the real Search Dataset, paid for, and the first one behind the Level 1 audit: "The Level 1 audit gates the 5% sample and the full dataset, not the two smaller samples." An applicant can inspect the free rows and the synthetic set before paying for an assurance report.

Eligibility is two conditions, both on Google's page. An applicant must "have had at least 50,000 monthly average users of its OSE services in the EU in the past year", where OSE means an online search engine, and must "either have provided OSE services in the EU for at least the last two consecutive years... or have been founded less than two years before applying but received more than EUR 50 million in capital investments". The Commission states outright that "AI chatbots offering search functionalities are eligible to receive shared data". Fees are capped at "the incremental costs of making the data available, together with a specified rate of return". No price exists yet.

Nobody has publicly applied. No company has confirmed applying, being found eligible, or receiving a single row, and the first agreements have not gone out. Google disagreed with the decision in public on the day it was adopted and has not confirmed an appeal, though a specification decision can be challenged before the General Court. Read the dates as the plan of record rather than a certainty.

In Ireland, 94% of the clicks belong to one engine

Google took 94.01% of Irish search in August 2026, against 3.51% for Bing, 0.93% for DuckDuckGo and 0.34% for Ecosia. Nearly every search-driven click an Irish customer made in August happened on Google. Those clicks, anonymised and stripped as described above, are the thing now being licensed.

The Commission spells out the consequence for site owners: "Improve indexing: search engines rely on an index of websites that can be searched efficiently to respond to a search query. The dataset can help recipients understand which websites to prioritise." The dataset carries the URLs users viewed and the position those results held. A recipient can see which pages European users actually opened, and prioritise crawling them.

Where that leads next is an argument rather than a finding, so treat it as one. Nobody has published a measurement of how much a rival engine improves once it holds this data, or a fixed-prompt test of how well today's answer engines handle Irish questions. What the rivals say about their own gap is suggestive. Ecosia's managing director puts it as coverage: "We need to go from answering two-thirds of queries to 100%". DuckDuckGo's lead public policy manager puts it as scale: "We are building our own index step by step, and we face a huge barrier in catching up, given Google's scale". Henna Virkkunen named the Commission's hope for the remedy: "Thanks to these measures we hope to see emerging alternatives to Google Search and Google's AI services, such as Gemini".

The other half of this story landed on 31 August, when the Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, covered in the EU just called ChatGPT a search engine. One regulator now treats an AI assistant as a search engine in one file and as a candidate recipient of Google's search data in another. If your reporting still rests on Google positions alone, the gap described in you can rank #1 on Google and be invisible to AI is the one to watch, because the engines it describes are the ones queueing for this data.

What you actually do

1. Stop reporting Google share as your visibility. The 94.01% figure describes where Irish clicks happen on traditional search engines, and counts no chatbot queries at all. It tells you nothing about whether an answer engine can find your pages.

2. Take the cross-engine measurement now. Pick the questions your customers really ask, run them across several engines on a fixed schedule, and save the result. 16 November is the first date anything changes hands, and a comparison needs a picture from before that.

3. Look after the pages Europeans already click. The dataset carries URLs viewed and ranked position, so the pages earning European clicks on Google today give a recipient's crawler its strongest reason to fetch them tomorrow. Keep them crawlable and quick, and answer the question in the first paragraph.

4. Widen the list of engines you check. An AI chatbot with search functionality can qualify for this data, so the same European click behaviour will soon inform several engines instead of one. Measuring Google alone describes a shrinking share of how customers find you.

Measure your visibility across engines

Questions people ask about this

Does this mean rival AI chatbots get to train on Google's search data?
No. The decision permits refining retrieval and ranking, web crawling, index building, and grounding for AI chatbots, and states that training the general purpose LLM underlying a chatbot is excluded from the permitted purpose. Recipients are also barred from consumer profiling, advertising, and systematically replicating Google's results.
Which companies are getting the data?
None that anyone can name. As of 8 September 2026 no company has publicly confirmed applying, being found eligible, or receiving data, and the first licence agreements only start dispatching on 17 September. The published test is 50,000 monthly average EU users in the past year, plus either two consecutive years of EU trading or, for a company under two years old, more than EUR 50 million raised.
Does any of this apply outside Europe?
No. The Search Dataset is Google Search data generated in the European Economic Area, and it is licensed for improving search services aimed at EEA users. It does nothing directly for a rival engine's coverage of the US or the rest of the world, and it does not touch Google Search outside Europe.
Is my own website's data in there?
Your pages appear in it as URLs European users viewed from search results, with the position those results held. There is no account information and no search histories, the precise timestamp is stripped, location is generalised to a NUTS 3 region, and the URLs of paid results are removed. There is no opt-in or opt-out for a site owner.
Is the shared data anonymous?
Contested, and by both sides. Google's programme page says the data will still contain personal data, and Brave, one of the rivals the remedy is meant to help, argues the release is not provably anonymous. The Commission's answer is k-anonymity with a floor of 1,000 users per group, 95% of users in groups of at least 25,000, and mandatory independent audits of every recipient.
Could the timeline slip?
It could. Google opposed the decision publicly on the day it was adopted, and a specification decision can be challenged before the General Court, though Google has not confirmed an appeal. The published dates are the plan of record, not a guarantee. Build your measurement around them rather than betting a campaign on them.

Sources: Google Search Central's European Search Dataset Licensing Program page, the European Commission's DMA developer-portal Q&A on the Alphabet specification proceedings, press release IP/26/1634, its Shaping Europe's Digital Future announcement and the annex to the decision in case DMA.100209, StatCounter Global Stats, Search Engine Journal, TechRadar and the Google Keyword blog. Market-share figures are StatCounter's August 2026 panel snapshot and move month to month.

References (11)
  1. "understand which websites to prioritise" (digital-markets-act.ec.europa.eu)
  2. 27 January 2026 (digital-markets-act.ec.europa.eu)
  3. 89.69% of European search in August 2026 (gs.statcounter.com)
  4. 94.01% (gs.statcounter.com)
  5. "ranking, query, click, and view data from Google Search in the European Economic Area" (developers.google.com)
  6. "web crawling and index building" (ec.europa.eu)
  7. "is not, in our view, provably anonymous, and its release would impose severe privacy risk on hundreds of millions of European citizens" (techradar.com)
  8. "Europeans' private searches would be exposed to unfamiliar companies" (blog.google)
  9. "must start sharing search data with eligible search engine providers from January 2027" (ec.europa.eu)
  10. "The Level 1 audit gates the 5% sample and the full dataset, not the two smaller samples." (searchenginejournal.com)
  11. "Thanks to these measures we hope to see emerging alternatives to Google Search and Google's AI services, such as Gemini" (digital-strategy.ec.europa.eu)

See how AI search sees your site

Free AI visibility check in under a minute

Check my site
George
Online
0%