# Publishers versus AI scrapers: the fight over content

[Skip to content](#lm-inhoud)Network/[NL](/en/uitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content)EN[Hubhub.llmnet.nlCompare models on task, language, cost and license.](https://hub.llmnet.nl/en/)[Communitycommunity.llmnet.nlPrompt techniques, patterns and system prompts.](https://community.llmnet.nl/en/)[APIapi.llmnet.nlLLMs in production: rate limits, routing, structured output.](https://api.llmnet.nl/en/)[Consultancyconsultancy.llmnet.nlRolling out AI in an organization, pilot to production.](https://consultancy.llmnet.nl/en/)[Newsnieuws.llmnet.nlAI developments, explained for the Netherlands.](https://nieuws.llmnet.nl/en/)[Benchmarkbenchmark.llmnet.nlMeasure AI quality yourself, on your own tasks.](https://benchmark.llmnet.nl/en/)[Careersvacatures.llmnet.nlAI roles, salaries and career paths in the Netherlands.](https://vacatures.llmnet.nl/en/)[Learnleren.llmnet.nlAI concepts in plain language, beginner to builder.](https://leren.llmnet.nl/en/)[Guidegids.llmnet.nlRun AI privately on your own Mac, PC, NAS or home server.](https://gids.llmnet.nl/en/)[Directorydirectory.llmnet.nlMapping the AI ecosystem: tools, models, companies.](https://directory.llmnet.nl/en/)[Radarradar.llmnet.nlSignals from X, research and communities for indie developers.](https://radar.llmnet.nl/en/)[Appsapps.llmnet.nlReviews of AI apps and open-source repos, with tips for builders.](https://apps.llmnet.nl/en/)[llmnet.nl — main site](https://llmnet.nl/en/)[](https://x.com/intent/post?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content&text=Publishers%20versus%20AI%20scrapers%3A%20the%20fight%20over%20content)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content)[](https://www.reddit.com/submit?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content&title=Publishers%20versus%20AI%20scrapers%3A%20the%20fight%20over%20content)[](#)[](https://x.com/intent/post?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content&text=Publishers%20versus%20AI%20scrapers%3A%20the%20fight%20over%20content)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content)[](https://www.reddit.com/submit?url=https%3A%2F%2Fnieuws.llmnet.nl%2Fen%2Fuitgevers-versus-ai-scrapers-de-strijd-om-de-nederlandse-content&title=Publishers%20versus%20AI%20scrapers%3A%20the%20fight%20over%20content)[](#)

# Publishers versus AI Scrapers: The Battle for Dutch Content

By Ivo Donker — compiled with AI assistance (Claude & Gemini) · 19 August 2026

The tension between creators of high-quality editorial content and developers of large language models has reached a critical phase in the Dutch-language region. Where AI labs were able to index the public web virtually unhindered for years to feed language models, news organizations, publishing houses, and individual authors are erecting increasingly strict barriers. The stakes of this confrontation are twofold: on one hand, compensation for the intellectual property that serves as training material; on the other, preserving a functioning digital distribution model in a network environment where direct answers are eroding traditional click-through traffic.

The dossier on [copyright and AI models](https://nieuws.llmnet.nl/en/ai-en-auteursrecht) the legal foundations of data mining and intellectual property took center stage; in this article we look at the practical and technical dynamics playing out between automated scrapers and the administrators of Dutch websites. Publishers and editorial teams are finding themselves forced to fundamentally overhaul their content architecture to counter unauthorized extraction.

## The Anatomy of Automated Data Extraction

To understand where the friction arises, we need to distinguish between the different ways AI developers collect web data. Not every crawler has the same purpose or the same technical characteristics. In practice, we see roughly three categories of crawlers active on Dutch servers:

 
- Classic search engine crawlers: Bots that index web pages with the primary goal of displaying search results, including a click-through link to the source page. For decades there was a tacit symbiosis here: the publisher supplied content, the search engine supplied visitors.
 
- Offline training scrapers: Bots that harvest web pages en masse to compile large datasets for the pre-training or fine-tuning of future foundation models. This activity generates zero referral traffic for the website owner.
 
- Real-time synthesis and retrieval bots: Crawlers that, the moment an end user asks a question, fetch live web pages, summarize them, and present them as a synthesis. The user is handed the answer directly, so the need to visit the original publication virtually disappears.

This shift breaks the historical balance of the open web. Where indexing used to result in measurable traffic and advertising revenue or subscriber acquisition, the publisher increasingly functions in today's landscape merely as a free raw-material supplier for systems that then compete with that same publisher.

## The Legal Reality: TDM Exceptions and Opt-Outs

At the European level, the discussion is governed by the Directive on Copyright in the Digital Single Market (DSM Directive, Directive (EU) 2019/790), which has been implemented in the Dutch Copyright Act. Article 4 of this directive contains an exception for text and data mining (TDM) for commercial purposes, provided rightsholders have not made an explicit reservation.

Under the law, this reservation must be formulated in an 'appropriate manner,' and for content that is publicly accessible online, this must be done in a machine-readable way. This leads to a legal and technical cat-and-mouse game. When a publisher states in its terms and conditions that scraping is prohibited, this is not always legally sufficient if no machine-readable instruction is present. At the same time, legal experts are asking whether advanced deep-learning training is even fully covered by the original definition of text and data mining, which was once designed for academic pattern research in large text files.

For a broader look at how artists, authors, and designers within the Dutch-language region are affected by these frameworks, the analysis on [the impact of AI on the creative sector](https://nieuws.llmnet.nl/en/ai-en-creatieve-sector) offers deeper insight into collective rights and enforcement. For news publishers, the burden of proof is particularly heavy: once a model has been trained, it is extremely difficult to conclusively demonstrate that specific articles were used without permission, unless the model reproduces literal fragments.

 
 
 Crawler / Bot Type | 
 Owner / Initiative | 
 Primary function | 
 Machine-Readable Opt-Out Status | 
 

 
 
 
 GPTBot | 
 OpenAI | 
 Training future AI models | 
 Respected via robots.txt | 
 

 
 OAI-SearchBot | 
 OpenAI | 
 Real-time search results & links | 
 Respected via robots.txt | 
 

 
 ClaudeBot | 
 Anthropic | 
 Model training & data collection | 
 Respected via robots.txt | 
 

 
 PerplexityBot | 
 Perplexity AI | 
 Live retrieval for answer synthesis | 
 Mixed compliance reported | 
 

 
 CCBot | 
 Common Crawl | 
 Open web archiving & open datasets | 
 Respected via robots.txt | 
 

 
 Bytespider | 
 ByteDance | 
 LLM training & analysis | 
 Historically frequent noncompliance incidents | 
 

 

## Technical Defense: From robots.txt to Web Application Firewalls

Publishers' first line of defense is the age-old robots.txt-protocol. A significant portion of leading news sites have configured their robots file to explicitly deny access to AI bots. A typical configuration looks like this:

# Voorbeeld van machineleesbare opt-out voor AI-crawlers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

# Zoekmachines voor reguliere vindbaarheid blijven toegestaan
User-agent: Googlebot
Allow: /

Although this protocol is simple to implement, it has fundamental limitations. Robots.txt is a voluntary protocol; it relies on the crawler's willingness to respect the instructions. Parties that operate less transparently can simply change their User-Agent string to that of a regular desktop browser or a generic mobile client.

For this reason, media companies have shifted to active blocking via Web Application Firewalls (WAFs). These systems analyze network behavior, TLS fingerprints, and IP reputation. When a bot attempts to request tens of thousands of pages per hour using headless browsers, the WAF intervenes with JavaScript challenges or rate limiting. This technical arms race brings additional infrastructure costs for publishers, who must secure servers against scraping volumes that regularly peak at levels comparable to light DDoS attacks.

## The Dilemma of Search Engine Dominance and AI Overviews

The situation becomes even more complex for publishers when major search engines blend traditional search functionality with generative overviews. In AI overviews, answers are placed directly above the organic search results. This puts publishers in a strategic bind.

Anyone who blocks a crawler via robots.txt to prevent articles from being summarized in AI answers risks disappearing entirely from the regular search index. For commercial media, this can lead to a significant drop in organic incoming visitor volume. Search engines do offer more granular tags (such as nosnippet or max-snippet), but these do not solve the core problem: a platform can rarely choose selectively to remain findable via classic blue links while not contributing input to the underlying generative summaries, without losing reach.

How this shift in organic discoverability and user behavior measurably plays out for online platforms is something we analyze in the guide on [the dynamics of SEO and AI content](https://radar.llmnet.nl/en/seo-ai-content-juli-2026), which breaks down the structural pressure on organic click volumes. The traditional content funnel — in which informative public articles serve as a gateway for paying subscribers — is thereby thrown off balance.

## The Impact of Browser Assistants and Zero-Click Searches

Beyond search engines on the web, integrated AI assistants in browsers form a second challenge. Modern browsers can read the content of an open web page directly from memory, summarize it, and rephrase it for the reader, without the page itself still being viewed intensively.

In the background article on [AI assistants in the browser and media impact](https://nieuws.llmnet.nl/en/ai-assistenten-in-de-browser) it is explained how local and server-based software layers bypass advertisements and interpret structured data. This reinforces the phenomenon of 'zero-click consumption': the consumer reads the facts and analysis but does not visit the website long enough to generate ad impressions or be confronted with a subscription offer.

For Dutch news provision, which relies heavily on a mix of digital subscriptions and advertising revenue, this directly undermines the economic foundation of investigative journalism. Producing an in-depth reconstruction requires significant editorial capacity, while an AI assistant summarizes the article within seconds into concise bullet points that are consumed elsewhere.

## International Deals versus the Dutch-Language Market

At the international level, some major media companies are responding by entering into licensing agreements. Large international corporations have signed contracts with AI labs. In exchange for fees, the AI companies gain authorized access to current and historical archives via dedicated API connections.

For the Dutch market, however, this dynamic is different:

 
- Scale limitation: The Dutch-language region has a relatively modest number of speakers. For international tech companies, the commercial necessity of striking separate, costly deals with individual Dutch publishers is much smaller than for English- or Spanish-language sources.
 
- Market concentration: The private Dutch news market has a high degree of concentration among a few large corporations. This gives them negotiating leverage on the one hand, but on the other hand risks leaving smaller, independent titles and regional media out in the cold under broad blanket agreements.
 
- Public media: Public broadcasters operate with government funds. Their content is in principle intended for society, but having publicly funded databases taken over free of charge by commercial tech companies runs into fundamental societal objections.

 
 
 Publisher Strategy | 
 Advantages | 
 Drawbacks & Risks | 
 

 
 
 
 Full blockade (WAF + robots.txt) | 
 Protects intellectual property; saves server load. | 
 Risk of losing online visibility and relevance. | 
 

 
 Entering into direct licensing deals | 
 Direct revenue stream; controlled access via API. | 
 Reinforces dependence on platforms; often only feasible for large players. | 
 

 
 Strict Paywalls & Dynamic Shielding | 
 Forces users and scrapers into authentication. | 
 Reduces reach among occasional readers; scrapers sometimes find alternative routes. | 
 

 
 Collective advocacy | 
 Collective rates via collective management organizations. | 
 Slow processes; cross-border enforcement remains complex. | 
 

 

## The Training Data Paradox: Synthetic Data versus Quality Journalism

A clear paradox is now emerging in the AI sector. As more publishers lock down their content and the public internet becomes saturated with automatically generated text of varying quality, the availability of authentic human training data declines. This forces developers to search for alternative data sources.

Those wanting to know how developers are trying to make up for this shortage of authentic text can turn to the article on [open datasets for language models](https://nieuws.llmnet.nl/en/open-data-initiatieven-voor-ai), which discusses initiatives around public domain and heritage data. However, without a continuous influx of current, fact-checked journalistic articles, language models quickly lose touch with societal current events and changing terminology. Synthetic data can refine patterns but does not shed light on new facts about societal events or policy developments.

This gives quality publishers a substantive position in the longer term: authentic data is becoming scarcer and therefore more relevant. The question remains, however, to what extent editorial teams have the operational room in the short term to wait for structural market compensation for this continuous data supply.

## Future Outlook and Possible Ways Out

The battle for Dutch content is gradually shifting toward regulatory enforcement and new technical distribution standards. Three possible directions of development are emerging:

First, the rise of standardized micro-licenses and automated payment protocols. Instead of binary choices (block everything or make everything freely available), platforms can develop technical interfaces through which AI agents pay a fee per requested article or token. However, this requires broadly supported technical standards that are still very much in development at this time.

Second, a more prominent role for collective management organizations. Just as in the music industry and with reprographic rights, a system could emerge for training and querying AI systems on copyrighted material. The proceeds could then be paid out to affiliated publishers and independent creators via distribution formulas.

Third, further enclosure of the web. Websites will increasingly disappear behind authentication walls, restricting access without an account or bot verification. In this way, the traditional open web is gradually giving way to shielded environments.

The outcome of this process is decisive for the provision of information. If editorial journalism becomes harder to finance because reach and revenue are skimmed off by external aggregation layers, this ultimately also affects the reliability of the sources that AI systems themselves rely on.
