AI Scrapers are automated software agents (web crawlers, retrieval bots, and content-fetching pipelines) operated by foundation-model laboratories, AI search vendors, retrieval-augmented-generation services, dataset aggregators, and synthetic-media producers for the purpose of harvesting publicly…

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:UserAgentString))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:IPRangeManifest))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:CrawlFrontier))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:PolitenessPolicy))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:RobotsTxtParser))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:ContentExtractionPipeline))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:hasPart ai:DeduplicationFilter))

## Dependency Relationships
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:requires ai:NetworkConnectivity))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:requires ai:HTTPClient))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:requires ai:PublicWebContent))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:requires ai:ComputeInfrastructure))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:requires ai:IdentityProvenance))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:dependsOn ai:CommonCrawl))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:dependsOn ai:InternetInfrastructure))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:dependsOn ai:DNS))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:dependsOn ai:PublicWebArchitecture))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:dependsOn ai:CDNLayer))

## Capability Relationships
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:FoundationModelPreTraining))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:RetrievalAugmentedGeneration))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:SearchIndexConstruction))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:KnowledgeGraphExtraction))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:SyntheticDataCuration))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:enables ai:FineTuningDatasetAssembly))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:supports ai:LargeLanguageModelTraining))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:supports ai:MultimodalModelTraining))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:supports ai:GenerativeAI))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:supports ai:AISearch))

## Implementation Relationships
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:implements ai:RobotsExclusionProtocol))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:implements ai:HTTPProtocol))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:implements ai:PolitenessThrottling))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:implements ai:DeDuplicationHashing))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:implements ai:ContentTypeNegotiation))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:uses ai:HeadlessBrowser))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:uses ai:HTTPLibrary))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:uses ai:SitemapsXML))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:uses ai:BloomFilter))

## Reduction Relationships
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:reduces ai:TrainingDataAcquisitionCost))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:reduces ai:ManualDataCuration))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:reduces ai:LicensingNegotiationOverhead))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:reduces ai:DatasetFreshnessLatency))

## Association Relationships
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:relatedTo ai:BotManagement))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:relatedTo ai:CopyrightLaw))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:relatedTo ai:DataProtection))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:relatedTo ai:TrainingData))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:contrastsWith ai:ClassicSearchCrawler))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:contrastsWith ai:WebArchiveCrawler))
SubClassOf(ai:AIScrapers
  ObjectSomeValuesFrom(ai:contrastsWith ai:AcademicCrawler))

## Data Properties (Characteristics)
DataPropertyAssertion(ai:hasIdentifier ai:AIScrapers "AI-1043"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:AIScrapers "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:botTrafficShare2024 ai:AIScrapers "0.496"^^xsd:decimal)
DataPropertyAssertion(ai:badBotShare2024 ai:AIScrapers "0.32"^^xsd:decimal)
DataPropertyAssertion(ai:botManagementMarketUSD2025 ai:AIScrapers "1400000000"^^xsd:integer)
DataPropertyAssertion(ai:trainingDataLicensingMarketUSD2025 ai:AIScrapers "1200000000"^^xsd:integer)
DataPropertyAssertion(ai:cumulativeClearviewAIFinesEUR ai:AIScrapers "100000000"^^xsd:integer)
DataPropertyAssertion(ai:rfc9309Year ai:AIScrapers "2022"^^xsd:integer)
DataPropertyAssertion(ai:euAIActArticle53Year ai:AIScrapers "2025"^^xsd:integer)

## Property Constraints
SubClassOf(ai:AIScrapers
  DataMinCardinality(1 ai:hasUserAgent xsd:string))
SubClassOf(ai:AIScrapers
  DataAllValuesFrom(ai:respectsRobotsTxt xsd:boolean))
SubClassOf(ai:AIScrapers
  DataSomeValuesFrom(ai:declaredOperator xsd:string))

## Annotations
AnnotationAssertion(rdfs:label ai:AIScrapers "AI Scrapers"@en)
AnnotationAssertion(rdfs:comment ai:AIScrapers "Automated web agents operated by foundation-model labs, AI search vendors and retrieval services to harvest publicly accessible content at scale for pre-training corpora, fine-tuning data and retrieval indexes. Identified by named user-agents (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Amazonbot, Applebot-Extended, PerplexityBot, Diffbot, FacebookBot), governed by voluntary robots.txt under IETF RFC 9309 (2022) and supplementary 2024-2026 proposals (AIPREF Working Group, IAB Tech Lab LLM Content Ingest API, Cloudflare pay-per-crawl, Spawning ai.txt, C2PA, NLWeb). Policed by enterprise bot-management vendors (Cloudflare, Akamai, Imperva, DataDome, Kasada, HUMAN). Confronted by major copyright litigation (NYT v OpenAI 2023, Getty v Stability UK 2025, Authors Guild MDL, RIAA v Suno/Udio 2024, News Corp v Perplexity 2024) and GDPR enforcement (Clearview AI cumulative €100M+ fines). Subject to EU AI Act Article 53 GPAI training-data transparency from 2 August 2025 and Code of Practice signed July 2025. Contested by defensive tooling (Glaze, Nightshade, Have I Been Trained, ai.robots.txt, Kudurru) and commercial licensing settlements (Reddit/Google $60M, Stack Overflow/OpenAI, NYT/Apple, AP/OpenAI, FT/Anthropic) constituting a $1.2B 2025 training-data market growing to $5.4B 2030."@en)
AnnotationAssertion(dcterms:identifier ai:AIScrapers "AI-1043"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:AIScrapers "Web Crawling, Training Data, Copyright, Bot Management, AI Governance, Data Protection"@en)

## Property Characteristics
AsymmetricObjectProperty(ai:requires)
AsymmetricObjectProperty(ai:enables)
AsymmetricObjectProperty(ai:implements)
AsymmetricObjectProperty(ai:contrastsWith)
TransitiveObjectProperty(ai:dependsOn)
FunctionalDataProperty(ai:rfc9309Year)
FunctionalDataProperty(ai:botTrafficShare2024)

About AI Scrapers

  • AI Scrapers are the automated agents through which the contemporary generative-AI industry acquires the substrate of its products. They sit at the intersection of three previously distinct disciplines—web crawling (an engineering tradition dating to ALIWEB 1993, WebCrawler 1994 and Googlebot 1998), data engineering (the ETL pipelines feeding pre-training warehouses), and intellectual-property practice (publishing, music licensing, photographic stock libraries)—and have, by becoming the primary mechanism by which copyrighted material enters foundation-model weights, made the previously theoretical question of whether machine learning constitutes copying a live commercial and legal dispute.
  • The defining distinction between an AI scraper and a classic search-engine crawler is not technical but teleological: both fetch HTML; both respect (or pretend to respect) robots.txt; both build large indexes. But where Googlebot indexes to direct user traffic back to the source page, an AI scraper ingests to populate model weights that subsequently answer the user’s question without ever sending traffic back. This disintermediation of the publisher is what has driven publisher pushback from the 2023 introduction of GPTBot through the November 2025 Getty v Stability AI judgment.
  • The population of named AI crawlers has grown from one (CCBot, operating quietly since 2008 to feed Common Crawl) to several dozen as of 2026, with the Dark Visitors registry (Dark Visitors, 2024) tracking 200+ AI-relevant user-agents. The crawlers visible to publishers are only the obedient ones; the invisible crawlers are the ones that strip user-agents, rotate residential proxies, and present as ordinary browser traffic—the practice exposed by Wired and Robb Knight in June 2024 with respect to PerplexityBot, and the engineering reality of every commercial scraping-as-a-service vendor (Apify, Bright Data, ScraperAPI, Oxylabs) whose products are marketed for legitimate competitive intelligence but consumed indistinguishably for AI training.

Components and Architecture of an AI Scraper

An industrial-scale AI scraper is not a single program but a distributed pipeline comprising several discrete subsystems, each with its own engineering and legal characteristics.

Crawl Frontier and Seed Selection

The crawl frontier is the prioritised queue of URLs awaiting fetch. Seed URLs typically come from three sources: existing indexes (Common Crawl, Bing webmap, sitemaps), domain-rank lists (Tranco, Majestic Million, Cisco Umbrella), and social-graph crawls (Reddit submission listings, Hacker News, Twitter/X firehose). Modern AI scrapers prioritise frontiers by expected training-value-per-byte: high-quality long-form English text from reputable domains is weighted above thin-content SEO farms. ByteDance’s Bytespider notoriously over-weighted Q&A sites (Stack Exchange family, Quora, Reddit) leading to the disproportionate request-rate observed by 2024 publisher dashboards.

HTTP Client and Politeness Policy

Production crawlers run concurrent HTTP/2 and HTTP/3 clients with per-host connection pooling. Politeness is enforced through token-bucket rate limiters keyed by origin domain, typically 1-10 requests per second per domain for established crawlers (GPTBot, ClaudeBot publish their rate limits), with adaptive backoff on 429 Too Many Requests and 503 Service Unavailable responses. Non-compliant scrapers (Bytespider pre-2024, allegedly PerplexityBot per the 2024 investigations) operate without effective per-host throttling, producing the kinds of accidental denial-of-service that took down iFixit and Freelancer.com in summer 2024.

robots.txt Parser

Per RFC 9309, the parser must implement: case-sensitive User-agent matching with longest-prefix-wins semantics; UTF-8 path matching with $ end-of-path and * wildcard support; case-sensitive path comparison; and Crawl-delay interpretation (though Crawl-delay is non-standard and inconsistently honoured). Sophisticated AI scrapers also parse non-standard extensions including noai and noimageai from Spawning’s ai.txt proposal and the X-Robots-Tag HTTP response header which permits per-resource opt-out without robots.txt modification.

Content Extraction and Boilerplate Removal

Raw HTML is processed through boilerplate removal (Readability, Trafilatura, Resiliparse) to isolate main-text content from navigation, comments, ads and footers. Common Crawl’s WARC dumps include both raw HTML and extracted text; training pipelines typically apply additional filtering (FastText language classification, perplexity filtering against a reference LM, deduplication via MinHash). The extraction stage is also where licence/CMI metadata is most commonly stripped, which is the basis for the NYT lawsuit’s DMCA Section 1202 claim.

Deduplication

Foundation-model training corpora are aggressively deduplicated to avoid memorisation. MinHash with locality-sensitive hashing (LSH) at 5-gram or document-level granularity removes near-duplicates; the SemDeDup technique (Abbas et al. 2023) extends this to embedding-space deduplication removing semantic duplicates that lexical hashing misses. Deduplication is also a defensive measure against the Glaze/Nightshade-style poisoning attack: poisoned samples diluted in massive deduplicated corpora have reduced per-sample impact.

Storage and Versioning

Output is stored in WARC (Web ARChive) format (ISO 28500), typically compressed with zstd and chunked at the 1 GiB level for parallel S3 read. Versioning matters legally: the NYT v OpenAI discovery dispute hinged on whether OpenAI retained training-corpus snapshots at temporal granularity sufficient to reconstruct what was ingested when. Judge Sidney Stein’s May 2024 preservation order represented the first judicial intervention into AI-training corpus retention.

Identity and Provenance Layer

Increasingly important: bot-attestation cryptographic identity is moving from proposal to deployment. Cloudflare’s verified-bot programme issues signed manifests binding User-agent to operator-controlled IP ranges and public keys; the IETF’s HTTP Message Signatures (RFC 9421, 2024) provides the underlying signature mechanism. By 2027 it is expected that unsigned crawler traffic will be presumptively low-trust at major CDN edge.

The Named AI Crawler Census (2026)

The publicly documented AI training and retrieval crawlers, as of Q1 2026, can be grouped by operator and intent.

OpenAI

  • GPTBot (announced August 2023): training crawler for successor models to GPT-4. User-agent contains GPTBot/1.x; +https://openai.com/gptbot. Published IP ranges in JSON manifest. Honoured by publishers via User-agent: GPTBot disallow directives. NYT, BBC, CNN, Reuters and 90% of US top-100 newspapers blocked within 48 hours of the August 2023 announcement.

  • OAI-SearchBot (2024): retrieval crawler for ChatGPT search (the product that became SearchGPT and eventually merged into ChatGPT). Distinct from GPTBot. Not used for training—used to render result snippets in real time. Opting out blocks ChatGPT search results but does not affect existing model weights.

  • ChatGPT-User: on-demand fetch agent triggered by user action inside ChatGPT (browse-the-web tool). Each request is keyed to a user prompt; not a bulk pre-training pipeline.

    This three-tier architecture (training crawler, search crawler, user fetcher) became the template that Anthropic, Google and others subsequently adopted.

    Anthropic

  • ClaudeBot: training crawler. User-agent: ClaudeBot/1.0 (+claudebot@anthropic.com). Caused notable load on iFixit and Freelancer.com in summer 2024 — iFixit CEO Kyle Wiens publicly tweeted hourly request counts in the tens of millions, prompting Anthropic to publish improved guidance and a respect-retroactively undertaking honouring later opt-outs by removing previously scraped pages from training corpora.

  • Claude-User: on-demand fetch agent (parallel to ChatGPT-User).

  • Claude-SearchBot: retrieval crawler for Claude’s web-search feature.

  • anthropic-ai: legacy UA token now retired in favour of ClaudeBot; many publisher robots.txt files retain it for historical opt-out continuity.

    Google

  • Google-Extended (October 2023): a pseudo-user-agent — no crawler actually identifies itself with this string. It is a purely declarative opt-out signal: disallowing Google-Extended in robots.txt blocks the use of Googlebot-crawled content in Gemini, Vertex AI and Bard training pipelines, without affecting classic Googlebot ranking. This separation of crawling from training-use is uniquely Google’s; OpenAI/Anthropic instead use distinct crawlers per purpose.

  • Googlebot and other Google product crawlers (AdsBot, Image, etc.) remain governed by the older robots.txt conventions and are out-of-scope as AI scrapers per se, but their output feeds Vertex AI by default unless Google-Extended: Disallow is set.

    Common Crawl

  • CCBot: operated since 2008 by the non-profit Common Crawl Foundation. Produces monthly 250 TiB WARC dumps published as CC-BY licensed datasets on AWS Open Data. The single most consequential AI scraper in existence: the pre-training corpora of GPT-3, RedPajama, Falcon, MPT, LLaMA-1, Mistral, and dozens of academic models derive directly from Common Crawl. Its 2008-era robots.txt-respecting design predates the AI-training use case; opt-outs are honoured but cannot be retroactively removed from existing dumps.

    Meta

  • meta-externalagent/1.1: training crawler for Llama models. Introduced 2024 to provide a distinct opt-out path from the legacy facebookexternalhit link-preview fetcher (which dates to the early 2010s open-graph era).

  • FacebookBot: search-related; partial overlap with training pipelines.

    ByteDance

  • Bytespider: training crawler feeding Doubao (China-market LLM). Aggressive crawling pattern with frequent robots.txt non-compliance prior to mid-2024, triggering mass publisher blocking via Cloudflare and Akamai fingerprint rules rather than UA-based directives. The most-blocked AI crawler in the Cloudflare AI Audit dashboard (2024-2025).

    Amazon

  • Amazonbot: feeds Alexa+ and forthcoming Olympus/AGI-tier models. Published JSON IP-range manifest. Respects robots.txt. Distinguished from amazon-product-search-bot and product-listing crawlers.

    Apple

  • Applebot-Extended (introduced June 2024): opt-out signal — like Google-Extended — gating the use of Applebot-crawled content in Apple Intelligence and Siri LLM training. Apple’s late introduction of this signal (June 2024, eight months after Google-Extended) was driven by publisher pressure ahead of the September 2024 iOS 18 / Apple Intelligence launch.

  • Applebot: classic search crawler underlying Siri and Spotlight; not an AI scraper per se but feeds Apple Intelligence by default unless Applebot-Extended: Disallow is set.

    Diffbot

  • Diffbot: commercial knowledge-graph extraction crawler operating since 2008. Sells extracted entity graphs (Diffbot Knowledge Graph: 10B+ entities, 1T+ facts) to enterprises including Salesforce, NASDAQ, Cisco. Controversial because the output is licensed rather than the raw HTML, blurring the fair-use line: courts have not yet ruled on whether extraction-and-relicensing of public web data constitutes derivative use.

    Perplexity

  • PerplexityBot: retrieval crawler. Subject of the June 2024 Wired and Robb Knight investigations that demonstrated UA stripping, residential-IP rotation and direct fetching of pages explicitly disallowed in robots.txt. The exposé coined the phrase ‘AI scraper laundering’ and triggered Cloudflare’s emergency UA-independent block rule, the News Corp/Dow Jones SDNY lawsuit (October 2024) and the New York Post’s complaint.

    Smaller Operators

  • Cohere-AI (Cohere training/retrieval)

  • YouBot (You.com retrieval)

  • DuckAssistBot (DuckDuckGo’s DuckAssist)

  • ImagesiftBot (Hive Moderation’s reverse-image-search and CSAM-detection pipeline)

  • Kangaroo Bot (ByteDance secondary crawler)

  • Timpibot (Timpi distributed search)

  • Awario and Crawlspace (commercial scraping-as-a-service operators consumed indistinguishably for AI training and brand monitoring)

Historical Trajectory: From Search to Generative Training

The lineage from search-engine crawling to AI training is a thirty-year arc punctuated by four distinct eras.

Era 1 — Discovery Crawling (1993-1998): ALIWEB (Martijn Koster, 1993), the World Wide Web Wanderer (Matthew Gray, MIT, 1993), WebCrawler (Brian Pinkerton, U Washington, April 1994), Lycos (CMU, 1994), Excite (Stanford, 1994), AltaVista (Digital Equipment, 1995), Inktomi (UC Berkeley, 1996) and finally Googlebot (Brin and Page, 1998). Crawlers were academic experiments that became commercial products. Volumes were modest — Googlebot circa 1999 processed roughly 25M pages — and the relationship between crawler and publisher was straightforwardly mutualistic: pages got indexed in exchange for traffic returned via search results. robots.txt emerged in 1994 to manage transient overload, not to ration access to a valuable input.

Era 2 — Web 2.0 and the API Decade (2005-2015): As social platforms accumulated user-generated content, the dominant data-acquisition pattern shifted from open-web crawling to authenticated API consumption. Twitter’s Streaming API (2008), Facebook Graph API (2010), and Reddit’s PRAW (2010) let third parties access content at scale under contractual terms. Web scraping persisted but became a grey-market activity — the hiQ Labs v. LinkedIn dispute (filed 2017) is best understood as a late-Era-2 conflict over whether publicly accessible HTML retains contractual restrictions.

Era 3 — Common Crawl and the Quiet Build-Up (2008-2019): Gil Elbaz and Nova Spivack founded the Common Crawl Foundation in 2007; CCBot began producing monthly WARC dumps in 2008. For a decade Common Crawl was a research-and-archival utility, not understood as a commercial input. Its corpus underpinned key NLP datasets (the GloVe word embeddings, Sebastian Nagel’s annual crawl statistics) without controversy. The 2019 announcement that GPT-2 was trained on a Common-Crawl-derived corpus drew minimal attention; GPT-3’s 2020 disclosure that 60 percent of its 570 GiB training set came from a filtered Common Crawl subset (with WebText2, Books1, Books2 and Wikipedia making up the balance) similarly passed without major copyright dispute, in part because the model was not yet commercially launched.

Era 4 — The ChatGPT Moment and the Bot Wars (2022-): ChatGPT’s November 2022 launch made foundation-model training a consumer-facing reality. Within nine months, OpenAI had published GPTBot (August 2023), the New York Times had filed (December 2023), and the bot-management industry had pivoted to AI-scraper detection as its primary growth vector. The defining feature of Era 4 is that the publisher-crawler relationship has become explicitly adversarial in the case of AI training: every fetch is read by the publisher as a potential rights infringement and by the crawler operator as a contested but defensible fair use. The legal and economic resolution of this conflict is still unfolding.

Robots Exclusion Protocol: Origins and Limitations

The Robots Exclusion Protocol — robots.txt — was authored by Martijn Koster in February 1994 after his web server was overrun by early crawlers, principally ALIWEB and the Plymouth Engineering Web Crawler. The original specification is a 4 KB plain-text file at /robots.txt declaring User-agent: ... and Disallow: /path directives. It carried no enforcement mechanism, no cryptographic identity binding, and no provision for per-purpose discrimination. For 28 years it remained a de-facto convention without formal status.

In September 2022 it was finally formalised as IETF RFC 9309 (‘Robots Exclusion Protocol’) by Henner Zeller (Google), Lizzi Harvey (Google) and Gary Illyes (Google), with co-implementer review by Microsoft and Yandex engineers. RFC 9309 codified the parsing rules (longest-match wins on path; case-sensitive; UTF-8) and made the protocol citable in compliance audits — but did not extend its expressive power.

The fundamental limitations of robots.txt for the AI-scraping era are:

  1. Voluntary compliance. Non-honourable bots silently ignore directives. The User-agent header is a self-declared HTTP value trivially spoofable. Enforcement requires fingerprint-based bot management at network ingress (Cloudflare, Akamai), not protocol semantics.
  2. No per-purpose discrimination. A publisher cannot say ‘allow Googlebot for ranking but block Google’s use of my content for training’ in robots.txt syntax alone. Google-Extended is a single-vendor workaround that other vendors must individually replicate (Applebot-Extended); this scales poorly.
  3. Aggressive caching. Crawlers fetch robots.txt at the start of a session and cache for hours-to-days; opt-out latency is 24-72h typical. By contrast, RAG retrieval can pull a freshly-published page within minutes.
  4. No cryptographic identity. Anyone can write a robots.txt rule for any User-agent name; nothing binds a UA string to an operator. Cloudflare’s 2024 Bot Management cryptographic-attestation proposals attempt to address this.
  5. Entry-point only. Many AI crawlers fetch sub-URLs without re-checking the root robots.txt for each, especially when fed URLs from external indexes (Common Crawl, Bing, Reddit listing pages).
  6. No monetary expressiveness. robots.txt cannot say ‘£0.001 per fetch via HTTP 402’; it can only express binary allow/deny.

These limitations have driven a 2024-2026 wave of supplementary protocol proposals:

  • IETF AIPREF Working Group (chartered late 2024, co-chaired by Mark Nottingham and Suresh Krishnan): developing machine-readable preference signals beyond binary allow/deny. Drafts include HTTP response headers (Content-Usage:), URI-list manifests, and per-purpose vocabulary (training, retrieval, summarisation, agent-action).
  • IAB Tech Lab LLM Content Ingest API (draft 2024): structured monetisation channel with HTTP 402 Payment Required and authenticated bot identity.
  • Cloudflare Crawler Hints / pay-per-crawl (proposal 2025): per-fetch micropayments via Cloudflare’s edge with revenue share to origin domain.
  • Spawning.ai ai.txt (2023-): extends robots.txt grammar with model/use-case granularity. Adopted by Stability AI for SD3 training cut.
  • C2PA Content Credentials: cryptographic provenance manifests binding licence terms to media manifest; not a crawler-control protocol but provides upstream signal.
  • Microsoft NLWeb (2025): site-level AI-agent disclosure standard letting publishers describe site structure to AI agents in a machine-readable way, blurring the boundary between ‘opt-out’ and ‘opt-in with API’.
  • TDM Reservations under the EU’s 2019 Digital Single Market Directive Article 4(3): rights holders can reserve TDM in machine-readable form; the format is left to industry, creating the gap that ai.txt and AIPREF are racing to fill.

Bot Management Industry

Enterprise bot management is now a 3.8B by 2030 (Grand View Research), with AI-scraper-related demand the principal growth driver post-2023. The industry sits at network ingress, applying behavioural ML, TLS fingerprinting, JavaScript challenges and IP reputation scoring to classify each request in single-digit milliseconds.

Cloudflare

Cloudflare protects roughly 20 percent of all internet traffic, giving it unmatched visibility into AI-scraper behaviour. Its AI-bot product timeline:

  • July 2024: one-click ‘Block AI Bots’ toggle for all 27M+ Cloudflare-protected domains, including the free tier. Within 30 days, 1M+ domains had activated the block.

  • September 2024: AI Audit dashboard surfacing per-bot request rates per domain, letting publishers see exactly which AI scrapers are hitting them and at what volume.

  • March 2025: AI Labyrinth — actively serving synthetically-generated content to identified non-compliant scrapers, poisoning their training corpora with model-generated text. This was the first major deployment of adversarial content as a network-layer defence.

  • July 2025: default-block of all unknown AI crawlers on the free tier, requiring AI vendors to enrol in Cloudflare’s verified-bot programme to gain access. This shifted the burden of justification from publishers (who previously had to opt out) to AI vendors (who must now opt in).

  • 2025-2026: Cloudflare pay-per-crawl proposal letting publishers charge per fetch with Cloudflare as payment intermediary.

    Other Vendors

  • Akamai Bot Manager: JS challenge + TLS fingerprint + behavioural ML scoring; protects ~15-25 percent of Fortune 500 web traffic.

  • Imperva (Thales subsidiary post-2023): Advanced Bot Protection. The Imperva Bad Bot Report 2024 measured 49.6 percent of all internet traffic as bots, 32 percent classified as bad bots — the highest figure ever recorded — and identified LLM-training scrapers as the fastest-growing bad-bot subcategory.

  • DataDome (Paris, $42M Series C 2023): ML model classifying each request in <2ms; deployed at Foot Locker, Reddit, Tripadvisor.

  • Kasada (Sydney): cryptographic proof-of-work client challenges resistant to headless-browser bots.

  • HUMAN Security (formerly White Ops, acquired PerimeterX 2022): BotGuard / MediaGuard deployed at Disney, Verizon, US DoD.

The unresolved legal question — does training a foundation model on copyrighted material constitute infringement, or is it transformative fair use — has produced an unprecedented litigation wave. As of Q1 2026, 30+ active lawsuits in the US, UK, EU and Canada carry combined claimed damages exceeding $50 billion.

New York Times v. OpenAI/Microsoft (SDNY, 2023-)

Filed in the Southern District of New York on 27 December 2023, the NYT complaint alleges:

  • Training-data infringement: OpenAI ingested NYT archives without licence

  • Output regurgitation: ChatGPT reproduces NYT articles verbatim when prompted

  • DMCA Section 1202 CMI removal: training pipeline strips copyright-management information

  • Unfair competition and trademark dilution

    OpenAI’s response in January 2024 argued fair use, characterised regurgitation as a ‘bug’ and asserted that ‘we support journalism, partner with news organisations, and the New York Times’ lawsuit is without merit.’ Discovery in May 2024 produced a dispute over OpenAI’s deletion of training-corpus snapshots; Judge Sidney Stein ordered preservation. Motion-to-dismiss was largely denied in March 2025, allowing the core infringement and DMCA claims to proceed. Trial is scheduled for 2026. Settlement remains possible but the NYT has publicly rejected pre-trial offers.

    Getty Images v. Stability AI (UK High Court, 2023-2025)

    Case CL-2023-000080 in the Commercial Court. Trial June-July 2025, partial judgment November 2025 by Mrs Justice Joanna Smith:

  • Training-infringement claim failed on territoriality: Stable Diffusion training occurred outside the UK (predominantly on AWS US-East), so UK copyright law’s territorial reach did not extend to the training acts.

  • The court found Stable Diffusion does not store or reproduce Getty’s training images at inference time — the model’s weights are not a ‘copy’ in the copyright sense.

  • Trademark infringement WAS found: Stable Diffusion outputs reproduced Getty’s watermark in some generations, constituting trademark use without authorisation.

  • Database-right claim partially succeeded for UK-resident data fetching.

    The mixed result has been read both ways: AI vendors as a narrowing of UK training-infringement risk; rights holders as a confirmation that watermark/trademark exposure remains live and that the territoriality loophole motivates statutory reform. A US parallel case in Delaware is ongoing.

    Authors Guild v. OpenAI (Consolidated MDL, SDNY 2023-)

    Combines the September 2023 Authors Guild class action with parallel suits by Sarah Silverman, Paul Tremblay (Cara), Michael Chabon, George RR Martin, Jodi Picoult, John Grisham and 17+ authors. Discovery has focused on the Books3 dataset (a 196,640-book corpus sourced from Bibliotik shadow library) used in training Llama-1, Bloom and other open models. Class certification motion pending 2026.

    Music Labels v. Suno / Udio (RIAA, June 2024)

    Filed by Universal Music Group, Sony Music Entertainment and Warner Music Group through the Recording Industry Association of America in June 2024 against Suno (Cambridge MA) and Udio (Uncharted Labs, NYC), the leading text-to-music generative startups. Allegations: training on copyrighted master recordings without licence. Suno admitted in August 2024 court filings that it trained on commercial recordings, invoking fair use as defence. Settlement negotiations reported Q1 2026 with industry-wide impact: the precedent will define whether musical-style mimicry is licensable.

    News Corp / Dow Jones v. Perplexity (SDNY, October 2024)

    Filed by News Corp (Wall Street Journal, New York Post) following the June 2024 Wired exposé. Alleges UA-stripping circumvention of robots.txt, output regurgitation of WSJ articles, and unfair competition. Significant because it targets a retrieval-AI rather than a pre-training laboratory, testing whether RAG citation-with-snippet constitutes fair use.

    Reddit v. Anthropic (June 2025)

    Filed after Reddit’s February 2024 commercial deal with Google did not extend to Anthropic; Reddit alleges unauthorised scraping post-deal. First major case where a platform sues an AI lab for scraping despite having a commercial licensing programme.

    Commercial Licensing Settlements

    Pre-litigation pressure has produced a parallel commercial-deal market:

  • Reddit / Google: $60M per year, February 2024, Gemini training access (disclosed in Reddit’s pre-IPO S-1).

  • Stack Overflow / OpenAI: undisclosed 8-figure value, May 2024 (OverflowAPI commercial licensing). Sparked moderator strike over user-content monetisation.

  • NYT / Apple: licensing deal December 2024 for Apple Intelligence training, despite NYT’s parallel OpenAI lawsuit.

  • Associated Press / OpenAI: July 2023 — the first major news-org training deal.

  • Financial Times / OpenAI: April 2024, multi-year archive licence.

  • Le Monde / OpenAI: March 2024, French archive.

  • News Corp / OpenAI: May 2024, $250M five-year deal across WSJ, NY Post, The Australian, Times of London.

  • Vox Media / OpenAI, The Atlantic / OpenAI: May 2024.

  • Axel Springer / OpenAI: December 2023, Bild and Welt archives.

  • Shutterstock / OpenAI / Meta / Apple / Google: image-licensing pacts.

    Combined, these deals constitute a 5.4B by 2030 (LSEG estimates).

Data Protection Enforcement

Where copyright is the US-and-UK frontline, EU data-protection law has been the European front. The Clearview AI saga illustrates the cumulative reach:

  • Italy Garante: €20M fine, February 2022, for scraping facial images without lawful basis under GDPR Articles 5, 6, 9.

  • CNIL (France): €20M, October 2022, plus order to delete data and €100K/day penalty for non-compliance.

  • Hellenic DPA (Greece): €20M, July 2022.

  • ICO (UK): £7.5M provisional notice October 2021; overturned 2023 on jurisdiction grounds (Clearview’s UK customer base too narrow to trigger UK GDPR extraterritorial reach).

  • Dutch DPA: €30.5M, September 2024, plus €5M per-officer penalties for non-cooperation — a novel deployment of director personal liability.

    Cumulative Clearview AI exposure across EU regulators exceeds €100M. The enforcement template — scraping public facial images is processing of biometric data requiring explicit consent or substantial public interest — applies in principle to any AI scraper harvesting personal data, though enforcement against general-purpose LLM training has been slower and more procedurally cautious (the Garante’s brief ChatGPT ban in spring 2023 was lifted after OpenAI added transparency notices and a deletion mechanism).

    UK ICO guidance issued through 2024 in a consultation series covering lawful basis, purpose limitation, accuracy, data subject rights and transparency. The ICO’s May 2025 response confirmed that legitimate interest CAN cover training scraping but that the balancing test must include foreseeable harms and effective ability to opt out — a more permissive stance than some European DPAs while still requiring concrete safeguards.

EU AI Act Article 53 and the GPAI Code of Practice

Regulation (EU) 2024/1689 — the EU AI Act — entered force 1 August 2024. Article 53 imposes obligations on providers of general-purpose AI models (GPAI), the central one being a ‘sufficiently detailed summary about the content used for training the general-purpose AI model’, applicable from 2 August 2025.

The implementing GPAI Code of Practice was signed in July 2025 by OpenAI, Anthropic, Google, Microsoft, Mistral, Amazon, IBM and others. Meta signed with reservations, citing competitive concerns about training-data disclosure. The Code requires:

  • Summary of training-data sources by category (web, books, code, licensed, synthetic)

  • Respect for opt-out signals expressed in machine-readable form (robots.txt, ai.txt, X-Robots-Tag)

  • Process for handling rights-holder complaints with documented turnaround

  • Disclosure of major data licensing arrangements

    Article 53 enforcement is the responsibility of the AI Office within DG CONNECT; fines for non-compliance can reach 3 percent of global annual turnover or €15M, whichever is higher.

Defensive Tooling Ecosystem

A counter-industry has emerged providing creators with tools to resist or detect AI scraping.

Glaze and Nightshade (University of Chicago)

Developed by Ben Zhao and Heather Zheng’s SAND Lab at the University of Chicago:

  • Glaze (February 2023): adversarial perturbations protecting artist style from style-mimicry fine-tuning. An imperceptible texture overlay shifts the CLIP feature embedding so that fine-tuning on the protected image produces a model that has learned an irrelevant pseudo-style rather than the artist’s true style. 4M+ downloads by 2025.

  • Nightshade (October 2023): goes further — actively poisons models that train on protected images by shifting class associations (e.g. dog images that the model perceives as cats in its CLIP latent space). 250K downloads in the first week. The Nightshade paper was the first peer-reviewed demonstration that small numbers of poisoned samples can corrupt diffusion-model training at scale.

    Both tools are free and run client-side, making them accessible to individual artists. The Glaze/Nightshade research underpins broader academic interest in adversarial training-data poisoning as a creator-empowerment paradigm.

    Spawning AI and Have I Been Trained

    Spawning AI (Mathew Dryhurst, Holly Herndon, Patrick Hoepner) operates:

  • Have I Been Trained: search tool over LAION-5B and other public training sets letting creators check whether specific images appear. Hosts an opt-out registry honoured by Stability AI for SD3 training, Hugging Face dataset filtering, and Shutterstock’s AI training pipeline, covering 80M+ opted-out images as of early 2026.

  • Do-Not-Train HTTP header (X-Robots-Tag: noai, noimageai) and ai.txt extension to robots.txt.

  • Kudurru: real-time scrape-detection network that identifies and blocks active LAION/Common Crawl scraping while it is in progress, distributed across opt-in publisher servers.

    ai.robots.txt

    Community project maintained by Robb Knight and Cory Doctorow providing a canonical robots.txt block-list of known AI crawler user-agents, regenerated continuously as new crawlers are identified. 50K+ websites use the generated file directly via GitHub raw URL, making it the de-facto opt-out lingua franca for small publishers and personal blogs.

    Dark Visitors

    Registry tracking 200+ AI-relevant user-agents with crawler purpose, operator, and recommended robots.txt directives. Subscription service for enterprise publishers.

    IIPC Memento and Archive Provenance

    The International Internet Preservation Consortium’s 2025 working group on preserving robots.txt-aware archive crawls vs training crawls aims to ensure that the Internet Archive’s Wayback Machine and academic web archives retain their distinct legal and ethical footing relative to commercial AI training.

Publisher Pushback Timeline (2023-2026)

  • August 2023: OpenAI publishes GPTBot UA. NYT, BBC, CNN, Reuters block within 48h. By month-end, 26 percent of top-1000 websites had added GPTBot disallow.
  • September 2023: Authors Guild + Silverman sue OpenAI.
  • October 2023: Google introduces Google-Extended pseudo-UA.
  • December 2023: NYT files lawsuit; OpenAI publishes blog response.
  • February 2024: Reddit-Google $60M licensing deal disclosed in S-1; Reddit blocks all unlicensed AI scraping post-IPO March 2024.
  • May 2024: Stack Overflow-OpenAI licensing deal; moderator strike over user-content monetisation.
  • June 2024: Wired/Robb Knight investigation exposes PerplexityBot UA stripping. Apple introduces Applebot-Extended.
  • July 2024: Cloudflare ships one-click AI bot block for 27M+ domains.
  • August 2024: News Corp sues Perplexity SDNY; Anthropic ClaudeBot causes iFixit and Freelancer outages.
  • October 2024: Dow Jones, NY Post sue Perplexity. NYT, WSJ, Reuters, BBC joint statement on AI training transparency.
  • December 2024: UK DSIT/IPO Copyright-and-AI consultation opens.
  • January 2025: Cloudflare AI Labyrinth ships — synthetic content served to non-compliant scrapers.
  • February 2025: Cross-publisher ‘Make It Fair’ campaign by The Times, The Telegraph, The Guardian, The Mirror, The Daily Mail.
  • March 2025: EU AI Act GPAI Code of Practice draft published.
  • July 2025: Cloudflare default-blocks unknown AI crawlers on free tier; EU AI Act GPAI obligations begin enforcement; GPAI Code of Practice signed.
  • November 2025: Getty v Stability AI UK judgment — mixed result.
  • Q1 2026: NYT/OpenAI, Suno/RIAA, Authors Guild/OpenAI settlement negotiations underway.

Academic Context: Research, Theory and Critique

The academic literature on AI scraping spans law, computer science and science-and-technology studies.

Computer Science Foundations

  • Brewster Kahle / Internet Archive (1996-): established the public web-archiving paradigm whose collisions with AI scraping motivate the IIPC’s 2025 provenance work.

  • Common Crawl Foundation (Gil Elbaz, 2008-): the institutional crawler whose monthly dumps became the unintentional foundation of LLM pre-training. Common Crawl’s governance crisis 2023-2024 — whether to add post-hoc opt-outs — illustrates the irreversibility problem of public training corpora.

  • Henner Zeller, Lizzi Harvey, Gary Illyes (Google): principal authors of IETF RFC 9309 (2022).

  • Mark Nottingham: long-running IETF HTTP-area editor and AIPREF Working Group co-chair.

  • Ben Zhao, Heather Zheng (U Chicago SAND Lab): Glaze, Nightshade, and the broader research programme on adversarial perturbations as creator-empowerment.

  • Andres Guadamuz (Sussex / QMUL Centre for Commercial Law Studies): leading UK-Spanish copyright-and-AI scholar; blog TechnoLlama widely cited in policy submissions.

  • Sandra Wachter and Brent Mittelstadt (Oxford Internet Institute): legal-and-ethical work on automated decision-making, the right to explanation, and IP implications of LLM training.

  • Pamela Samuelson (UC Berkeley): foundational US fair-use scholarship now applied to LLM training; her 2024 papers on ‘Generative AI Meets Copyright’ widely cited in NYT v OpenAI briefing.

  • Matthew Sag (Emory): ‘Copyright Safety for Generative AI’ framework analysing memorisation and output filtering.

  • Lilian Edwards (Newcastle): AI-IP-data-protection triangle; key UK academic voice in the 2024 IPO Code of Practice working group.

  • Marisa McVey (UCL): intellectual property and creator rights.

    Science and Technology Studies

  • Iyad Rahwan (Imperial College London, formerly MPI Berlin): machine behaviour as a research field treating AI scrapers and the systems they feed as objects of empirical study.

  • Shannon Vallor (Edinburgh Futures Institute): technomoral framing of automation, including LLM-training extractivism.

  • Stephen Cave and Kanta Dihal (Cambridge Leverhulme CFI): AI imaginaries shaping how public and policymakers understand scraping practices.

    Policy and Advocacy

  • Jess Northend and Jeegar Kakkad (Tony Blair Institute for Global Change): 2023-2025 papers advocating a permissive UK TDM regime to support AI-sector competitiveness against the EU AI Act and US fair-use ambiguity. Influential on Labour government 2024-2026 AI policy.

  • Royal Society (Sir Paul Nurse era): ‘Science in the Age of AI’ (May 2024) and Copyright report (September 2024) recommending bounded TDM exception with transparency obligations.

  • Cory Doctorow (writer, activist, ai.robots.txt co-maintainer): trenchant public-interest critique of scraping economics under the framing of enshittification.

Current Landscape (2026)

As of Q1 2026, the AI-scraping landscape exhibits five characteristics that distinguish it from the pre-2023 status quo:

  1. Crawler-vendor stratification. Foundation-model labs now operate distinct crawlers for distinct purposes (training, search, on-demand fetch) with corresponding opt-out signals. The single-bot era is over.
  2. Default-block edge enforcement. Cloudflare’s July 2025 default-block for unknown AI crawlers means that for the 27M+ Cloudflare-protected domains the burden of justification has flipped: AI vendors must enrol in verified-bot programmes rather than publishers having to actively block.
  3. Commercial licensing as steady state. The $1.2B 2025 licensing market means major publishers have monetised the access channel rather than litigated it; litigation continues but as a leverage tool for licensing negotiation rather than as the primary remedy.
  4. Regulatory codification. EU AI Act Article 53 has set a regulatory floor for training-data transparency that is propagating to the UK (DSIT consultation), Canada (AIDA), and arguably the US (state-level AI accountability legislation in California, Colorado, Texas).
  5. Adversarial countermeasures going mainstream. Glaze, Nightshade, AI Labyrinth and ai.robots.txt have moved from research curiosities to deployed defences with measurable impact on training-data quality.

Quantitative landscape: as of Q1 2026, the Cloudflare Radar AI-bots dashboard tracks 26 named AI crawler families with aggregate request volumes ranging from GPTBot’s ~4 billion requests per day across protected domains to long-tail crawlers (Cohere-AI, YouBot, DuckAssistBot) at the 100K-1M daily request scale. Bytespider and ClaudeBot together account for roughly one-third of all blocked AI requests, reflecting both their aggressive crawl patterns and the default-block policies most publishers have adopted against them. The Imperva 2024 Bad Bot Report cross-section showed that AI-training crawlers were the only bad-bot category to grow more than 100 percent year-on-year, with vendors reporting that approximately 28 percent of all blocked bot traffic in 2024 was attributable to AI-training intent — a category that did not meaningfully exist in their 2021 reports.

Outstanding tensions:

  • Common Crawl’s pre-2023 dumps remain in circulation and continue to be ingested by new models, meaning opt-outs cannot be made retroactive without either Common Crawl Foundation cooperation or per-vendor undertakings (Anthropic’s voluntary respect-retroactively undertaking is the only public example).
  • Open-source models (LLaMA-3, Mistral, Falcon) cannot meaningfully unmix scraped content from weights, leaving them more exposed to injunctive remedies than closed-source vendors.
  • The territoriality loophole exposed by Getty v Stability AI remains until either statutory reform (mooted in EU and UK consultations) or contrary precedent in another jurisdiction.
  • Bot fingerprinting arms race: as Cloudflare improves detection, scrapers improve evasion (residential proxies, browser-real automation via Puppeteer, distributed crawls). Each side’s investment is now measured in eight-figure annual budgets.

UK Context: Academic, Industrial and Policy

Academic Centres

  • Imperial College London: Iyad Rahwan’s machine-behaviour group; Centre for Languages, Culture and Communication AI-ethics group; collaboration with GE Healthcare on consent-aware medical-data pipelines.

  • UCL: Centre for Artificial Intelligence (Mirco Musolesi); Marisa McVey on IP-and-AI.

  • University of Edinburgh: Centre for Technomoral Futures (Shannon Vallor); Edinburgh Futures Institute hosting cross-disciplinary AI governance work.

  • University of Cambridge: Leverhulme Centre for the Future of Intelligence (Stephen Cave, Kanta Dihal); Bennett Institute for Public Policy work on AI regulation.

  • University of Oxford: Internet Institute (Sandra Wachter, Brent Mittelstadt); Reuben College.

  • University of Manchester: Christabel Pankhurst Institute for Public Policy Research on AI governance; Alan Turing Institute Manchester hub.

  • Newcastle University: Lilian Edwards; Open Lab on participatory data governance.

  • Queen Mary University of London (QMUL): Centre for Commercial Law Studies; Andres Guadamuz (visiting).

  • University of Sussex: Andres Guadamuz; Sussex Centre for the Digital Humanities.

    Northern English Industrial

  • Manchester: Co-op Group’s data ethics team operating consent-aware customer-data pipelines; BAE Systems Manchester applied AI division holding defensive-scraping-detection contracts with UK government.

  • Leeds: Sky Betting & Gaming deploying DataDome bot management at scale; Channel 4 AI policy team.

  • Sheffield: University of Sheffield Speech and Hearing group running consented audio-scraping pilots for ASR training; Tata Steel UK applied AI for industrial-vision data licensing.

  • Newcastle: Sage Group (FTSE 100) on enterprise AI data licensing; Newcastle University Open Lab.

  • Liverpool: University of Liverpool’s Computer Science Department on robotics simulation data.

    UK Policy Saga

  • February 2024: UK Government abandons the IPO Code of Practice process after a nine-month working group (BPI, Publishers Association, Society of Authors, BSAC vs DeepMind, Stability, OpenAI, Anthropic) fails to reach consensus on opt-out vs opt-in default.

  • December 2024 - February 2025: DSIT/IPO consultation ‘Copyright and Artificial Intelligence’ proposing a TDM exception with opt-out for rights holders. 11,500 responses, majority opposed. Cross-publisher ‘Make It Fair’ campaign by The Times, The Telegraph, The Guardian, The Mirror, The Daily Mail with simultaneous front-page wrappers and full-page editorials.

  • 2025: Government response delayed pending House of Lords scrutiny; cross-party amendments to the Data (Use and Access) Bill mandating training-data transparency.

  • Tony Blair Institute continues advocacy for permissive TDM regime; Royal Society reports counterbalance with bounded-exception position.

Future Directions (2026-2030)

The 2026-2030 trajectory is shaped by five vectors:

  1. Protocol layer maturation. IETF AIPREF expected to publish RFC drafts 2026-2027 with HTTP-header and machine-readable preference vocabularies replacing the binary robots.txt directive. IAB Tech Lab LLM Content Ingest API or analogous monetisation channel likely to ship at scale by 2027 with major-CDN integration.
  2. Crawler attestation. Cryptographically signed bot identity (via TLS extensions, HTTP signature headers, or DNS-bound key publication) becoming standard, eliminating UA-spoofing as a viable scraper strategy. Cloudflare’s verified-bot programme is the de-facto early implementation.
  3. Per-fetch micropayments. Cloudflare and analogous edge vendors enabling per-fetch HTTP 402 settlement, potentially via stablecoins or web-monetisation APIs. This shifts the licensing layer from annual-deal-with-major-publishers to continuous-micro-payment-to-long-tail.
  4. Regulatory convergence. EU AI Act Article 53 enforcement creating de-facto global standards (the Brussels Effect) as foundation-model providers comply globally rather than maintain separate EU-vs-rest training pipelines. UK and US likely to align on training-data transparency by 2028.
  5. Adversarial training-data poisoning normalisation. AI Labyrinth-style synthetic-content defences, Glaze/Nightshade-style perturbation defences, and Kudurru-style real-time detection becoming standard CDN and CMS features. Foundation-model labs investing in detection-of-detection (filtering poisoned content from training corpora), creating an arms race that elevates training-data quality assurance to a first-class research domain.

Risks and tensions:

  • Open-source / open-weights divergence: if regulatory or licensing burdens fall disproportionately on closed-source vendors, open-source models become the cheap-but-legally-exposed alternative, creating uneven playing field.
  • Web fragmentation: if AI crawlers are systematically blocked, the open web ceases to be a representative training substrate, and foundation-model performance degrades on long-tail and current-events queries unless aggressively licensed.
  • Synthetic-data substitution: as scraping becomes legally and operationally costly, synthetic-data substitution (model-generated training material) accelerates, raising the spectre of model collapse documented by Shumailov et al. (Nature 2024).

Research & Literature

Key references and primary sources:

  • Koster, M. (1994). ‘A Standard for Robot Exclusion.’ Original robots.txt specification.
  • Zeller, H., Harvey, L., Illyes, G. (2022). IETF RFC 9309: Robots Exclusion Protocol. https://datatracker.ietf.org/doc/html/rfc9309
  • Imperva (2024). Bad Bot Report 2024: 49.6 percent bot traffic, 32 percent bad bots.
  • Cloudflare (2024-2025). AI Audit, AI Labyrinth and default-block product launches. https://blog.cloudflare.com/ai-audit/
  • OpenAI. GPTBot, OAI-SearchBot, ChatGPT-User documentation. https://platform.openai.com/docs/bots
  • Anthropic. ClaudeBot crawler documentation and respect-retroactively undertaking 2024.
  • Dark Visitors registry. https://darkvisitors.com/agents
  • ai.robots.txt project (Knight, R., Doctorow, C.). https://github.com/ai-robots-txt/ai.robots.txt
  • Shan, S., Zhao, B., et al. (2023). ‘Glaze: Protecting Artists from Style Mimicry’. USENIX Security 2023.
  • Shan, S., Cryan, J., Wenger, E., Zheng, H., Hanocka, R., Zhao, B. (2023). ‘Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models’. arXiv:2310.13828.
  • Samuelson, P. (2023). ‘Generative AI Meets Copyright’. Science 381(6654).
  • Sag, M. (2023). ‘Copyright Safety for Generative AI’. Houston Law Review.
  • New York Times Company v. Microsoft Corporation et al. (S.D.N.Y. 2023). Case 1:23-cv-11195.
  • Getty Images (US) Inc. v. Stability AI Ltd. UK High Court CL-2023-000080, judgment November 2025.
  • Authors Guild et al. v. OpenAI Inc. (S.D.N.Y. 2023, consolidated MDL).
  • UMG Recordings et al. v. Suno Inc. (D. Mass. 2024); UMG et al. v. Uncharted Labs (S.D.N.Y. 2024).
  • News Corp et al. v. Perplexity AI Inc. (S.D.N.Y. 2024).
  • Garante per la protezione dei dati personali (Italy). Clearview AI provvedimento 2022.
  • CNIL (France). Clearview AI decision SAN-2022-019.
  • Autoriteit Persoonsgegevens (Netherlands). Clearview AI decision September 2024.
  • Regulation (EU) 2024/1689 (EU AI Act). https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  • GPAI Code of Practice (signed July 2025).
  • UK ICO (2024-2025). Generative AI consultation series response, May 2025.
  • UK DSIT/IPO (December 2024). Copyright and Artificial Intelligence consultation.
  • Royal Society (May 2024). ‘Science in the Age of AI’.
  • Royal Society (September 2024). Copyright report.
  • Tony Blair Institute (2023-2025). Northend, J., Kakkad, J. UK AI competitiveness papers.
  • Spawning AI. Have I Been Trained. https://haveibeentrained.com/
  • Shumailov, I., et al. (2024). ‘AI models collapse when trained on recursively generated data’. Nature 631.
  • Common Crawl Foundation. Monthly WARC dumps. https://commoncrawl.org/

Metadata

  • This page was enriched as part of the Phase 6 ontology refinement sprint, pivoting from the pre-enrichment stub (which inadvertently catalogued developer-side scraping libraries — Crawl4AI, Firecrawl, Scrapegraph-AI — that are scraping tools rather than scrapers as a regulated bot category) to the policy-and-legal frame appropriate for the AI Scrapers concept. The pre-enrichment dev-tooling content has been allowed to lapse: those libraries belong on a separate Web Scraping Tools or LLM-Friendly Crawler Libraries page rather than under the AI Scrapers concept which denotes the bot category encountered by publishers.
  • Domain: artificial-intelligence (unchanged — correct under the AI bot-and-training-data reading).
  • IRI correction: from ngm#AIScrapers to artificial-intelligence#AIScrapers reflecting accurate ontological domain prefix.
  • legacy-term-id: AI-1043 (slot following AI-1042 GANs from the previous Phase 6 batch).
  • Authority score: 0.87 reflecting depth of regulatory, legal and industry citation coverage with named cases, named regulators, named statutes, named academic researchers and quantitative market data.
  • Cross-references: deliberately bridges to Training Data, Bot Management and Copyright as the three downstream concept clusters likely to be enriched in subsequent batches.

Provenance

  • domain-correction: iri prefix moved from ngm# to artificial-intelligence# to reflect correct ontological domain; preferred-term retained as ‘AI Scrapers’ with alternative-terms expanded (AI Crawlers, AI Training Crawlers, LLM Bots, Generative-AI Web Bots, Foundation-Model Scrapers, AI Web Harvesters)