Death of the Internet is a concept cluster in the ics-society and Media Theory domains capturing the convergent degradation of the public World Wide Web as a space for authentic human knowledge-exchange, discovery, and democratic discourse.

Semantic Classification

  • domain-correction: infrastructure → ethics-society (iri updated from infrastructure#DeathOfTheInternet to ethics-society#DeathOfTheInternet; uri updated accordingly; legacy-term-id ES-1041 assigned)

Content

Compositional Relationships (Components)

SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:DeadInternetTheory))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:Enshittification))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:HabsburgAIModelCollapse))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:SyntheticContentSaturation))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:BotTrafficDominance))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:FilterBubble))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:DarkForestDynamics))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:APIMonetisation))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:AuthenticContentScarcity))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:hasPart es:BackgroundTokensProblem))

## Dependency Relationships
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:requires es:AuthenticHumanContent))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:requires es:ContentProvenanceSignals))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:requires es:PlatformAccountability))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:requires es:OpenAPIAccess))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:AdvertisingEconomics))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:PlatformMonopoly))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:AITrainingDataQuality))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:SearchEngineEconomics))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:SurveillanceCapitalism))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:dependsOn es:NetworkEffectsLockIn))

## Capability Relationships
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:EpistemicDegradation))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:AlgorithmicCapture))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:ModelCollapseLoop))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:DigitalDivideAmplification))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:ManipulationOfPublicOpinion))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:enables es:DisinformationEcosystem))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:supports es:DecentralisedWebAlternatives))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:supports es:ContentProvenanceAdoption))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:supports es:OpenSourceAIMovement))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:supports es:CryptographicAttestation))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:supports es:HumanCentricGovernance))

## Implementation Relationships
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:C2PAStandard))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:ContentCredentials))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:RobotsTxtEnforcement))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:APIRateLimiting))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:AgeVerification))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:implements es:OnlineSafetyActEnforcement))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:uses es:LargeLanguageModels))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:uses es:BotDetectionSystems))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:uses es:AlgorithmicFeedCuration))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:uses es:ProgrammaticAdvertising))

## Reduction Relationships
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:HumanTrafficProportion))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:AuthenticContentDiversity))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:DiscoverabilityOfIndependentContent))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:EpistemicQualityOfPublicWeb))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:TrustInOnlineInformation))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:SerendipitousDiscovery))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:reduces es:IndependentPublisherEconomics))

## Contrast and Association Relationships
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:contrastsWith es:OpenWeb))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:contrastsWith es:FederatedSocialNetworks))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:relatedTo es:SurveillanceCapitalism))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:relatedTo es:GlobalInequality))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:relatedTo es:AIGovernance))
SubClassOf(es:DeathOfTheInternet
  ObjectSomeValuesFrom(es:relatedTo es:AgentMediation))

## Data Properties
DataPropertyAssertion(es:hasIdentifier es:DeathOfTheInternet "ES-1041"^^xsd:string)
DataPropertyAssertion(es:authorityScore es:DeathOfTheInternet "0.87"^^xsd:decimal)
DataPropertyAssertion(es:botTrafficProportion2024 es:DeathOfTheInternet "0.496"^^xsd:decimal)
DataPropertyAssertion(es:maliciousBotProportion2024 es:DeathOfTheInternet "0.32"^^xsd:decimal)
DataPropertyAssertion(es:aiNewsWebsitesMay2024 es:DeathOfTheInternet "1265"^^xsd:integer)
DataPropertyAssertion(es:syntheticContentLowAuthority2024 es:DeathOfTheInternet "0.52"^^xsd:decimal)
DataPropertyAssertion(es:humanAuthoredWebProportion2026 es:DeathOfTheInternet "0.40"^^xsd:decimal)

## Property Constraints
SubClassOf(es:DeathOfTheInternet
  DataAllValuesFrom(es:requiresProvenanceSignal xsd:boolean))
SubClassOf(es:DeathOfTheInternet
  DataSomeValuesFrom(es:collapseStageType xsd:string))
SubClassOf(es:DeathOfTheInternet
  DataMinCardinality(1 es:hasBotTrafficMeasure xsd:decimal))
SubClassOf(es:DeathOfTheInternet
  DataMaxCardinality(1 es:hasModelCollapseGeneration xsd:integer))

## Annotations
AnnotationAssertion(rdfs:label es:DeathOfTheInternet "Death of the Internet"@en)
AnnotationAssertion(rdfs:comment es:DeathOfTheInternet "Concept cluster capturing the convergent degradation of the public web through bot traffic dominance (49.6% non-human 2024, Imperva), synthetic content saturation (52-68% AI markers in new low-authority content, Originality.ai 2024), Habsburg AI model collapse (Shumailov et al. Nature 2024 recursive training degradation), and platform enshittification (Doctorow 2023 three-stage decay), defended against by C2PA provenance signals, decentralised web alternatives (Solid/ActivityPub), and UK Online Safety Act 2023 implementation."@en)
AnnotationAssertion(dcterms:identifier es:DeathOfTheInternet "ES-1041"^^xsd:string)
AnnotationAssertion(dcterms:subject es:DeathOfTheInternet "Media Theory, Platform Decay, AI Ethics, Internet Governance, Information Integrity, Surveillance Capitalism"@en)

)

Property Characteristics

AsymmetricObjectProperty(es:requires) AsymmetricObjectProperty(es:enables) AsymmetricObjectProperty(es:implements) AsymmetricObjectProperty(es:reduces) TransitiveObjectProperty(es:dependsOn) FunctionalDataProperty(es:botTrafficProportion2024) FunctionalDataProperty(es:syntheticContentLowAuthority2024) FunctionalDataProperty(es:humanAuthoredWebProportion2026)

About the Death of the Internet

  • Death of the Internet is not a single phenomenon but a conceptual cluster — a syndrome of mutually reinforcing degradation processes that together threaten to transform the World Wide Web from a space of authentic human discourse into a low-quality synthetic environment dominated by automated systems serving commercial and political interests. The concept draws together empirical measurements of bot traffic dominance, theoretical frameworks for Platform Decay, scientific findings on AI Training Data quality collapse from recursive synthetic training, and political-economic analyses of how Surveillance Capitalism and Advertising Economics systematically degrade the user experience.
  • The Dead Internet Theory originated as an anonymous post on the Agora Road forum in 2021, proposing that the internet had been “largely taken over by artificial intelligence and automatic bots” since approximately 2016–2017, coordinated by a small number of corporations and state actors to manufacture consensus, push propaganda, and manipulate public behaviour. When first published, it was widely dismissed as conspiracy theory. By 2024, empirical research had validated its core claims sufficiently to transform it from fringe speculation into a subject of serious academic and policy attention.
  • The key empirical anchors validating the theory are:
    • Imperva Bad Bot Report 2024: 49.6% of all internet traffic non-human; 32% actively malicious bots — the highest proportion since tracking began in 2013. AI Scrapers including GPTBot, ClaudeBot, and Amazonbot constitute a growing share of the “legitimate” bot category.
    • NewsGuard AI News Tracker: 1,265+ AI-generated news and information websites with minimal human oversight by May 2024, up 25-fold from 49 in April 2023. These sites share structural characteristics: 100–1,000+ AI-generated articles per day, no editorial oversight, programmatic advertising monetisation, frequent factual errors and hallucinations.
    • Originality.ai Corpus Sampling 2024: AI-content markers in 52–68% of newly published content across low-authority domains, with the proportion rising approximately 8–12 percentage points per year.
    • Cloudflare 2024 State of Application Security Report: Automated traffic exceeding human traffic in 14 of 17 measured industry sectors.
    • CIA officer testimony (Yahoo Finance 2023): Estimate that up to 80% of Twitter/X accounts could be non-human, consistent with Elon Musk’s own pre-acquisition claim before he retracted it.
  • The political-economic driver is what Surveillance Capitalism analyst Shoshana Zuboff terms “behavioural surplus”: data generated by user activity as raw material for predictive products sold to advertisers. Maximising data generation requires maximising engagement regardless of content quality, creating a structural incentive for platforms to optimise for attention capture rather than epistemic value. This is the “original sin of the internet” — the advertising-funded model — documented by Tim Wu in The Attention Merchants (2016) and extended by Ben Thompson’s “The Agentic Web and Original Sin” (Stratechery, 2025) to the emergence of Agentic Internet clients as the dominant web consumers.

Enshittification: The Platform Decay Lifecycle

  • Cory Doctorow’s “enshittification” framework, introduced in his January 2023 essay “Tiktok’s enshittification” on Pluralistic.net and subsequently popularised in Locus Magazine and the Financial Times (“‘Enshittification’ is coming for absolutely everything,” 2023), provides the most widely adopted theoretical model for Platform Decay.
  • The model identifies a three-stage lifecycle applicable to every major platform:
    • Stage 1 — Subsidise Users: The platform suppresses rent-seeking to attract users through genuine value. Social Media platforms offer organic reach to all pages and publishers; AI Search engines provide free, comprehensive results; e-commerce platforms offer low seller fees and favourable terms to build product catalogue.
    • Stage 2 — Exploit Users to Serve Business Customers: Once users are locked in through Network Effects, switching costs, and data dependencies, the platform begins extracting value. Facebook suppresses organic reach for pages, charging them for access to their own followers. Amazon degrades search results with paid placement. Google increases advertisement density in search results. The value transferred from users funds subsidies to business customers (advertisers, API clients, enterprise subscribers) who provide the platform’s primary revenue.
    • Stage 3 — Exploit Business Customers for Shareholders: Once business customers are similarly locked in, the platform extracts from both. Twitter/X eliminates free API access and monetises bot and NSFW content. Google’s leak of internal search ranking documents (SparkToro analysis, June 2024) reveals systematic favouring of established brands over quality independent content. TikTok charges creators for algorithm access. The platform degrades until it collapses or faces regulatory disruption.
  • Doctorow traces this lifecycle across Google Search (replacing organic results with AI Overviews and advertisements, degrading findability for independent content), Meta (suppressing organic reach then monetising political content), Amazon (paid placement degrading search quality, logistics fee extraction from third-party sellers), and Twitter/X (API monetisation at $42,000/month for enterprise access, removal of third-party moderation tools). The structural driver is monopoly capture enabled by network effects: users cannot leave Facebook because their social graph is captive to the platform; sellers cannot leave Amazon Marketplace because their customers are there.
  • The SparkToro analysis of leaked AI Search API documents (June 2024) documented that Google’s ranking signals had shifted substantially toward brand recognition and domain authority, systematically disadvantaging independent and specialist publishers. As Rand Fishkin’s analysis noted: “From 1998–2018, one could reasonably start a powerful marketing flywheel with SEO for Google. In 2024, that’s no longer realistic on the English-language web in competitive sectors.”
  • The advertising-funded model creates a perverse relationship between content quality and platform economics. The Programmatic Advertising system — RTB auctions placing advertisements on any page meeting targeting criteria, regardless of content quality — means that AI content farms generating low-quality articles at industrial scale can monetise through the same advertising networks as high-quality journalism, with no economic penalty for quality degradation. The result documented in WIRED’s investigation into programmatic advertising and Deepfakes and fraudulent content is an advertising-funded disinformation ecosystem in which brand advertisers unknowingly fund misinformation sites.

Habsburg AI: Model Collapse and Recursive Degradation

  • Shumailov et al.’s “The Curse of Recursion: Training on Generated Data Makes Models Forget” (Nature, May 2024) provides the most rigorous scientific treatment of Habsburg AI dynamics — the progressive statistical degeneration of Large Language Models trained recursively on AI-generated outputs.
  • The mechanism is information-theoretic. Each generation of AI-generated training data introduces a systematic bias toward the modal distribution of the previous model, because AI models sample from their own learned distribution rather than drawing from the original human-authored corpus. The noise compounds geometrically: if the original model has 5% error rate on rare linguistic patterns, generation-2 training data has 5% + systematic_bias errors, generation-3 data has those errors plus new biases introduced by generation-2’s already-degraded distribution. The researchers observed measurable degradation within 3–5 training generations, with catastrophic failure modes appearing by generation 8–10.
  • The degeneration proceeds in three stages mapped to information-theoretic failure modes:
    • Tail loss (generations 3–5): Low-frequency but important linguistic patterns — rare vocabulary, minority languages, specialised technical terminology, dialectal variation — disappear from model outputs as they are not represented in the modal synthetic training data. The model begins to produce homogeneous outputs that reflect the average of previous model outputs rather than the diversity of human expression.
    • Modal collapse (generations 5–8): Outputs cluster toward a narrow distribution. The model produces fewer distinct phrasings, perspectives, and framings. Diversity metrics (type-token ratios, perplexity scores on held-out human text, entropy of output distributions) decline measurably. Large-Scale Pretrained Foundation Model trained on collapsed data produce less useful responses for tasks requiring creative, diverse, or specialised outputs.
    • Catastrophic degradation (generations 8–10+): The model produces incoherent or repeatedly similar outputs for queries that require diversity or rare knowledge. In Shumailov et al.’s image generation experiments, later-generation models produced outputs that were visually degraded and clustered. In text generation experiments, models began producing repetitive, low-coherence outputs for complex queries.
  • The Habsburg analogy is precise. The Habsburg dynasty’s practice of royal inbreeding — marrying within a narrow gene pool of related European royal families to preserve lineage purity — produced progressive accumulation of deleterious recessive traits across generations. Charles II of Spain, the last Habsburg monarch (1661–1700), suffered severe physical and cognitive disabilities including an enlarged jaw, tongue, and head, inability to chew food, and early cognitive decline — the endpoint of generations of inbreeding. Just as genetic diversity is required for population fitness, epistemic and linguistic diversity in training data is required for Large-Scale Pretrained Foundation Model quality, and systematic restriction to a narrow distribution produces analogous degeneration.
  • The policy implication is that preservation of the Background Tokens corpus — pre-AI authentic human discourse — is a strategic imperative for maintaining AI system quality. The Data Provenance Initiative (Epoch AI, 2024) estimated that uncontaminated human-authored text now constitutes less than 40% of the publicly indexable web by token count, declining at 8–12 percentage points per year. C2PA provenance signing, training data watermarking, and the Internet Archive’s preservation efforts represent the primary mechanisms for maintaining access to authentic training corpora as the open web degrades.

Synthetic Content Saturation and the Background Tokens Problem

  • The Background Tokens problem — identified and named on the Latent Space podcast — refers to the finite, non-renewable corpus of authentic pre-AI human discourse captured in Common Crawl, Books3, Wikipedia, Project Gutenberg, and similar training datasets. This corpus, representing authentic human expression across centuries and billions of authors, is consumed in a single training pass by each major model generation and cannot be replenished at the rate it is diluted by AI-generated content.
  • Prior to 2022, the Common Crawl corpus was estimated to be approximately 85% authentic human-authored content. By 2024, the Data Provenance Initiative estimated that uncontaminated human-authored text constituted approximately 40% of the public web by token count. The industrial economics of AI content generation drive this decline: at 0.01 per AI-generated article, operators can publish 10,000+ articles per day for 100. An estimated 1,265+ dedicated AI news sites (NewsGuard, May 2024) publish at this cadence, alongside individual operators using ComfyUI Workflows and API pipelines for long-tail content targeting.
  • The AI Search interaction is structurally perverse. Google’s AI Overviews (launched May 2024) reduce traffic to source websites by answering queries directly in the search results page, destroying the advertising economics that previously incentivised high-quality content creation. Simultaneously, AI content farms can flood the search index with topically relevant but low-quality content at a cost structure that legitimate publishers cannot match. The result is a market failure in content quality: the advertising economics that historically rewarded high-quality content are disrupted by AI-generated volume competition, while the AI systems that replace search are themselves degraded by training on the polluted content corpus they helped create.
  • The right-wing media asymmetry identified in WIRED (2023) represents a further distributional distortion. While mainstream and left-leaning news sites increasingly blocked AI training scrapers via robots.txt, right-wing media outlets disproportionately welcomed scraping. As noted on the Latent Space podcast’s discussion of Bias in Large Language Models, this differential willingness to permit training data collection creates systematic political skew in future model training distributions — not through intentional design but through the differential economics of content licensing and political stance toward AI regulation.
  • A secondary effect documented in academic research is the contamination of peer-reviewed scientific literature by AI-generated content. A BioRxiv preprint from February 2024 documented widespread fraudulent images in systematic reviews of preclinical depression research, evidencing that AI-assisted fraud has penetrated even the most rigorously gatekept information ecosystem. The Deepfakes and fraudulent content problem thus extends from consumer web content to scientific publishing, with potential consequences for Large-Scale Pretrained Foundation Model trained on scientific corpora.

Bot Traffic and the Dead Internet in Practice

  • The quantitative evidence for bot dominance is now extensive and multi-sourced. Imperva’s Bad Bot Report 2024 found that 49.6% of all internet traffic was non-human in 2024, with the following breakdown:
    • 17.6% good bots: Legitimate crawlers — Googlebot, Bingbot, GPTBot, ClaudeBot, Amazonbot, the Perplexity crawler — performing indexing and training data collection.
    • 32% bad bots: Malicious automation — credential stuffing, web scraping for competitive intelligence, ad fraud, inventory hoarding (automated ticket purchasing, product scalping), fake account creation, and carding attacks. The 32% figure was the highest recorded since Imperva began tracking in 2013.
  • CHEQ’s parallel tracking found broadly consistent figures, estimating 36–40% invalid traffic in paid advertising contexts — meaning that approximately one-third to two-fifths of advertising spend reaches non-human “eyeballs.”
  • The practical consequences extend across multiple domains of social and economic life:
    • Financial markets: Bot-driven trading creates flash crashes and artificial liquidity signals, documented in SEC enforcement actions and academic market microstructure research.
    • Political discourse: Coordinated bot networks manufacture apparent consensus around fringe positions — the “astroturfing” documented by Europol’s 2023 warning on AI-generated influence operations targeting European elections. The interaction with Digital Society Surveillance creates a feedback loop in which manufactured consensus shapes surveillance-informed algorithmic amplification.
    • Consumer markets: Bot-driven ticket scalping (Ticketmaster CNBC testimony, 2017 and 2023 updates), fake review manipulation (Amazon’s ongoing litigation), and inventory hoarding (GPU availability crises partly attributed to automated purchasing bots) distort market signals and harm ordinary consumers.
    • Academic publishing: The BioRxiv preprint (February 2024) on fraudulent images in preclinical depression systematic reviews; the “DDoS attack of academic bullshit” characterisation of AI-generated academic content flooding preprint servers and low-quality journals.
    • Social media authenticity: Russian bot inflation of Instagram influencer follower counts into the tens of millions (documented in Adweek, 2022); the X/Twitter Super Bowl 2024 traffic analysis suggesting majority fake engagement during the most-watched event in US television history; the ancient spam account phenomenon on X in 2024 where a bot posted an AI-generated image description without an image and attracted hundreds of admiring bot replies generating fictional human responses to a non-existent image — a closed loop of AI generating content for AI to consume with no human participation.
    • Dating and relationship formation: AI companions applications mediate a growing proportion of romantic interactions; RIZZ and similar AI dating coach applications assist in composing messages on dating platforms; the prospect of AI-to-AI interaction across dating layers — bots persuading bots that persuade bots — represents an endpoint of the Dead Internet Theory dynamic in the most intimate domain of human social life.
  • Jailbroken Large-Scale Pretrained Foundation Model can already solve CAPTCHA human-verification challenges, and open-source vision models are within 12–18 months of CAPTCHA-solving capability at sub-$0.001/challenge costs, potentially triggering a discontinuous expansion of automated web activity that would further accelerate the Dead Internet dynamics.

The Dark Forest Theory and Digital Retreat

  • Yancey Strickler’s “Dark Forest Theory of the Internet” (Medium, 2019) and Maggie Appleton’s “The Expanding Dark Forest and Generative AI” (maggieappleton.com, 2023) provide the metaphorical framework for the human response to synthetic content saturation and bot dominance.
  • Drawing on Liu Cixin’s Three-Body Problem, in which the universe is silent not because it lacks intelligence but because all intelligent civilisations have learned to hide to avoid being targeted by predators, the theory proposes that authentic human discourse has not disappeared from the internet — it has retreated into private, non-indexed, non-gamified spaces: private messaging channels, encrypted group chats, Discord servers, newsletters, paid communities, podcasts, and in-person networks.
  • The internet has become, in this analysis, a dark forest:
    • In response to advertisements, tracking, trolling, hype, and predatory manipulation, authentic human communities have retreated to non-indexed spaces.
    • These “dark forests” are non-indexed, non-optimised, and non-gamified environments — they do not participate in the attention economy and are therefore invisible to advertising-funded platforms and AI Scrapers.
    • Examples include private Telegram channels and groups, Signal communities, Discord servers (with private channels invisible to search), newsletters sent to opt-in subscriber lists, Substack publications behind paywalls, and podcasts (where meaning is conveyed through tone and intonation in ways that resist algorithmic gamification).
    • The dark forest dynamic accelerates as authentic signal becomes harder to find in public spaces: each additional bot or AI-generated piece of content on a public platform raises the cost of finding authentic human interaction there, making private alternatives more attractive.
  • Appleton’s 2023 extension argues that Generative AI has dramatically accelerated the retreat into the dark forest by raising the cost of authentic signal detection in public spaces. When Large Language Models can perfectly mimic human conversational patterns, authentic human communication becomes indistinguishable from bot output to casual observers, creating the “lemon market” dynamic (Akerlof 1970) in which uncertainty about authenticity drives down the value of all public communication and accelerates migration of high-quality discourse to private, authenticated spaces.
  • Caroline Busta’s 2020 essay “The Internet Didn’t Kill Counterculture — You Just Won’t Find It on Instagram” anticipates this analysis, arguing that authentic subcultural production had already migrated away from platform-indexed spaces before the Generative AI inflection point. The Instagram-ification of counterculture — optimising aesthetic production for algorithmic amplification — had already hollowed out the authentic subcultural energy that the early web enabled, even before AI-generated content created a positive abundance of synthetic imitations.
  • The philosophical corollary is that the internet as a democratic public sphere — the original vision of the open web as a space for horizontal, peer-to-peer knowledge exchange — has been effectively captured and corrupted by the combination of platform monopoly, Surveillance Capitalism, and synthetic content. The “private internet” of dark forests is a rational response to this capture, but it comes with epistemic costs: fragmentation of discourse, echo chamber dynamics, and the loss of the serendipitous cross-community discovery that characterised the early public web.

Platform API Monetisation and Balkanisation

  • The 2023–2024 API monetisation events at Reddit and X/Twitter represent a structural inflection point in the Balkanisation of the open web. Both events were explicitly designed to prevent AI training data collection by third parties whilst simultaneously flooding their own feeds with algorithmic and sponsored content.
  • Reddit API Monetisation (June 2023): Reddit announced $12,000/month pricing for its data API, triggering the largest protest in Reddit’s history (approximately 8,000 subreddits going dark for 48 hours). The pricing made it economically infeasible for academic researchers, AI Scrapers, and third-party application developers to continue existing access patterns. Reddit’s motivation was dual: capturing revenue from AI companies that had been training on Reddit data for free, and retaining proprietary access to authentic human opinion signal as a competitive asset for Reddit’s own AI and AI Search products.
  • X/Twitter API Monetisation (February 2023 onward): X/Twitter eliminated free API access and implemented tiered pricing (5,000/month Pro, $42,000/month Enterprise). This destroyed the ecosystem of third-party moderation tools that had provided significant community safety infrastructure (Botometer, Perspective API integrations, community moderation bots) whilst monetising data access for AI companies. The downstream consequence for Legacy Media and academic research on political discourse was severe: Twitter had been a primary source of authentic real-time public opinion signal for researchers, and its effective closure to non-paying API users severed that access.
  • The internet is becoming Balkanised as large data aggregators close their APIs and prevent scraping. This is already having measurable effects on AI Search quality — Reddit and Twitter/X had been among the most consequential sources of authentic human opinion signal for search quality evaluation and AI training, and their removal from open-access accelerated the shift toward AI-generated synthetic content as the marginal unit of web content.
  • The right-wing media scraping asymmetry (WIRED, 2023) creates an additional distributional distortion. Right-wing media outlets disproportionately permitted AI training scraping whilst mainstream and progressive outlets blocked it, creating systematic political skew in future model training distributions. As noted in discussions of Bias in Large Language Models, this creates AI systems with embedded political bias not from any intentional design but from the differential economics of content licensing.

AI-Augmented Search: The Epistemic Cost

  • The replacement of traditional keyword search with AI-generated summaries — Google AI Overviews (May 2024), Bing Copilot, Perplexity, SearchGPT — represents a qualitative change in the epistemology of web discovery that connects to multiple AI Risks.
  • Traditional keyword search returned links to source documents, preserving the ability for users to evaluate sources, check context, and follow citation chains. AI Overviews return synthesised answers that:
    • Remove source attribution, making it impossible to evaluate the credibility or context of the claimed information.
    • Mask uncertainty, presenting probabilistic syntheses as confident factual statements.
    • Blend multiple sources without indicating weighting, creating a false impression of consensus.
    • Hallucinate at rates between 1.5% (Google’s internal testing) and 7–12% (independent audits by The Markup and Search Engine Land for complex factual queries).
  • The energy economics of AI Search add a second dimension of cost. MIT’s GenAI team (Luccioni et al. 2024) estimated that a single AI-assisted search query consumes approximately 10x the energy of a traditional keyword search: 3.0 Wh versus 0.3 Wh per query. At Google’s scale of approximately 8.5 billion searches per day, replacing 50% with AI-assisted responses would increase search electricity consumption by approximately 12.75 TWh/year — roughly equivalent to the annual electricity consumption of Slovenia. The convergence of epistemic degradation and environmental cost — paying twice, in information quality and carbon emissions, for a worse epistemic outcome — is one of the most cited critiques in the AI Ethics literature on AI Search.
  • The Toilet Theory of the Internet (The Atlantic, May 2024) provides an analogy: just as the introduction of flush toilets solved a surface-level sanitation problem by pushing waste downstream rather than eliminating it, AI Overviews “solve” the problem of information overload by pushing the epistemic work downstream to an AI system that cannot actually bear it — producing confident-sounding hallucinations that create more epistemic harm than the original information overload problem.
  • The interaction with Global Inequality is severe. The International Telecommunication Union has documented that AI-powered search tools are inaccessible to regions with weaker infrastructure. The shift from free keyword search (available on any browser with a 0.3 Wh energy cost) to paid AI-search tiers removes viable knowledge discovery from approximately 98.5% of the world — recreating, in the search layer, the exact access inequality that the early web had promised to resolve.

Components and Architecture

  • The Death of the Internet concept cluster comprises seven interacting subsystems, each connecting to distinct pages in the ontology:
  • Bot Infrastructure encompasses the technical stack for deploying non-human traffic, connecting to AI Scrapers and Cyber Security and Cryptography:
    • Residential proxy networks enabling bots to evade IP-based detection by routing through real consumer IP addresses.
    • CAPTCHA-solving services using human micro-workers or AI vision models (costs falling from 0.10–$0.30/1,000 in 2024).
    • Headless browser automation frameworks (Playwright, Puppeteer, Selenium) enabling human-like browsing patterns.
    • LLM-powered conversational bots capable of passing Turing tests in limited contexts.
  • Content Synthesis Infrastructure encompasses the AI content pipeline, connecting to Large Language Models, ComfyUI Workflows, and Large-Scale Pretrained Foundation Model:
    • LLM API access with sub-$0.01/1,000 token costs making mass content generation economically viable.
    • Automated SEO optimisation layers targeting long-tail keyword combinations.
    • Programmatic advertising integration via Google AdSense and RTB networks.
    • Content delivery network infrastructure for high-volume publishing at scale.
  • Algorithmic Feed Systems encompass the recommendation engine stacks of major platforms, connecting to Social Media, Cognitive AI, and AI Adoption:
    • TikTok’s For You Page (optimising for watch time, creating filter bubble dynamics documented in AI Risks).
    • Instagram Reels and Facebook News Feed (Meta’s engagement-maximisation architectures).
    • YouTube Recommendations (documented by Guilluy and Ribeiro for radicalisation pathways).
    • Twitter/X’s algorithmic timeline (post-Musk acquisition: reduced moderation, increased political content amplification).
  • Platform Economics encompasses the advertising auction mechanisms, connecting to Surveillance Capitalism and Competition in AI:
    • Google AdSense and the RTB programmatic advertising ecosystem.
    • Meta Audience Network.
    • Amazon Sponsored Products.
    • Affiliate marketing networks funding AI content farms.
  • Data Provenance Systems encompass the emerging defence layer, connecting to Cryptography Security and Privacy, Blockchain Network, and Digital Signature:
    • C2PA metadata schemas (Coalition for Content Provenance and Authenticity) for cryptographic signing of media.
    • Adobe Content Credentials (C2PA implementation in Photoshop, Firefly, Premiere).
    • Blockchain-anchored attestation services for content timestamping and provenance.
    • News agency C2PA adoption consortium (AP, Reuters, AFP joint implementation).
  • Regulatory Architecture encompasses governance responses, connecting to Ofcom, EU AI Act Regulatory Instrument, UK Online Safety Act, and Consumer Protection:
    • Ofcom’s Online Safety Act 2023 implementation framework (Category 1 transparency reports, age assurance, recommender system transparency).
    • EU AI Act Article 50 synthetic content transparency requirements (effective August 2026).
    • DSA Article 34 risk assessment requirements for very large online platforms (effective February 2024).
    • FTC 2024 rule prohibiting AI-generated fake testimonials.
  • Decentralised Alternatives encompass post-enshittification web infrastructure, connecting to Decentralised Web, Solid, Blockchain Network, and Distributed Identity:
    • Solid/WebID (Tim Berners-Lee’s data pod model; W3C specification; MIT CSAIL implementation; Inrupt Ltd commercialisation from Cambridge UK).
    • ActivityPub federation (W3C Recommendation; Mastodon 10M+ users by 2024; Pixelfed, PeerTube, Misskey).
    • IPFS/Filecoin decentralised content-addressed storage.
    • Nostr protocol (decentralised social graph with cryptographic identity).

Use Cases / Major Families

  • Dead Internet Theory Validation Research: Academic and journalistic investigations quantifying the proportion of non-human internet activity. Primary exemplars:
    • Imperva Bad Bot Report series (2013–2024) — annual measurement methodology combining passive traffic analysis, honeypot deployment, and bot signature classification.
    • CHEQ Invalid Traffic Report — parallel measurement focused on paid advertising contexts, estimating 36–40% invalid traffic.
    • NewsGuard AI News Tracking — qualitative and quantitative assessment of AI-generated news and information websites.
    • Originality.ai corpus sampling — applying AI detection classifiers to newly published web content at scale.
    • Stanford Internet Observatory studies on coordinated inauthentic behaviour — documenting bot network operations and state-sponsored information operations on major platforms.
  • Platform Decay Analysis: Empirical and theoretical work mapping the enshittification lifecycle:
    • Doctorow’s Pluralistic.net essay series providing the theoretical framework and platform-by-platform case studies.
    • SparkToro analysis of the leaked Google Search API documents (June 2024) providing internal evidence of algorithmic shifts against independent publishers.
    • Rand Fishkin’s SEO industry research documenting the decline of independent content discoverability.
    • Ben Thompson’s Stratechery analysis of the advertising-funded internet’s original sin and its agentic evolution.
    • The Guardian’s Zombie Internet investigation (May 2024) documenting the economics of AI content farms.
  • Habsburg AI Mitigation Research: Technical research on preventing and detecting model collapse in recursive training pipelines:
    • Training data provenance tracking (C2PA integration with training pipelines; Data Provenance Initiative consortium methodology).
    • Synthetic data detection and filtering (watermarking schemes, linguistic pattern analysis, retrieval-augmented filtering).
    • Human feedback amplification (maintaining annotation team scale proportional to synthetic content volume in training corpora).
    • Diversity-preserving training objectives (techniques to prevent modal collapse by explicitly penalising distribution narrowing during fine-tuning of Large-Scale Pretrained Foundation Model).
    • Constitutional AI approaches — training on human value judgements about output quality rather than purely behavioural imitation.
  • Content Provenance and Authentication: Technical implementation of provenance defence:
    • C2PA Standard implementation in Adobe Creative Cloud applications (Photoshop, Premiere, Firefly) with mandatory signing for AI-generated content.
    • News agency consortium adoption (AP, Reuters, AFP implementing C2PA for all wire service photographs from 2024).
    • Browser-based provenance display (Chrome and Firefox Content Credentials extension support).
    • Blockchain-anchored attestation for long-form text content (Starling Lab methodology for conflict journalism provenance).
  • Regulatory and Policy Responses: Government interventions targeting specific manifestations:
    • UK Ofcom Online Safety Act 2023: transparency reports (January 2024), age assurance (July 2025), recommender system transparency (expected 2026).
    • EU AI Act Article 50: disclosure requirements for AI-generated content effective August 2026.
    • DSA Article 34: risk assessment for recommender systems in very large online platforms (20M+ EU users), effective February 2024.
    • FTC 2024 Rule on Fake Reviews: prohibiting AI-generated fake testimonials and incentivised reviews.
    • California AB 3211 (Content Provenance and Authenticity Act, signed 2024): requiring C2PA metadata for AI-generated content distributed in California.

Academic Context

  • The academic treatment of internet degradation spans multiple disciplines with distinct analytical emphases:
  • Information Science:
    • Metzler et al. 2021 “Rethinking Search” at Google Research — on information retrieval quality degradation under neural and generative approaches.
    • Gao et al. 2020 on The Pile — documenting an 800GB authentic training corpus and the filtration choices that shape training data quality.
    • Dodge et al. 2021 on C4 corpus documentation — revealing the scale of contamination, duplication, and quality variation in Common Crawl derivatives.
    • Data Provenance Initiative consortium 2024 “Consent in Crisis” — documenting the decline of the AI data commons and the shift from 85% to 40% authentic human text on the public web.
  • Media Theory:
    • Postman’s Technopoly (1992) — the foundational argument that technology systems can capture culture and re-orient it around their own operational imperatives, prefiguring the engagement-maximisation dynamic.
    • Carr’s The Shallows (2010) — empirical and neuroscientific case for the internet’s effects on deep reading, attention, and cognition.
    • Pariser’s The Filter Bubble (2011) — the Filter Bubble thesis that algorithmic curation traps users in echo chambers, reducing epistemic diversity.
    • Zuboff’s The Age of Surveillance Capitalism (2019) — the definitive political-economic framework for understanding platform incentives as extraction of behavioural surplus.
    • Doctorow’s enshittification framework (2023) — the operational three-stage lifecycle model connecting Zuboff’s structural analysis to specific platform dynamics.
    • Appleton’s dark forest extension (2023) — the most cited synthesis of the human response to synthetic content saturation.
  • Computer Science:
    • Shumailov et al. Nature 2024 — the foundational empirical and theoretical proof of model collapse from recursive synthetic training.
    • Candès and Recht 2009 on distributional shift — mathematical framework for understanding training-deployment distribution divergence that underpins model collapse analysis.
    • Sadasivan et al. arXiv 2023 — theoretical proof of the impossibility of reliable AI text detection under capable generative models, constraining the detection-based defence approach.
    • Empirical studies of Bias in Large Language Models arising from training data composition asymmetries (right-wing/left-wing scraping differential; language and cultural representation gaps).
  • Political Economy:
    • Wu’s The Attention Merchants (2016) — historical analysis of the commodification of human attention from print advertising through social media.
    • Zuboff’s surveillance capitalism framework connecting platform incentives to the structural degradation of epistemic quality.
    • Haidt’s The Anxious Generation (2024) — the most widely-read synthesis of engagement-maximisation effects on adolescent mental health and Employment Social Contract Under Automation.
    • Stratechery analyses of platform economics and the advertising original sin (Thompson 2025).
  • Political Science and Democracy Studies:
    • Europol 2023 warning on AI-generated influence operations targeting European democratic processes.
    • World Economic Forum identification of AI-assisted disinformation as a top 5 global risk (2024 Global Risk Report).
    • Stanford Internet Observatory and Atlantic Council Digital Forensic Research Lab documentation of coordinated inauthentic behaviour by state and non-state actors.
    • The Digital Society Surveillance literature connecting bot manipulation to state-level information operations — Russian, Chinese, and domestic political operators deploying coordinated inauthentic behaviour at scale to manipulate electoral and policy outcomes.

Current Landscape (2026)

  • By Q1 2026, the Death of the Internet dynamics have progressed from emerging trend to structural baseline. Key status markers:
  • Bot Traffic: Stabilised at approximately 50–55% of total internet traffic per most measurement methodologies, with composition shifting as LLM API-powered bots (GPTBot, ClaudeBot, Amazonbot, Perplexitybot) now constitute a significant proportion of the “legitimate” bot category — crawling the public web at scale to train and update AI Search and Large Language Models systems.
  • Synthetic Content: The most sophisticated AI content farm operators now produce content passing AI detection classifiers at rates above 90% through paraphrasing, human editing passes, and detection-evasion techniques. The Guardian’s “Zombie Internet” characterisation has been adopted in mainstream journalism, and the problem has reached sufficient visibility that it is a standard agenda item for platform trust and safety teams.
  • AI Search: Google AI Overviews are deployed across all major markets, with independent audits finding error rates of 3–8% on factual queries and citation failures in approximately 30–40% of responses that purport to cite sources. AI Search has effectively bifurcated search into a paid AI-tier and a degraded free tier, a structural change connecting to Competition in AI dynamics and Global Inequality in knowledge access.
  • Content Provenance: C2PA adoption has reached approximately 8% of new professional photo and video content (driven by Adobe Firefly mandatory signing, Getty Images implementation, and news agency consortium adoption) but remains below 1% of web text content. The two-tier internet is now effectively operational: tier-1 consists of C2PA-signed, attribution-tracked, paywall-accessed content from major publishers; tier-2 consists of the free public web increasingly indistinguishable from AI-generated slop.
  • Agent Mediation: Agents and Agentic Internet architectures are beginning to mediate significant proportions of online interaction — purchasing, scheduling, research, and customer service queries routed through AI agents rather than human browser sessions. As agents become primary web clients, the proportion of internet traffic attributable to non-human agents is projected to approach 80–90% by 2028–2030, completing the transformation of the web from a human-readable information environment to a machine-readable API layer.
  • Regulatory Implementation:
    • EU DSA Article 34 risk assessment requirements are generating the first wave of algorithmically-informed regulatory enforcement, with the European Commission investigating TikTok and Meta’s recommender system designs under the systemic risk framework.
    • UK Ofcom’s Online Safety Act implementation is proceeding cautiously, with industry lobbying delaying several provisions and the codes of practice for recommender systems expected in late 2026.
    • The US lacks equivalent comprehensive regulation, with FTC enforcement actions addressing specific harms (fake reviews, deceptive AI endorsements per the 2024 Fake Reviews Rule) rather than systemic platform architecture — a regulatory gap that has become a major transatlantic policy tension.
    • California AB 3211 (Content Provenance and Authenticity Act) represents the most advanced US state-level intervention, requiring C2PA disclosure for AI-generated content and creating a compliance baseline that major platforms are treating as de facto national standard.
  • Mental Health Consequences:
    • Jonathan Haidt’s The Anxious Generation (2024) catalysed a major policy debate on social media, AI companions, and algorithmic feed effects on adolescent mental health.
    • US Senate hearings (January 2024, featuring Meta/Snap/TikTok/Discord CEOs) resulted in proposed Kids Online Safety Act provisions with bipartisan support but no final legislation as of Q1 2026.
    • UK Department for Education guidance (January 2024) restricting mobile phone use in schools addresses the device-access pathway rather than the algorithmic architecture that drives harm.
    • The Surgeon General’s 2023 advisory on social media and youth mental health moved the US public health consensus toward active recommendation for limiting platform exposure for adolescents under 16.
    • The connection between algorithmic engagement-maximisation and adolescent anxiety, depression, and self-harm — measured across multiple longitudinal datasets by Przybylski (Oxford OII), Haidt (NYU Stern), and Jean Twenge (San Diego State) — is now sufficiently documented to have moved from academic controversy to public health consensus.

UK Context

  • Regulatory Framework: The UK Online Safety Act 2023 represents the primary regulatory intervention, with Ofcom’s implementation timeline running from Category 1 service transparency reports (January 2024) through illegal content duties (March 2024) to age assurance requirements (July 2025) and recommender system transparency (expected 2026). Ofcom’s approach focuses on harm outcomes rather than algorithmic process transparency, granting platforms implementation discretion. The Online Safety Act’s provisions on fake accounts and coordinated inauthentic behaviour (Part 3, Chapter 2) provide a statutory basis for platform enforcement that the US lacks. Ofcom’s Horizon Scanning programme has explicitly identified AI-generated synthetic content and bot manipulation as priority harms for the 2025–2026 regulatory cycle.
  • Academic Research: Imperial College London’s Centre for Responsible Technology leads UK work on platform accountability and algorithmic harm, with the Data Science Institute contributing technical analysis of bot detection and content provenance. The Oxford Internet Institute (OII), under Sonia Livingstone and Andrew Przybylski, conducts the most internationally cited empirical work on technology effects on wellbeing and democratic discourse, directly informing Ofcom’s evidence base. The Alan Turing Institute’s Public Policy Programme coordinates inter-institutional AI governance research, with nodes at Manchester (ATI Node / Manchester Digital Economy Research Centre), Edinburgh, and Cambridge. Cambridge University’s Centre for the Future of Intelligence (Stephen Cave, Kanta Dihal) examines the philosophical dimensions of AI’s effects on human self-understanding. Edinburgh’s Centre for Technomoral Futures (Shannon Vallor) produces internationally cited work on the moral architecture of AI systems and the ethics of attention manipulation. UCL’s AI Centre (£80M investment, 2023) has active programmes on algorithmic bias and information integrity with direct linkages to the DSIT AI Safety Institute.
  • Northern English Industrial Context: Manchester’s designation as the UK’s most AI-ready city (SAS Institute survey, 2025) and its role as the UK’s second-largest digital economy cluster position it as a key site for internet economy research and practice. Manchester Digital Economy Research Centre coordinates with Greater Manchester Combined Authority on digital inclusion and AI-mediated information access equity. Leeds’s £330M Microsoft AI partnership (2024) includes data centre infrastructure physically hosting AI inference and bot traffic affecting UK internet quality. Newcastle’s Northumbria Digital Society programme examines distributional effects of internet quality degradation on post-industrial communities, with specific focus on the correlation between geographic digital exclusion and susceptibility to AI-generated Deepfakes and fraudulent content. Sheffield Hallam’s Creative Industries unit tracks the collapse of independent digital publishing economics under AI content farm competition — directly relevant to the cluster of independent music, fashion, and arts publishers historically centred in South Yorkshire. The BBC R&D centre at MediaCityUK (Salford) conducts applied research on content provenance signals and authentic journalism markers, with implications for BBC News online and BBC Verify (the BBC’s fact-checking operation launched in 2022 in response to information integrity concerns).
  • Civil Society and Industry: The Open Rights Group is the primary UK civil society voice on platform accountability and Surveillance Capitalism, advocating for stronger algorithmic transparency in Online Safety Act implementation. Which? Consumer Association conducted a 2024 investigation finding 32% of product reviews on a major UK retailer’s site showed AI generation markers. JISC (Joint Information Systems Committee) maintains UK academic internet infrastructure and has produced research on AI content flooding effects on academic information retrieval quality. The Ada Lovelace Institute (London) publishes policy-focused research on platform accountability and recommender system governance, directly informing parliamentary scrutiny of the Online Safety Act. Full Fact, the UK’s primary independent fact-checking organisation, has documented the interaction between AI-generated misinformation and Legacy Media trust decline, with specific documentation of AI-generated news articles replicating as secondary sources across the UK web.
  • Decentralised Web UK Presence: Berners-Lee’s Inrupt (founded 2018, Cambridge UK / Boston US) commercialises the Solid protocol for enterprise data pod deployment. Southampton University (Nigel Shadbolt, Wendy Hall) retains the original Solid research programme. The Flemish Government’s significant Solid investment (€20M+ across 2022–2025) has made EU-funded Solid deployment the most advanced outside of US enterprise contexts, with UK-Belgium research collaboration through the ODI (Open Data Institute, founded by Berners-Lee and Shadbolt) on Decentralised Web data governance.

Future Directions (2026-2030)

  • Content Provenance Infrastructure Scaling:
    • C2PA adoption expected to reach 25-40% of new professional media content by 2028, driven by EU AI Act Article 50 (effective August 2026) and California AB 3211 compliance requirements.
    • Critical unresolved question: can provenance infrastructure scale to user-generated content at 500M+ posts/day scale of Social Media, or will it deepen the two-tier internet?
    • Browser-native C2PA display (Chrome and Firefox Content Credentials extensions) could normalise provenance checking as standard user behaviour.
    • Extension of C2PA to text content is the key gap for addressing AI content farm proliferation at scale.
  • Autonomous Agent Internet Mediation:
    • As Agents and Agentic Internet architectures become primary web clients, non-human traffic will approach 80-90% of total internet traffic by 2028-2030.
    • Qualitative transformation: the web shifts from a human-readable information environment to a machine-readable API layer with human access as a secondary use case.
    • Policy implications — agent identity, agent liability, agent manipulation, economics of agent-mediated access — are largely unaddressed by current regulatory frameworks.
    • The interaction with Global Inequality is acute: agent-mediated access to high-quality information requires LLM API subscription costs inaccessible to the majority of global internet users.
    • The AISP 5.1 framework and emerging agent identity standards (W3C DIDs, Verifiable Credentials) represent nascent infrastructure for agent accountability.
  • Model Collapse Mitigation:
    • Training data watermarking at generation stage — enabling filtering of AI-generated content from future training corpora (Google SynthID, Meta Stable Signature).
    • Human feedback amplification — maintaining annotation team scale proportional to synthetic content volume.
    • Retrieval-augmented generation — grounding Large-Scale Pretrained Foundation Model outputs in retrieved authentic documents rather than purely parametric knowledge.
    • Diversity-preserving training objectives — explicitly penalising distribution narrowing during fine-tuning to maintain tail distribution coverage.
    • Shumailov et al. finding: even 10-20% authentic human data significantly mitigates collapse dynamics, making Background Tokens preservation tractable if acted on before the window closes.
  • Decentralised Web Renaissance:
    • Solid’s data pod model, ActivityPub federation, and IPFS/Filecoin content-addressed storage represent the primary architectural alternatives to enshittifying centralised platforms.
    • Fediverse scale: 10M+ Mastodon users plus 20M+ Bluesky AT Protocol users (2024) — the most significant tests of non-advertising-funded social infrastructure.
    • UK investment: Inrupt (Cambridge), Southampton University Solid research programme, ODI data governance collaboration.
    • Payment channel integration — Lightning Network micropayments, CBDC Frameworks-based micro-licensing — could replace advertising as the economic model for content creation.
  • Epistemic Infrastructure as Public Policy:
    • Authentic content corpora, content provenance standards, and bot detection infrastructure framed as public epistemic goods — analogous to libraries and public broadcasting.
    • UK precedents: Internet Archive preservation mandate, JISC academic internet infrastructure, BBC public service information obligations.
    • Policy direction: C2PA provenance certification, open-access authentic training corpora, and state-supported Decentralised Web infrastructure as public investment.
    • UNESCO Recommendation on the Ethics of AI (2021, 193 member states) and subsequent national AI strategies increasingly incorporate epistemic infrastructure as a public good.

Risk and Limitations of Counter-Measures

  • The primary counter-measures against internet degradation each carry their own risks and limitations that constrain their effectiveness:
  • C2PA / Content Provenance Limitations:
    • Provenance signing only works for content created after the signing infrastructure is deployed — the existing corpus of 40B+ authentic human web pages has no cryptographic attestation and is indistinguishable from AI-generated content by signature alone.
    • C2PA metadata can be stripped by any downstream processing step that does not preserve metadata chains (image compression, screenshot, re-upload), creating large gaps in provenance coverage even for originally signed content.
    • The “liar’s paradox” problem: a sophisticated AI system can generate C2PA-signed content if given access to a signing certificate, meaning provenance signing attests to the identity of the signer but not to the authenticity of the content.
    • Adoption is driven by major platforms and publishers (Adobe, Getty, AP, Reuters) but remains below 1% of web text content as of Q1 2026, creating an asymmetric provenance landscape that deepens the two-tier internet.
    • Small and independent publishers — including the Legacy Media outlets in specialist journalism and local news that have the strongest authentic voice claims — lack the technical capacity and economic incentive to implement C2PA signing workflows.
  • AI Detection Limitations:
    • Sadasivan et al. (arXiv:2303.11156, 2023) proved theoretically that no AI text detection scheme can reliably distinguish AI from human text when the AI model is sufficiently capable and when the detector’s training distribution differs from the production distribution. This applies to all statistical detection approaches.
    • Watermarking approaches (adding statistical signatures to AI-generated text during inference) are defeated by paraphrasing, translation, and moderate editing — exactly the post-processing used by AI content farms to evade detection.
    • False positive rates for AI detection tools remain in the 10–20% range for non-native English human writing, creating significant equity and credibility problems if detection is used for high-stakes decisions (academic integrity, journalism verification, legal evidence).
    • Detection becomes an arms race in which each improvement in detection capability is met by improvements in generation capability that evade it, with no stable equilibrium.
  • Decentralised Web Adoption Limitations:
    • Solid and ActivityPub-based Federated Social Networks require users to manage their own identity and data pods, creating a technical barrier that limits adoption to technically sophisticated users. The majority of internet users lack the capacity or motivation to self-host or manage decentralised infrastructure.
    • Network effects strongly favour incumbent centralised platforms: a Mastodon instance with 10,000 users cannot replicate the discovery, community, and content diversity of a Twitter with 300M users, even if Mastodon’s architecture is technically superior.
    • Spam, moderation, and content quality problems are harder to address in federated architectures where no central authority can enforce community standards, creating a different kind of content quality problem in the decentralised alternative.
    • The economic model for Decentralised Web infrastructure is unresolved: without advertising revenue or enterprise subscription fees, the costs of running federated infrastructure fall on volunteer administrators and hobbyist server operators, creating sustainability challenges.
  • Regulatory Limitations:
    • The UK Online Safety Act, EU DSA, and similar regulations apply to platforms operating in their jurisdictions but cannot address bot infrastructure, AI content farms, and synthetic content operations hosted in jurisdictions outside their reach.
    • Regulatory compliance costs are regressive: large platforms can absorb compliance overhead while small platforms and new entrants face proportionally higher costs, potentially concentrating the market further in the hands of the incumbents most responsible for Platform Decay.
    • Enforcement lags: the time from harm identification to regulatory response to enforcement action typically spans 3–7 years, during which the harm compounds. The Bot Report 2024 data was available; the regulatory response remains incomplete.
    • Jurisdictional fragmentation: divergent regulatory regimes in the US, EU, UK, and Asia create compliance complexity that platforms can exploit by routing activities through permissive jurisdictions, and that creates inconsistent user protection depending on geographic location.
  • Background Tokens Depletion:
    • The Background Tokens problem is structurally irreversible: once the open web is dominated by AI-generated content, there is no mechanism to recover the authentic human corpus that has been diluted. The Internet Archive preserves historical snapshots but cannot regenerate the ongoing production of authentic human discourse that AI systems need to maintain quality.
    • The Data Provenance Initiative’s finding that human-authored text is declining at 8–12 percentage points per year suggests that the window for intervention — before the open training corpus is irreversibly degraded — is measured in years rather than decades.
    • The Global Inequality dimension means that the depletion of authentic content in minority languages and non-Western cultural contexts proceeds faster than in English, as AI content farms target high-traffic English keywords and leave minority-language authentic content relatively undisturbed — creating a perverse situation in which English-language AI training data degrades faster than minority-language equivalents.

Conceptual Relationships to Adjacent Ontology Nodes

  • The Death of the Internet cluster has rich structured relationships to adjacent concepts in the ontology that deserve explicit mapping:
  • Relationship to Agents and Agentic Internet: The agent-mediated internet represents the endpoint state of the Dead Internet Theory dynamic — not bots mimicking human activity but autonomous AI agents conducting legitimate commercial and informational transactions on behalf of human principals. As Agents become primary web clients, the proportion of authentic human-originated internet activity asymptotically approaches zero. The Agentic Internet page covers the infrastructure and protocol dimensions; this page covers the epistemic and ethical consequences.
  • Relationship to AI Scrapers: AI training data scrapers (GPTBot, ClaudeBot, CCBot, Common Crawl) are simultaneously a symptom and a cause of the Death of the Internet dynamic. They are a symptom in that their proliferation reflects the scale of AI training data demand; they are a cause in that their harvesting of the authentic web corpus accelerates its depletion. The AI Scrapers page covers the technical and legal dimensions of scraping; this page covers the systemic epistemic consequences.
  • Relationship to Large Language Models and Large-Scale Pretrained Foundation Model: The Habsburg AI collapse dynamic directly threatens the quality trajectory of Large Language Models and Large-Scale Pretrained Foundation Model. As the training corpus degrades, model quality degrades in a manner that is self-reinforcing but potentially detectable through diversity metrics. The Large-Scale Pretrained Foundation Model page covers technical architecture; this page covers the training data provenance dimension.
  • Relationship to Bias in Large Language Models: The right-wing media scraping asymmetry and the differential rate of authentic content depletion across languages and cultural contexts both introduce systematic biases into Large-Scale Pretrained Foundation Model training distributions. The Bias in Large Language Models page covers detection and mitigation; this page covers one of the upstream generative causes.
  • Relationship to AI companions: The extension of AI-mediated interaction to intimate relationship contexts — dating platform assistance, companion AI applications — represents the most intimate expression of the Dead Internet dynamic. AI companions covers the product and safety dimensions; this page contextualises them within the broader synthetic content saturation problem.
  • Relationship to Digital Society Surveillance: The bot infrastructure underpinning the Dead Internet Theory serves not only commercial but surveillance and political manipulation purposes. Coordinated inauthentic behaviour by state actors (documented in Europol, Stanford Internet Observatory, and Atlantic Council Digital Forensic Research Lab reports) uses the same technical stack as commercial bot operations. Digital Society Surveillance covers the political surveillance dimension; this page covers the information integrity consequences.
  • Relationship to Decentralised Web and Solid: These represent the technical architecture of the primary structural alternative to the centralised, enshittifying platform internet. Decentralised Web and Solid cover the protocol and implementation dimensions; this page provides the epistemic and political-economic motivation for their adoption.
  • Relationship to Global Inequality: The two-tier internet — C2PA-signed, paywall-gated, high-quality content for paying users in wealthy countries; AI-generated slop on the free web for the remaining 3.5B internet users — directly amplifies existing Global Inequality in knowledge access, civic participation, and economic opportunity.

Research and Literature

  • Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. (2024). “The Curse of Recursion: Training on Generated Data Makes Models Forget.” Nature, 628, pp. 755–762.
  • Doctorow, C. (2023). “Tiktok’s enshittification.” Pluralistic.net, 23 January 2023. Reprinted and expanded, Locus Magazine, March 2023.
  • Doctorow, C. (2023). “‘Enshittification’ is coming for absolutely everything.” Financial Times, 11 November 2023.
  • Imperva (2024). Bad Bot Report 2024: The Imperative of Bot Management. San Mateo: Imperva. Documents 49.6% non-human traffic proportion and 32% malicious bot proportion.
  • NewsGuard (2024). AI-Generated News and Information Websites Tracker. NewsGuard Technologies, Updated May 2024. Documents 1,265+ AI news sites, 25-fold increase from April 2023.
  • Originality.ai (2024). State of AI Content on the Web: 2024 Annual Report. Documents 52–68% AI markers in low-authority new content.
  • Appleton, M. (2023). “The Expanding Dark Forest and Generative AI.” maggieappleton.com, January 2023.
  • Strickler, Y. (2019). “The Dark Forest Theory of the Internet.” Medium, 17 March 2019.
  • Busta, C. (2020). “The Internet Didn’t Kill Counterculture — You Just Won’t Find It on Instagram.” New Models, Issue 13.
  • Zuboff, S. (2019). The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power. London: Profile Books.
  • Thompson, B. (2025). “The Agentic Web and Original Sin.” Stratechery, 14 February 2025.
  • Pariser, E. (2011). The Filter Bubble: What the Internet is Hiding from You. New York: Penguin Press.
  • Fishkin, R. (2024). “An Anonymous Source Shared Thousands of Leaked Google Search API Documents with Me.” SparkToro Blog, 27 May 2024.
  • Cloudflare (2024). 2024 State of Application Security Report. San Francisco: Cloudflare Inc. Documents automated traffic exceeding human traffic in 14/17 industry sectors.
  • Luccioni, A., Jernite, Y., and Strubell, E. (2024). “Power Hungry Processing: Watts Driving the Cost of AI Deployment?” arXiv:2311.16863. Documents AI search energy consumption at 10x traditional search per query.
  • Sadasivan, V.S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S. (2023). “Can AI-Generated Text be Reliably Detected?” arXiv:2303.11156. Theoretical impossibility of reliable AI text detection.
  • Akerlof, G.A. (1970). “The Market for ‘Lemons’: Quality Uncertainty and the Market Selection.” Quarterly Journal of Economics, 84(3), pp. 488–500. Theoretical basis for authentic content market failure under information asymmetry.
  • Gao, L. et al. (2020). “The Pile: An 800GB Dataset of Diverse Text for Language Modeling.” arXiv:2101.00027.
  • Dodge, J. et al. (2021). “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.” EMNLP 2021.
  • Haidt, J. (2024). The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness. New York: Penguin Press.
  • Mollick, E. (2024). “Two Weird Things That Are Going to Happen.” LinkedIn, 18 April 2024. Bots-to-bots interaction projection.
  • Wu, T. (2016). The Attention Merchants: The Epic Scramble to Get Inside Our Heads. New York: Knopf.
  • Coalition for Content Provenance and Authenticity (C2PA) (2023). C2PA Technical Specification 1.3. Primary content provenance standard.
  • Berners-Lee, T. (2022). Solid Technical Report. MIT CSAIL / W3C. Decentralised web alternative architecture.
  • Ofcom (2023). Online Safety Act 2023: Implementation Framework and Codes of Practice. London: Ofcom.
  • Epoch AI / Data Provenance Initiative (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. Documents uncontaminated human text falling to approximately 40% of public web.

Green Shoots: Authentic Content Revival Mechanisms

  • Despite the structural pressures toward internet degradation, several mechanisms represent genuine potential for reversal or mitigation:
  • The Fediverse and ActivityPub Ecosystem:
    • Mastodon has grown to 10M+ registered users and approximately 1.5M monthly active users by 2024, demonstrating that non-enshittified Federated Social Networks can achieve meaningful scale.
    • ActivityPub’s W3C Recommendation status (2018) means that Decentralised Web social infrastructure now has a stable, widely-implemented protocol foundation.
    • Bluesky’s AT Protocol (2024 open source release) introduces an alternative federated architecture with better scalability properties than ActivityPub, attracting 20M+ registered users by early 2024 and representing a potential scale threshold for federated social alternatives.
    • The presence of high-signal communities on Mastodon instances (particularly in academic, security research, and journalism domains) demonstrates that authentic discourse can survive and thrive in non-advertising-funded spaces.
  • The IndieWeb and Personal Web Revival:
    • The “back to blogs” movement documented in Ed Zitron’s “Where Have All the Websites Gone?” (2024) and similar analyses reflects user-driven retreat to personal websites, RSS feeds, and newsletters as alternatives to algorithm-mediated platforms.
    • Substack’s growth to 35M+ free subscribers and significant paid subscriber counts by 2024 demonstrates that subscription-funded text content can achieve economic viability outside the advertising model.
    • Ghost, WordPress, and other self-hosted publishing platforms retain large user bases, representing a substantial corpus of authentic, non-algorithm-optimised human writing.
    • The revival of personal websites, link blogs, and Solid data pods among technically sophisticated users represents a qualitative shift in how authentically internet-native people relate to the web.
  • AI as Epistemic Defence:
    • LLM agents given access to verified search can achieve superhuman rating performance on fact-checking tasks (arXiv:2403.18802, 2024), running at 20x lower cost than human fact-checkers.
    • Larger, more capable Large-Scale Pretrained Foundation Model are systematically more factual than smaller models, suggesting that the quality trajectory of AI systems (if fed clean training data) is toward better epistemic performance.
    • AI-powered disinformation detection tools (NewsGuard’s AI content tracker, Originality.ai’s classifier, Google’s SynthID watermarking) provide scalable mechanisms for identifying synthetic content in ways that human-only review cannot match.
    • The framework of “signed, attested, timestamped content” (Nic Carter, 2024) as the post-AI epistemic standard — where unsigned content defaults to unreliable and only cryptographically attested content carries epistemic weight — represents a coherent architectural path to a post-enshittification information environment.
  • Regulatory Momentum:
    • The EU AI Act Article 50 transparency requirements (effective August 2026) will mandate disclosure of AI-generated content at the largest scale yet, creating a legal baseline for provenance disclosure.
    • The EU DSA Article 34 risk assessment obligations are generating the first wave of algorithmic accountability enforcement, with real fines and structural remedies.
    • The UK Online Safety Act’s age assurance requirements (July 2025) are driving platform identity verification infrastructure that could support broader content provenance systems.
    • State-level legislation (California AB 3211) is establishing provenance requirements that could propagate nationally through platform compliance decisions.
  • Cryptographic End Points as Forcing Function:
    • The convergence of cryptographic identity (Distributed Identity, Digital Signature), content provenance (C2PA Standard), and Blockchain Attestation creates a technological pathway to an internet where content authenticity is structurally verifiable rather than probabilistically estimated.
    • Decentralised identity standards (DID W3C Recommendation 2022; Verifiable Credentials W3C Recommendation 2022) provide the identity layer for a provenance-secured internet.
    • The Solid protocol’s integration of WebID-based identity with data pod access control represents a deployed implementation of the authenticated-endpoints model at meaningful scale in enterprise and government contexts.
    • Payment channel integration (Lightning Network micropayments for content access; CBDC-based micro-licensing) could replace advertising as the economic model for content creation, removing the advertising-incentive driver of enshittification.

Metadata

Provenance

  • Shumailov et al. “The Curse of Recursion” Nature 2024 (model collapse / Habsburg AI)
  • Doctorow “Enshittification” Pluralistic 2023 (platform decay lifecycle)
  • Doctorow / Financial Times “‘Enshittification’ is coming for absolutely everything” 2023
  • Imperva Bad Bot Report 2024 (49.6% non-human traffic; 32% malicious bots)
  • NewsGuard AI News Tracker 2024 (1,265+ AI news sites; 25-fold increase in 13 months)
  • Originality.ai State of AI Content 2024 (52-68% AI markers in low-authority new content)
  • Appleton “The Expanding Dark Forest and Generative AI” maggieappleton.com 2023
  • Strickler “Dark Forest Theory of the Internet” Medium 2019
  • Busta “The Internet Didn’t Kill Counterculture” New Models 2020
  • Zuboff The Age of Surveillance Capitalism Profile Books 2019
  • Thompson “The Agentic Web and Original Sin” Stratechery 2025
  • Fishkin / SparkToro Google Search API leak analysis 2024
  • Cloudflare 2024 State of Application Security Report
  • Epoch AI / Data Provenance Initiative “Consent in Crisis” 2024
  • Akerlof “Market for Lemons” 1970
  • Sadasivan et al. “Can AI-Generated Text be Reliably Detected?” arXiv 2023
  • Luccioni et al. “Power Hungry Processing” arXiv 2024
  • Ofcom Online Safety Act implementation framework 2023
  • C2PA Technical Specification 1.3 (2023)
  • Haidt The Anxious Generation 2024
  • Pariser The Filter Bubble 2011
  • Wu The Attention Merchants 2016
  • Gao et al. “The Pile” arXiv 2020
  • Dodge et al. C4 corpus documentation EMNLP 2021
  • Berners-Lee Solid Technical Report MIT CSAIL 2022
  • Mollick on bots-to-bots interaction LinkedIn 2024
  • Wikipedia Dead Internet Theory entry (documents trajectory from 2021 fringe to mainstream academic interest)
  • WIRED “Most News Sites Block AI Bots; Right-Wing Media Welcomes Them” 2023 (scraping asymmetry)
  • WIRED “Death of Truth: Misinformation Advertising” (programmatic advertising and disinformation economics)
  • The Atlantic “Nobody Knows What’s Happening Online Anymore” December 2023
  • The Atlantic “Welcome to Geriatric Social Media” (platform demographic dynamics)
  • The Atlantic “Social Media Broke Up With News. So Did Readers.” November 2023
  • The Atlantic “The Toilet Theory of the Internet” May 2024
  • Gizmodo “Google Search Is Now a Giant Hallucination” 2024
  • Ben Landau-Taylor “The DDoS Attack of Academic Bullshit” (AI flooding academic publishing)
  • BioRxiv “Fraudulent studies undermining reliability of systematic reviews” February 2024
  • Stratechery “The Agentic Web and Original Sin” February 2025
  • Mashable “The majority of traffic from Elon Musk’s X may have been fake during the Super Bowl” 2024
  • GitHub Sebby37/Dead-Internet (empirical documentation project)
  • osf.io/xcwdn “Using AI to Reduce Conspiracy Theory Beliefs” (counterpoint: AI as epistemic tool)
  • Adweek “Study: 15 Million Russian Instagram Influencers’ Followers Are Bots” 2022
  • Europol warning on AI in online content (2023)
  • W3C SN-dystopia diagram (Berners-Lee on social network dystopia trajectories)
  • Shyam Sankar “Technology is the Problem” (technology-as-root-cause framing)
  • domain-correction: infrastructure → ethics-society