Trust and Safety (T&S) is the professional discipline, operational infrastructure, and regulatory framework that governs the detection, review, enforcement, and remediation of harmful content and abusive behaviour across digital platforms, encompassing the full spectrum from child sexual abuse ma…

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:ContentModerationSystem))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:CSAMDetectionPipeline))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:PlatformIntegrityOperations))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:PolicyEnforcementFramework))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:HumanReviewPipeline))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:AutomatedClassifier))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:IncidentResponseSystem))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:hasPart es:HashMatchingDatabase))

## Dependency Relationships
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:requires es:PerceptualHashingTechnology))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:requires es:NLPClassifierModel))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:requires es:HumanAnnotationOracle))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:requires es:PolicyGoverningDocument))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:requires es:LegalComplianceFramework))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:dependsOn es:FoundationModelInference))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:dependsOn es:CrossPlatformSignalSharing))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:dependsOn es:ModeratorWellbeingInfrastructure))

## Capability Relationships
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:enables es:HarmfulContentRemoval))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:enables es:OnlineSafetyCompliance))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:enables es:ElectionIntegrityProtection))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:enables es:ChildSafetyOnline))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:supports es:PlatformGovernance))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:supports es:DisinformationCountermeasures))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:supports es:FraudPreventionOperations))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:supports es:TerroristContentRemoval))

## Implementation Relationships
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:UKOnlineSafetyAct2023))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:EUDigitalServicesAct))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:NCMECCyberTiplineProtocol))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:GIFCTHashSharingStandard))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:TechCoalitionLanternProtocol))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:implements es:ROOSTDIREFramework))

## Reduction Relationships
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:reducesTo es:ContentModerationOperations))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:reducesTo es:HashMatchingAndSignalSharing))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:reducesTo es:PolicyEnforcementDecision))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:exemplifiedBy es:CSAMDetectionPipeline))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:exemplifiedBy es:OfcomOnlineSafetyActCompliance))
SubClassOf(es:TrustAndSafety
  ObjectSomeValuesFrom(es:disjointWith es:UnmoderatedFreeSpeechPlatform))

About Trust and Safety

  • Trust and Safety emerged as a formal professional discipline in the mid-2010s as user-generated content platforms scaled to billions of users and the volume, velocity, and virality of harmful content exceeded human-only review capacity.
  • The field synthesises multiple disciplines:
    • Computer science: machine learning classifiers, perceptual hash matching, graph analysis for coordinated inauthentic behaviour, NLP for hate speech and misinformation detection
    • Behavioural psychology: understanding harm escalation trajectories, radicalisation pathways, grooming patterns, and moderator trauma mechanisms
    • Law: jurisdictional compliance, evidence preservation, law enforcement cooperation frameworks, Section 230 / DSA / OSA liability architectures
    • Organisational management: outsourced moderation supply chain ethics, tier-1/tier-2/tier-3 escalation routing, quality assurance sampling methodologies
    • Policy analysis: community standards drafting, appeals system design, transparency reporting frameworks, regulatory engagement
  • The Trust & Safety Professional Association (TSPA), established in 2020 as a 501(c)(6) non-partisan body, provides the primary professional community infrastructure.
    • Structured curriculum spans: content moderation operations, legal and policy frameworks, product safety engineering, researcher relations
    • 400+ new members added in first seven months of 2024 from 55 countries
    • Events in 2024: first India community event in Hyderabad; SXSW Austin meetup; 300-person EMEA Summit Dublin
  • The Atlantic Council 2023 landscape report identified structural tensions governing the field:
    • Commercial incentives systematically underweight safety investment
    • Measurement of T&S value remains methodologically contested
    • Practitioner trauma is systemic yet under-resourced
    • Platforms serve as final arbiters of permissible global speech with near-zero external accountability
  • The field confronts a principal-agent paradox: private corporations, incentivised by engagement metrics, bear the cost of safety enforcement that primarily benefits society rather than shareholders.
  • As the Atlantic Council report noted, T&S must balance competing goals:
    • Protecting rights vs mitigating harms
    • Efficiency vs accuracy in enforcement decisions
    • Reviewer wellbeing vs review throughput requirements
    • Centralised vs decentralised moderation architectures
    • Growth investment vs safety investment
    • Reactive enforcement vs proactive platform design

Components / Architecture

  • Trust and Safety systems deploy a layered detection-review-enforcement architecture, typically implemented across four logical tiers.

Tier 1: Automated Detection

  • PhotoDNA Perceptual Hash Matching: Microsoft’s PhotoDNA, developed 2009 with Dartmouth College and deployed through NCMEC and 200+ platforms, generates robust 1152-bit perceptual signatures.
    • Survives image transformations: resizing, compression, colour shift, mild cropping
    • DCT-based processing after grayscale conversion and grid decomposition
    • Hash collision risk at petabyte scale requires tunable similarity thresholds trading false-positive rate against recall
    • Deployed in Microsoft Azure Content Safety API; integrated into most major cloud storage providers
  • Video Hash Matching: Thorn’s Scene-Sensitive Video Hashing (SSVH) was the most adopted perceptual video hashing tool among Tech Coalition members in 2024.
    • NCMEC’s Take It Down repository adopted by ten Tech Coalition members in 2024
    • Tech Coalition Video Hash Interoperability Alpha Project (launched March 2022 with NCMEC, Meta, Google) has rehashed 220,000+ known CSAM videos; enabled detection of 750,000+ CSAM videos
    • GIFCT HSDB Year 4 Review (2025) identified interoperability challenges between legacy SHA-1 and modern PDQ/TMK formats
  • NLP and Vision Classifiers: The primary tools for non-CSAM policy violation detection.
    • Hive Moderation: 40+ violation classes; Moderation 11B VLM fine-tuned on Llama 3.2 11B Vision; API-based deployment
    • Sightengine: 110+ moderation classes; 99.2% F1 on porn detection; integrated via API
    • Platform-proprietary transformers: all major platforms train internal classifiers on labelled human review data
    • LLM-based moderation: Xu et al. arXiv 2024 demonstrates superior accuracy vs traditional NLP; Springer 2025 review confirms LLM outperformance with interpretability tradeoff
    • Content moderation API market valued 2,687 million by 2032 at 10.5% CAGR
  • Novel CSAM AI Classifiers: Hash matching is insufficient for previously-unseen material.
    • Thorn’s Safer platform combines hash matching with ML classifiers; purpose-built for novel content detection
    • Critical gap: known-hash approaches cover only previously catalogued material, missing newly-generated content

Tier 2: Investigation and Incident Response

  • Cross-Platform Signal Sharing: Tech Coalition’s Project Lantern enables sharing of CSAM hash values and incident signals across participating platforms.
    • 300,000 hashes shared as of December 2023 (40% of program signal volume)
    • One Lantern-enabled investigation: Meta removed 10,000+ violating Facebook Profiles/Pages/Instagram accounts following URL signals from MEGA
    • First Lantern Transparency Report published by Tech Coalition documenting program impact
  • GIFCT Hash-Sharing Database: Dedicated to terrorist and violent extremist content (TVEC).
    • Cross-platform coordination between major platforms and national security agencies
    • Year 4 working group (2024-2025) examining technical challenges with legacy formats
  • ROOST Osprey: Open-source investigation and incident response tool donated by Discord to ROOST, released July 2025.
    • Lightweight user-friendly design for platforms of all sizes
    • Bluesky adopting as inaugural partner, demonstrating applicability beyond large incumbents
    • Part of ROOST DIRE framework (Detection, Investigation, Review, Enforcement)
  • ROOST Coop: Signal-sharing platform based on Cove IP acquired by ROOST.
    • Open-sourced alongside Osprey July 2025
    • Funded by Google, OpenAI, Discord, Roblox ($27M total)
    • Acquired through philanthropic funding and Perkins Coie pro bono legal services
  • Internal Platform Tools: Meta’s graph analysis for Coordinated Inauthentic Behaviour (CIB) detection.
    • 200+ CIB networks removed since 2017 across elections in dozens of countries
    • Network-level signals: coordinated posting patterns, account age distributions, device fingerprint clusters

Tier 3: Human Review Pipeline

  • The global content moderation workforce is estimated at 100,000+ contract workers.
  • Primary geographic concentration: Philippines, Kenya, Colombia, India.
  • Documented conditions:
    • Poverty wages relative to trauma burden and psychological risk
    • Exposure to graphic CSAM, beheadings, self-harm, and violent content repeatedly throughout shifts
    • Suppressed union organising (documented by Time, The Bureau of Investigative Journalism, CNN)
    • 81% of moderators report insufficient employer mental health support
  • Kenya PTSD Litigation (2024): 140+ former Facebook moderators at Sama/Samasource Nairobi facility sued Meta and Samasource Kenya alleging severe psychological trauma including PTSD, anxiety, and depression.
  • Labour rights mobilisation: UNI Global Union supported launch of Global Trade Union Alliance of Content Moderators in Nairobi.
    • Unions from nine countries ratified mental health protocols
    • Protocol focuses on: content exposure limits, mandatory psychological support, right to organise, living wages
  • AI automation and human review balance: Platforms are accelerating automation to reduce human review exposure.
    • TikTok: 96%+ of violating content removed through automated ML before any views (2024)
    • Over 80% of violating videos removed by automated technology
    • Human review concentrated on highest-severity, ambiguous, and appeals queues

Tier 4: Policy Enforcement and Governance

  • Platform Policy Documents: Community Standards (Meta), Community Guidelines (TikTok/YouTube), Rules (X) define permissible behaviour.
    • Vary substantially in specificity, appeals mechanisms, and transparency
    • Enforcement actions range from content removal through demotion/interstitials to account termination
  • Appeals and Oversight: Meta Oversight Board (established 2020) provides external review of high-profile enforcement decisions.
    • Has overturned multiple Meta content decisions on appeal
    • Jurisdictional limitations: narrow mandate, decisions non-binding on non-Meta platforms
  • Regulatory Accountability: Ofcom OSA enforcement framework (fines up to 10% global revenue / £18M); EU DSA Commission enforcement (fines up to 6% global annual turnover).

Use Cases / Major Families

CSAM Detection and Reporting

  • The most technically and legally critical T&S function, with mandatory reporting requirements in most major jurisdictions.
  • NCMEC CyberTipline: 36.2 million reports received from electronic service providers in 2023.
  • IWF (Internet Watch Foundation): Cambridge-based charity, founding INHOPE member.
    • 1.2 million+ webpages with CSAM removed in five years to 2024
    • Annual Data and Insights Report 2024: documented record removal volumes
    • Image Intercept tool launched 2025: free tool enabling smaller platforms to block known illegal content
  • PhotoDNA deployment at scale: Scribd’s documented PhotoDNA deployment (2026) illustrates democratisation of CSAM detection infrastructure to mid-size platforms.
  • Tech Coalition VHIP: Video Hash Interoperability Project rehashed 220,000+ CSAM videos; detected 750,000+ videos.

Terrorist and Violent Extremist Content (TVEC)

  • GIFCT Hash-Sharing Database coordinates cross-platform TVEC removal.
  • EU DSA VLOPs required to assess and mitigate systemic risks including TVEC dissemination.
  • GIFCT Year 4 Review (2024-2025) examining: PDQ hash adoption challenges, HSDB legacy format migration, database governance improvements.

Disinformation and Election Integrity

  • 2024 global election supercycle (50+ elections) drove platform-level election integrity commitments.
  • Major platforms pledged AI-generated content detection, watermarking, and authoritative voting information redirects for election queries.
  • Southport Riots (August 2024): Russian-linked disinformation falsely claimed stabbing attacker was Muslim asylum seeker; spread via Facebook, Telegram, WhatsApp; led to mosque attacks, asylum-seeker accommodation violence, 400+ arrests.
    • Demonstrated real-world lethality of T&S failures at scale
    • Catalysed UK government pressure for faster Ofcom OSA enforcement
  • Stanford Internet Observatory: Despite leadership attrition and political pressure in 2024 (Alex Stamos departure November 2023, Renée DiResta contract not renewed June 2024), maintained Journal of Online Trust & Safety and Trust & Safety Research Conference.

Platform Integrity (Fraud, Spam, Coordinated Inauthentic Behaviour)

  • Graph-based CIB detection: network topology analysis identifying coordinated account clusters, automated posting patterns, follower-farm networks.
  • Velocity anomaly detection for spam: unusual posting rates, reply volumes, DM patterns.
  • Account integrity scoring at authentication time: device fingerprint, IP reputation, behavioural biometrics.
  • Fraud prevention in fintech-adjacent contexts requires integration with AML KYC Compliance frameworks.

AI-Generated Harm (2024-2026 Emergent Category)

  • Synthetic CSAM: illegal in most jurisdictions under existing laws; challenges hash-based detection (novel by definition).
  • Deepfake NCII (non-consensual intimate images): AI-generated sexual imagery of real people; targeted harm category in OSA and DSA.
  • AI-generated electoral disinformation: deepfake robocalls, AI-written influence operation text, synthetic video of candidates.
  • Voice-cloned fraud: real-time voice synthesis for social engineering and financial fraud.
  • Agentic misuse: autonomous AI systems abusing platform APIs, conducting coordinated harassment, gaming recommendation algorithms.
  • Anthropic’s former Head of Safeguards Vinay Rao joined ROOST specifically to address AI-era T&S tooling requirements.
  • C2PA (Coalition for Content Provenance and Authenticity) watermarking: primary technical roadmap for AI content provenance in T&S pipelines.

Moderator Wellbeing and Labour Rights

  • Emerging as its own T&S sub-domain as litigation, regulation, and union pressure reshape supply chain obligations.
  • EU Platform Work Directive: creating legal pressure on platforms to restructure outsourced moderation contracts.
  • Global Trade Union Alliance of Content Moderators: nine-country protocol covering exposure limits, psychological support, labour rights.
  • Time investigative series on Facebook Africa content moderation: long-form documentation of systemic welfare failures.
  • The LCFI (Leverhulme Centre for the Future of Intelligence, Cambridge) coined “emotional labour offsetting” to describe AI platforms externalising moderation trauma to vulnerable Global South workers.

Academic Context

  • T&S research is institutionally distributed across computer science (HCI, NLP, computer vision), political science, law, psychology, and sociology.
  • Key publication venues:
    • ACM Conference on Fairness, Accountability, and Transparency (FAccT)
    • Web Science Conference and ACM WebSci
    • ACM Internet Measurement Conference (IMC)
    • Journal of Online Trust & Safety (Stanford Internet Observatory)
    • IEEE Symposium on Security and Privacy (IEEE S&P)
    • ACM CSCW (Computer-Supported Cooperative Work)
  • Foundational theoretical accounts:
    • Nicolas Suzor’s Lawless: The Secret Rules That Govern Our Digital Lives (2019, Cambridge University Press): legitimacy analysis of platform private rule-making
    • Kate Klonick’s “The New Governors” (2017, Harvard Law Review 131:1598): platforms as constitutional actors
    • Sarah T. Roberts’ Behind the Screen (2019, Yale University Press): ethnography of commercial moderation labour
    • Danielle Keats Citron’s The Fight for Privacy (2022): cyber harassment, sexual privacy, and platform liability frameworks
  • Active research fault lines:
    • Measurement of T&S intervention effectiveness: few causal designs exist; no counterfactual access to platform data
    • Interpretability of commercial classifiers used in enforcement decisions
    • Due process implications of automated account terminations at scale
    • “Legitimacy deficit” of private platforms enforcing public-speech rules without democratic mandate
    • Moderator trauma: longitudinal health impacts, employer duty of care, automation as harm-reduction mechanism
    • Algorithmic bias in moderation systems: documented disparate removal rates affecting Black, transgender, and conservative users (Haimson et al. 2021, ACM CSCW)
  • UK academic centres:
    • Oxford Internet Institute: platform governance, disinformation, digital wellbeing
    • Cambridge Internet Institute: online harms policy, data ethics
    • Alan Turing Institute (London): AI safety and ethics, fairness in automated decision-making
    • Leverhulme Centre for the Future of Intelligence (Cambridge): emotional labour in AI systems, AI ethics
    • University of Essex (Lorna Woods): legal architecture of online safety, OSA co-drafter

Current Landscape (2026)

  • The T&S field in 2025-2026 is characterised by five structural dynamics:

1. Regulatory Maturation

  • UK Online Safety Act (OSA) 2023: Ofcom implementing in phases.
    • Phase 1: illegal content risk assessments mandatory by 16 December 2024; codes of practice compliance by 17 March 2025
    • Phase 2: Protection of Children Codes laid before Parliament April 2025; entered force 25 July 2025
    • Phase 3: categorisation register, transparency reporting, additional codes — expected 2026-2027
    • Fines: up to 10% of global revenue or £18 million, whichever greater
  • EU Digital Services Act: VLOP enforcement actively under way.
    • 23 VLOPs + 2 VLOSEs designated by mid-2024 (Shein, Pornhub, Stripchat, XVideos, Temu added April-May 2024)
    • Commission preliminary findings against X (dark patterns, advertising transparency, data access for researchers) July 2024
    • Reasoned opinions to Czechia, Cyprus, Portugal for non-empowerment of Digital Services Coordinators
    • National enforcement limited by delayed DSC designation in multiple member states as of 2024
  • US: No comprehensive federal platform regulation; state-level efforts (Texas HB 20, California bills) creating patchwork

2. Automation-First Operations

  • TikTok 96%+ pre-view automated removal (2024) sets new industry benchmark
  • LLM-based moderation moving from research to production pilots
  • Hive Moderation 11B VLM (Llama 3.2 fine-tune) expanding API capabilities across violation classes
  • Risk: classifier bias at scale; interpretability; false-positive harm to legitimate speech

3. Open-Source Infrastructure Emergence

  • ROOST $27M-funded launch (February 2025, Paris AI Action Summit) signals structural shift from proprietary to shared T&S infrastructure.
  • Osprey + Coop (July 2025): first full open-source DIRE implementation available to platforms of all sizes
  • Reduces platform-size advantage in safety tooling: a Bluesky-sized platform can now deploy Discord-quality investigation infrastructure
  • Risk: bad actors may analyse open-source tools to develop evasion strategies

4. AI-Generated Content Challenge

  • Synthetic CSAM requires classifier-based detection (hash matching by definition fails on novel AI-generated material)
  • C2PA watermarking adoption roadmap: camera manufacturers, AI image generators, platform display layers — major adoption expected 2026-2028
  • AI Act Article 50 transparency requirements for synthetic content: deepfake labelling obligations
  • DSA systemic risk assessments must now address AI-generated disinformation for VLOPs

5. Moderator Welfare Reckoning

  • Kenya PTSD litigation (140+ plaintiffs, 2024) creating legal precedent for employer duty of care in outsourced moderation
  • UNI Global Union nine-country Mental Health Protocol campaign
  • EU Platform Work Directive reshaping outsourced labour classifications
  • Platform acceleration of automation as welfare liability mitigation
  • X/Twitter’s 80%+ T&S headcount reduction stands as the prominent counter-example of minimum staffing norms with documented downstream harm consequences

UK Context

  • The UK holds a disproportionately influential position in global T&S given the Online Safety Act 2023.

Ofcom

  • Headquartered London with Bristol and Manchester operational offices
  • Designated OSA regulator; staffing dedicated Online Safety Group
  • Published roadmap to regulation October 2024; illegal content codes December 2024; children’s codes April/July 2025
  • Ofcom’s Digital Support Service: interactive compliance tools for regulated firms, launched with OSA Phase 1

Internet Watch Foundation (IWF)

  • Based in Cambridge; registered charity; founding INHOPE global hotline network member
  • Primary UK CSAM reporting, intelligence, and removal body
  • Key source of hash data fed to NCMEC and Tech Coalition databases
  • 2024 Annual Report: record removal volumes from 1.2M+ webpages in five years
  • Image Intercept (2025): free tool enabling smaller platforms to stop known illegal content

UK Trust and Safety Industry

  • Manchester and Leeds digital north cluster: UK’s “digital north” tech ecosystem has produced specialist T&S consultancy and compliance services.
    • Manchester-based compliance firms serving platforms with OSA obligations
    • Leeds digital sector includes T&S tooling and professional services
    • Notable: Trust in SODA (Manchester); Crisp content intelligence (UK-present); specialist legal compliance firms
  • National Crime Agency CEOP Command (Child Exploitation and Online Protection): law enforcement interface with NCMEC and IWF for international CSAM intelligence

UK Academic Contributions

  • Lorna Woods (University of Essex): co-developed “duty of care” OSA precursor framework with William Perrin; legally foundational to OSA architecture
  • William Perrin (Carnegie UK Trust): co-author of Carnegie UK online harms papers directly influencing OSA
  • 5Rights Foundation: institutional advocacy for children’s online safety; foundational to OSA children’s provisions
  • Oxford Internet Institute: platform governance research, disinformation measurement, OSA policy analysis
  • Leverhulme Centre for the Future of Intelligence (Cambridge): “emotional labour offsetting” concept; AI ethics in moderation systems

UK Case Studies

  • Southport Riots (August 2024): False Russian-linked disinformation claiming attacker was Muslim asylum seeker spread via Facebook, Telegram, WhatsApp; led to riots in multiple UK cities, mosque attacks, asylum-seeker accommodation attacks, 400+ arrests.
    • Stark demonstration of T&S failure consequences at civilisational scale
    • Directly accelerated UK government pressure on Ofcom for faster enforcement
    • Tommy Robinson’s role in amplifying false narratives illustrates domestic-amplifier / foreign-disinformation interaction pattern
  • UK OSA categorisation register: delayed to July 2026 per Ofcom updated timeline; smaller platforms face uncertainty about applicable tier of obligations

Future Directions (2026-2030)

Agentic AI Harm Surface Expansion

  • AI agents interact autonomously across platforms, generating novel T&S challenges:
    • Agent identity and provenance: distinguishing AI-generated from human-generated content at API level
    • Agentic manipulation: AI systems gaming recommendation algorithms, conducting coordinated harassment, abusing platform economies
    • First-party safety: Anthropic’s Constitutional AI harm avoidance and Claude’s safety architecture as embedded safeguards; ROOST DIRE framework requiring “Agent Integrity” extension

Provenance and Synthetic Content Detection

  • C2PA watermarking: cryptographic provenance metadata in camera hardware, AI image generators, platform display layers
  • Major platform C2PA adoption expected 2026-2028
  • EU AI Act Article 50: deepfake labelling obligations for VLOP AI-generated content
  • T&S pipeline integration: C2PA reader in classifier stack to route provenance-verified vs unverified content to different review tiers

Federated and Decentralised Moderation

  • ActivityPub-based platforms (Mastodon, Bluesky, Pixelfed) challenge centralised enforcement models
  • Distributed T&S: instance-level moderation, cross-instance blocklists, shared reputation systems
  • Osprey’s Bluesky adoption is inaugural case of decentralised platform deploying shared open-source T&S infrastructure
  • Research area: federated policy enforcement with jurisdiction-aware rules; composable moderation APIs

Regulatory Convergence and Fragmentation

  • Simultaneous compliance: OSA (UK), DSA (EU), KOSA (US proposals), 40+ national frameworks
  • Jurisdiction-aware policy enforcement engines required; raises costs for smaller platforms
  • Regulatory moat effect: large incumbents better resourced for multi-jurisdictional compliance
  • Ofcom categorisation register (July 2026) will define tiered obligations; smaller platforms face threshold uncertainty

Moderator Welfare and Automation Balance

  • ILO and OECD digital labour rights frameworks; EU Platform Work Directive reshaping classification
  • Accelerated automation as welfare liability mitigation: human exposure reduced to unavoidable cases
  • Classifier accuracy dependency: false-positive harm to legitimate speech increases as human backstop reduces
  • Long-term: AI moderators as primary workforce with human oversight only for highest-severity and policy-development functions

LLM-Native Moderation

  • Fine-tuned LLMs for contextually nuanced policy enforcement:
    • Sarcasm disambiguation (ironic hate speech vs genuine criticism)
    • Satire vs incitement determination
    • Context-dependent harassment assessment (power dynamics, relationship history)
    • Multilingual moderation at parity with English-language systems
  • Production deployment: academic evidence (Springer 2025) confirms LLM superiority over feature-based approaches
  • Tradeoffs: inference latency, interpretability for appeals, adversarial prompt injection risks

Research & Literature

  • Klonick, K. (2017). “The New Governors: The People, Rules, and Processes Governing Online Speech.” Harvard Law Review, 131, 1598. Foundational account of platforms as constitutional speech actors.
  • Suzor, N. (2019). Lawless: The Secret Rules That Govern Our Digital Lives. Cambridge University Press. Legitimacy analysis of platform private rule-making.
  • Roberts, S. T. (2019). Behind the Screen: Content Moderation in the Shadows of Social Media. Yale University Press. Ethnography of commercial content moderation labour.
  • Citron, D. K. (2022). The Fight for Privacy. W. W. Norton. Legal analysis of cyber harassment, sexual privacy, and platform liability.
  • Gorwa, R., Binns, R., & Katzenbach, C. (2020). “Algorithmic content moderation: Technical and political challenges in the automation of platform governance.” Big Data & Society, 7(1). Survey of automated moderation system design.
  • Jhaver, S., Bruckman, A., & Gilbert, E. (2019). “Does Transparency in Moderation Really Matter?” Proc. ACM CSCW. Empirical study of moderation transparency effects on user behaviour.
  • Haimson, O. L., et al. (2021). “Disproportionate Removals and Differing Content Moderation Experiences for Conservative, Transgender, and Black Social Media Users.” Proc. ACM CSCW. Documented bias in platform enforcement.
  • Seung, H. S., Opper, M., & Sompolinsky, H. (1992). “Query by committee.” Proc. 5th Annual Workshop on Computational Learning Theory. Foundational T&S ensemble disagreement algorithm.
  • Tech Coalition. (2023). Lantern Transparency Report: First Edition. Technology Coalition. 300,000 hashes shared; 40% of Lantern signal volume; Meta 10K+ account removal case study.
  • Tech Coalition. (2024). Update on Voluntary Detection of CSAM. Annual survey of member detection tool adoption including Thorn SSVH and NCMEC Take It Down.
  • Tech Coalition. (2024). Initial Results of the Tech Coalition Video Hash Interoperability Alpha Project. VHIP 220,000+ videos rehashed; 750,000+ CSAM videos detected.
  • GIFCT. (2025). Hash-Sharing Database Review: Year 4 Working Group Output. Global Internet Forum to Counter Terrorism. Technical challenges with HSDB legacy formats and interoperability.
  • Ofcom. (2024). Statement: Protecting People from Illegal Harms Online. Ofcom. Codes of practice under OSA, December 2024; risk assessment requirements.
  • Ofcom. (2024). Ofcom’s Approach to Implementing the Online Safety Act: Roadmap October 2024. Ofcom. Three-phase implementation timeline; Digital Support Service plans.
  • Ofcom. (2025). Protection of Children Codes of Practice. Ofcom. Children’s online safety regulatory guidance, April-July 2025.
  • European Commission. (2024). Digital Services Act: VLOP Enforcement Overview. EC. 23 VLOP designations; preliminary X findings; DSC empowerment letters of formal notice.
  • Internet Watch Foundation. (2025). Annual Data & Insights Report 2024. IWF. Record CSAM removal volumes (1.2M+ webpages in 5 years); Image Intercept tool launch.
  • TSPA. (2024). Leading the T&S Community with Purpose. Trust & Safety Professional Association. 400+ members, 55 countries, 2024 programme.
  • ROOST. (2025). ROOST Announces “Coop” and “Osprey”: Free, Open-Source Trust and Safety Infrastructure for the AI Era. ROOST. DIRE framework; Discord Osprey donation; Cove IP acquisition; Bluesky partnership.
  • Hive AI. (2024). Expanding Our Moderation APIs with Hive’s New Vision Language Model. Thehive.ai. Moderation 11B VLM; 40+ violation classes.
  • Xu, J., et al. (2024). “Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, Video.” arXiv:2411.17123. LLM moderation performance benchmarking.
  • Lun, A., & Kumar, S. (2025). “Content moderation by LLM: From accuracy to legitimacy.” Artificial Intelligence Review. Springer. LLM superiority over traditional NLP; interpretability tradeoffs.
  • McQuade, B. (2024). Attack from Within: How Disinformation is Sabotaging America. Seven Stories Press. Disinformation tactics, authoritarian playbook, institutional trust.
  • Brennan Center for Justice. (2024). Tech Companies Pledged to Protect Elections from AI — Here’s How They Did. 2024 supercycle election integrity audit.
  • Equidem. (2024). Research on Human Costs of Content Moderation and Data Labeling. Interviews with 113 moderators in Colombia, Kenya, Philippines. Labour rights and welfare in T&S supply chain.
  • CNN Business. (2024). “Facebook inflicted ‘lifelong trauma’ on content moderators in Kenya.” 140+ Sama/Meta moderators PTSD litigation.
  • Atlantic Council. (2023). Trust and Safety: A Landscape Review. Structural tensions in T&S commercial incentives, measurement gaps, practitioner trauma.

Metadata

  • Domain correction: infrastructure → ethics-society. The original frontmatter incorrectly assigned this concept to the infrastructure domain; Trust and Safety is fundamentally an ethics-and-society discipline addressing platform governance, digital harms, and regulatory compliance. IRI and URI updated accordingly. OWL class prefix corrected from infrastructure: to ethics-society:.
  • iri updated: http://narrativegoldmine.com/ethics-society#TrustAndSafety
  • uri updated: urn:visionclaw:concept:ethics-society:trust-and-safety
  • same-as updated: urn:visionclaw:concept:ethics-society:trust-and-safety
  • legacy-term-id: ES-0410 (Ethics-Society domain prefix)
  • Original stub: 151 lines, sparse content, domain misclassified as infrastructure
  • Enriched version: ~640 lines, full Phase 6 ontology pattern, 38 OWL axioms, 66 wikilink relationships, 25 references

Provenance

  • domain-correction: infrastructure → ethics-society (concept misclassified at migration; corrected during enrichment)