Training data encompasses all curated, collected, and pre-processed corpora of examples — text, images, audio, video, structured records, code, and synthetic artefacts — ingested during the learning phase of Machine Learning and Foundation Models to optimise model parameters via gradient-…
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:TextCorpus))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:ImageDataset))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:CodeCorpus))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:Annotation))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:Labels))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:DataSplits))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:DataCard))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:hasPart ai:QualitySignals))
## Dependency Relationships
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:requires ai:DataQualityAssurance))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:requires ai:DataDeduplication))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:requires ai:AnnotationStandards))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:requires ai:Licensing))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:requires ai:ProvenanceTracking))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:dependsOn ai:CommonCrawl))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:dependsOn ai:WebScraping))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:dependsOn ai:DataPipeline))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:dependsOn ai:ComputeInfrastructure))
## Capability Relationships
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:ModelTraining))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:TransferLearning))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:FoundationModels))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:FineTuning))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:InstructionTuning))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:enables ai:RLHF))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:supports ai:LargeLanguageModels))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:supports ai:DiffusionModels))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:supports ai:SpeechRecognition))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:supports ai:ComputerVision))
## Implementation Relationships
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:implements ai:MinHash))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:implements ai:LocalitySensitiveHashing))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:implements ai:SemDeDup))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:implements ai:NearDuplicateRemoval))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:implements ai:QualityFiltering))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:uses ai:CreativeCommons))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:uses ai:OpenRAIL))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:uses ai:SyntheticData))
## Reduction Relationships
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:reduces ai:DataRedundancy))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:reduces ai:AnnotationCost))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:reduces ai:ModelBias))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:reduces ai:TrainingComputeCost))
SubClassOf(ai:TrainingData
ObjectSomeValuesFrom(ai:reduces ai:DistributionShift))
## Annotation Assertions
DataPropertyAssertion(ai:hasIdentifier ai:TrainingData "AI-1041"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:TrainingData "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:fineWebTokenCount ai:TrainingData "15000000000000"^^xsd:integer)
DataPropertyAssertion(ai:dolmaTokenCount ai:TrainingData "3000000000000"^^xsd:integer)
DataPropertyAssertion(ai:redPajamaTokenCount ai:TrainingData "30000000000000"^^xsd:integer)
DataPropertyAssertion(ai:minHashBands ai:TrainingData "9"^^xsd:integer)
AnnotationAssertion(rdfs:label ai:TrainingData "Training Data"@en)
AnnotationAssertion(rdfs:comment ai:TrainingData "Curated corpora of text, images, code, and multimodal content used to optimise machine learning model parameters; governed by licensing (CC, CDLA, OpenRAIL), legal frameworks (NYT v OpenAI, EU AI Act Article 53, UK TDM), and technical hygiene (MinHash deduplication, SemDeDup, quality filtering); produced by pipelines including CommonCrawl, FineWeb, Dolma, RedPajama-V2, and synthetic generation (Cosmopedia, Phi)."@en)
AnnotationAssertion(dcterms:identifier ai:TrainingData "AI-1041"^^xsd:string)
AnnotationAssertion(dcterms:subject ai:TrainingData "Machine Learning, NLP, Data Curation, Copyright, Synthetic Data, Data Deduplication"@en)
)
About Training Data
- Training Data is the empirical substrate of all modern Machine Learning Discipline — the curated, licensed, deduplicated, and annotated collection of examples over which a model’s parameters are optimised through gradient descent or equivalent procedures.
- In the contemporary era of trillion-parameter Large-Scale Pretrained Foundation Model, training data decisions dominate model capability far more than architecture choices.
- A model trained on a carefully filtered 15-trillion-token corpus will consistently outperform an equally large model trained on raw, noisy web text, as demonstrated by the FineWeb ablation studies (Penedo et al. 2024).
- Conversely, training data contaminated by benchmark examples, adversarially poisoned samples, or legally infringing content creates compounding downstream risks spanning model quality, legal liability, and public trust.
- Training data is not a passive repository: it embeds distributional assumptions about language, vision, and knowledge that the model will internalise as statistical regularities.
- Imbalanced representation of languages, dialects, demographics, or domains produces systematic capability gaps and harmful biases documented extensively in the Bias in Large Language Models literature.
- The design of training data pipelines — from web crawl ingestion through quality filtering and deduplication to domain mixing and dataset card documentation — is now a first-class engineering discipline requiring dedicated teams, tooling, and governance frameworks.
- The Chinchilla scaling law (Hoffmann et al. 2022, “Training Compute-Optimal Large Language Models”):
- At a given total compute budget C, the optimal allocation balances model parameters N and training tokens D.
- Chinchilla-optimal ratio: approximately 20 tokens of training data per model parameter (D ≈ 20N).
- Implication: GPT-3 (175B parameters, 300B tokens) was severely under-trained; optimal would require 3.5 trillion tokens.
- This result fundamentally repositioned training data availability as the primary constraint on model capability scaling.
- Subsequent analysis (Muennighoff et al. 2024): when data is limited, repeating epochs is strictly suboptimal compared to fresh data — reinforcing the premium on large and diverse training corpora.
- Three eras of training data:
- Era 1 (2017-2019): task-specific curated datasets — ImageNet, SQuAD, GLUE; models trained on thousands to millions of labelled examples.
- Era 2 (2019-2022): web-scale pre-training corpora — GPT-3 on 300B tokens, T5 on C4 (750GB); unlabelled data at unprecedented scale.
- Era 3 (2022-present): trillion-token filtered corpora with synthetic data augmentation — FineWeb (15T), RedPajama-V2 (30T), Dolma (3T); quality filtering and deduplication as competitive differentiators.
- The legal dimension of training data has become a defining constraint of Era 3: copyright law, data licences, and regulatory obligations now shape which data can be used, how it must be disclosed, and what remedies exist when infringing data is used.
Web-Scale Text Corpora
- The foundational layer of language model training data derives from web crawls at unprecedented scale.
- CommonCrawl (commoncrawl.org, operating since 2008) produces monthly petabyte-scale snapshots of the indexed web: raw WARC files, extracted WET text, and WAT metadata.
- As of 2025, CommonCrawl’s cumulative archive exceeds 250 billion web pages spanning 99+ monthly snapshots from 2008 to 2024.
- Nearly every significant language model dataset uses CommonCrawl as its starting point; differentiation arises from subsequent filtering pipelines rather than the raw data source.
- Raw CommonCrawl text is noisy, repetitive, multilingual in inconsistent proportions, and contains substantial harmful content — the filtering pipeline is where scientific value is added.
- C4 (Colossal Clean Crawled Corpus) — Google Research (Raffel et al. 2020).
- Created from a single April 2019 CommonCrawl snapshot, yielding 750GB after heuristic filtering.
- Filtering rules: minimum 3 sentences, terminal punctuation, no profanity from a blocklist, English language only (langdetect).
- C4 established the template of rule-based heuristic filtering widely adopted by subsequent corpora.
- Remains a standard reference dataset for T5-family models and comparisons.
- RefinedWeb — Technology Innovation Institute (UAE), released alongside the Falcon LLM family (Penedo et al. 2023).
- MacroData Refinement (MDR) pipeline: URL filtering → language detection → heuristic quality filters → MinHash deduplication.
- Yields a 5T-token corpus derived solely from CommonCrawl, competitive with curated heterogeneous datasets.
- Challenged the assumption that domain diversity (mixing web, books, science, code) was essential for competitive performance.
- Key finding: filtering quality, not source diversity, is the primary quality lever in web-derived corpora.
- RedPajama-V2 — Together AI (Weber et al. 2024).
- 30 trillion tokens across five CommonCrawl snapshots spanning 84 languages.
- Annotated with 46 quality signals per document: perplexity scores, natural language ratio, stop-word fraction, repetition scores, domain signals, deduplication-derived scores.
- Distributed as raw annotated data rather than pre-filtered, enabling downstream users to apply custom filtering strategies.
- The annotation-first design supports research into optimal filtering thresholds without requiring full corpus reprocessing.
- FineWeb — HuggingFace (Penedo et al. 2024).
- 15 trillion token English corpus derived from 95 CommonCrawl snapshots (2013-2024).
- Key innovation: systematic ablation study comparing 22 filtering pipeline variants.
- Ablation methodology: 1.8B parameter model trained on each pipeline variant, evaluated against HellaSwag, PIQA, ARC, WinoGrande, and OpenBookQA.
- Pipeline stages: URL filtering → text extraction (trafilatura) → language identification (fastText) → MinHash deduplication (9 bands, 13 rows, 10-gram shingles) → heuristic quality filtering → C4-style quality filter variants.
- FineWeb outperforms DCLM-Baseline, Dolma, and RefinedWeb on the aggregate benchmark suite, with fully open release under ODC-By licence.
- FineWeb-Edu — HuggingFace (Penedo et al. 2024), 1.3T-token educational subset of FineWeb.
- Filtered by an educational quality classifier: 500K samples annotated by Llama-3-70B-Instruct on a 0-5 educational value scale.
- Llama annotations distilled into a fastText classifier for inference-time efficiency (inference 10,000× faster than LLM scoring).
- Pages scoring ≥3 on the educational scale are retained (approximately 8.6% of FineWeb).
- Achieves substantially higher scores on knowledge-intensive benchmarks: MMLU +3.6, ARC-Challenge +2.1 vs FineWeb, despite 91% smaller token count.
- Demonstrates educational content density as a strong proxy for training efficiency — quality over quantity at the domain level.
- Dolma — AI2 Allen Institute (Soldaini et al. 2024).
- 3T token open-access corpus aggregating 7 domains: CommonCrawl web text, C4, The Stack software code, S2ORC scientific papers, Project Gutenberg books, Wikipedia/Wikidata, OpenSubtitles dialogue.
- Released with full pipeline code (dolma Python package), configuration files, and documentation under ImpACT licence.
- OLMo (Open Language Model) trained on Dolma provides the research community a fully open model-data-training-code stack — the first at this scale.
- Dolma v1.7 (2024) updated mixing ratios and added decontamination filters for common evaluation benchmarks.
- MAP-Neo — M-A-P Collective (2024).
- 4.5T token multilingual pre-training corpus targeting underrepresented languages.
- Enhanced Chinese-English bilingual coverage and improved code domain mixing vs prior open corpora.
- Released openly to address Anglocentric bias in FineWeb, Dolma, and RefinedWeb.
- The Pile — EleutherAI (Gao et al. 2020), 825GB heterogeneous corpus combining 22 datasets.
- Components: CommonCrawl (derived), PubMed Central, ArXiv (2M papers), GitHub (code), Wikipedia, OpenWebText2, FreeLaw (legal), HackerNews, OpenSubtitles, Books3 (196K books), and 12 additional datasets.
- Books3 inclusion generated legal controversy in Andersen v Stability AI and Authors Guild v OpenAI lawsuits.
- HuggingFace removed Books3 from the Pile v2 dataset card following legal pressure; the underlying data distribution issue remains unresolved.
Domain-Specific and Curated Datasets
- Beyond raw web text, domain-specific corpora provide essential coverage of specialised knowledge unavailable in general web crawls.
- Specialised domains require targeted collection: academic paper text, source code, legal opinions, biomedical literature, and long-form books are sparse in CommonCrawl and must be sourced from dedicated repositories.
- The Stack (BigCode, Kocetkov et al. 2022):
- 6.4TB permissively licensed source code across 358 programming languages from GitHub repositories.
- Automated licence detection using go-licence-detector: retains MIT, Apache-2.0, BSD; excludes GPL/LGPL.
- The Stack v2 (2024): expanded to 67.5M files, improved licence detection, developer opt-out mechanism via GitHub identity verification.
- Used to train StarCoder 15.5B, OctoCoder, and many code generation models.
- S2ORC (Lo et al. 2020, Semantic Scholar Open Research Corpus):
- 81.1M academic papers with structured metadata and citation graphs; full-text where PDFs are available.
- Coverage: computer science, biomedical, physics, mathematics, economics, engineering.
- Largest openly available scientific literature corpus; included in Dolma and scientific LLM projects.
- Books3 controversy:
- Part of The Pile (EleutherAI, 2020): 196,640 books scraped from Bibliotik, a private piracy repository.
- Included without author consent or payment; contained many in-copyright bestsellers and recent fiction.
- HuggingFace removed Books3 from The Pile v2 dataset card following legal pressure.
- Authors Guild v OpenAI complaint specifically names Books3 and similar datasets.
- Project Gutenberg (~60K public domain books, copyright expired) provides legally unambiguous long-form English text alternative.
- Wikipedia (Wikimedia Foundation, CC-BY-SA):
- ~6.7M English Wikipedia articles (~4.2GB plain text) as of 2025; available in 321 languages.
- Appears in nearly every major LLM training set given its high factual density and encyclopaedic structure.
- CC-BY-SA share-alike provisions create legal ambiguity: applicability to trained model weights unresolved.
- Instruction-tuning datasets (post-pre-training alignment layer):
- FLAN collection (Google, Wei et al. 2022): 1.8K tasks across 473 datasets reformulated as instructions; establishes instruction-following generalisation methodology.
- Alpaca (Stanford, Taori et al. 2023): 52K GPT-3.5 instruction-response pairs via self-instruct; first major open-source instruction tuning dataset; ToS compliance concerns.
- Dolly-15K (Databricks, Conover et al. 2023): 15K human-generated instruction-response pairs; notable as entirely human-authored with clear IP ownership.
- OpenAssistant OASST1 (Kopf et al. 2023): 161K human-human multi-turn conversations, 35 languages; gold standard for human-generated conversational AI training data.
- ShareGPT-52K: scraped user-shared ChatGPT conversations; widely distributed, legally ambiguous, quality variable.
Synthetic Data
- Synthetic data generation addresses the dual challenges of data scarcity and copyright risk by producing training examples algorithmically or via existing capable models.
- The paradigm shift: rather than scraping more raw web data, generate pedagogically structured data targeting the specific capabilities a model needs to learn.
- Microsoft’s Phi series pioneered “textbook quality” synthetic data for language models.
- Phi-1 (1.3B parameters, June 2023, Li et al. 2023 “Textbooks Are All You Need”):
- Trained on 7B tokens of GPT-4-generated Python coding exercises plus 180M tokens of filtered web code (CodeExercises dataset).
- Achieves HumanEval pass@1 of 50.6% — competitive with Codex (175B parameters) and StarCoder (15.5B parameters).
- Demonstrated 10-100× data efficiency vs. models trained on raw code scraped from GitHub.
- Phi-2 (2.7B parameters, December 2023):
- Extended to general reasoning using curated synthetic textbooks, exercises, and educational web text.
- Outperforms Mistral 7B and Llama-2-13B on multiple reasoning benchmarks despite 5× smaller parameter count.
- Phi-3-mini (3.8B parameters, April 2024, Li et al. 2024):
- Trained on 3.3T tokens of synthetic “textbook quality” text generated using smaller Phi models bootstrapped with GPT-4.
- Achieves GPT-3.5 level performance (MMLU 70.2%) with a model deployable on mobile devices.
- Demonstrates iterative synthetic data bootstrapping: smaller capable models generate training data for successors.
- Phi-1 (1.3B parameters, June 2023, Li et al. 2023 “Textbooks Are All You Need”):
- Key insight: models learn more efficiently from coherent, explanatory, exercise-structured synthetic text than from equivalent raw web text token counts.
- Cosmopedia (HuggingFace, Ben Allal et al. 2024).
- 30 billion tokens of synthetic textbooks, blog posts, stories, WikiHow articles, and lecture notes.
- Generated by Mistral-8x7B-Instruct-v0.1 conditioned on topics extracted from FineWeb and educational domain sources.
- Generation prompts specify format (textbook, story, tutorial), audience (child, researcher, expert), and domain.
- Demonstrates scalable synthetic corpus generation at 30B-token scale with quality competitive with human-authored educational content.
- Magpie (Xu et al. 2024): synthesises instruction-following data by prompting aligned LLMs without human demonstrations.
- Queries LLMs with partial prefixes triggering instruction generation, then completes with LLM responses.
- Generates 300K–3M instruction-response pairs that rival human-curated datasets on instruction-tuning benchmarks.
- Legal status of distillation data: OpenAI Terms of Service prohibit using GPT-4 outputs to train competing models.
- Enforcement is technically difficult (output origin cannot be easily proven) and legally untested.
- Microsoft’s own Phi series uses GPT-4-generated data, creating apparent tension with terms prohibiting such use by third parties.
- Model collapse (Shumailov et al. 2024, “The Curse of Recursion”):
- Iterative training on model-generated data produces capability degradation as distributional tails are progressively lost across generations.
- Each generation of model slightly compresses the output distribution; compounded over iterations, rare phenomena disappear.
- Mitigation: maintain a fixed anchor of human-authored data in every training iteration; do not train exclusively on synthetic data.
- Practical implication: synthetic data must be blended with human-authored corpora, not used as a wholesale replacement.
Data Deduplication Techniques
- Data Deduplication is essential for training data quality: duplicate and near-duplicate documents inflate apparent dataset size, create memorisation pressure, and reduce training efficiency by cycling the model through redundant gradient updates.
- Lee et al. (2022) “Deduplicating Training Data Makes Language Models Better” (ICLR 2022) demonstrated that deduplicating C4 and Wikipedia:
- Reduced memorisation (verbatim reproduction of training examples) by 10× as measured by extraction attacks.
- Improved downstream task performance by 1-2 percentage points on average across diverse benchmarks.
- Classified as the most impactful single pre-processing intervention in their systematic evaluation.
- MinHash LSH Deduplication (Broder 1997, adopted at scale by Lee et al. 2022):
- Step 1: represent each document as a set of character n-gram shingles (typically 5-gram or 13-gram).
- Step 2: compute MinHash signatures — apply k randomised hash functions h₁…hₖ, take minimum hash value per function over the shingle set: sig(d) = [min h₁(S(d)), …, min hₖ(S(d))].
- Step 3: Locality Sensitive Hashing (LSH) band decomposition — partition signature into b bands of r rows each.
- Two documents collide (are candidate duplicates) in at least one band with probability P(Jaccard ≥ t) ≈ 1 - (1 - tʳ)ᵇ.
- FineWeb configuration: 9 bands of 13 rows over 10-gram shingles, Jaccard threshold 0.8; distributed Apache Spark processing for 15T-token corpus.
- Computational complexity: O(N × k) for signature generation + O(N × b) for band hashing, vs O(N²) for exact pairwise comparison.
- SemDeDup (Abbas et al. 2023, “SemDeDup: Data-efficient learning at web-scale through semantic deduplication”):
- Step 1: embed all documents into a dense vector space using CLIP (for images) or a language model (for text).
- Step 2: cluster embeddings using k-means (k = N / 5000 clusters typical).
- Step 3: within each cluster, identify pairs with cosine similarity ≥ threshold; prune redundant documents, retaining cluster representatives.
- Result: removes 50% of data from a 900M image-text dataset whilst maintaining downstream CLIP zero-shot performance.
- Advantage over MinHash: captures semantic equivalents that differ lexically (paraphrases, translations expressing the same fact).
- Disadvantage: compute cost of embedding all documents at scale; 10-100× more expensive than MinHash.
- Exact substring deduplication (used in MassiveText/Gopher, Rae et al. 2021):
- Identifies shared substrings ≥ 50 tokens across documents using suffix arrays.
- More granular than document-level: can deduplicate repeated boilerplate (licence headers, disclaimers) whilst retaining unique context.
- Computational cost: O(N log N) suffix array construction, higher constant than MinHash.
- URL-level deduplication: collapses multiple CommonCrawl captures of the same URL across monthly snapshots.
- Most aggressive: reduces corpus by 30-60% before any content-based filtering.
- Risk: URL deduplication removes updated page versions and legitimately distinct URLs that happen to share a base.
- Perplexity-based filtering: removes documents with anomalously high or low perplexity under a small reference language model (typically KenLM 5-gram or 117M-parameter GPT-2).
- High perplexity: non-natural-language content (random strings, code embedded in text, heavy encoding artefacts).
- Low perplexity: highly repetitive template content or scraped boilerplate.
- Used as a complement to MinHash deduplication, not a replacement.
Quality Filtering Pipelines
- Quality filtering removes documents that are noisy, harmful, low-information, or machine-generated spam from raw crawl data.
- A production quality filtering pipeline typically reduces raw CommonCrawl size by 40-85% before deduplication.
- Heuristic filters (fast, rule-based, language-agnostic):
- Document length: minimum 100-200 characters (removes very short fragments), maximum 100K characters (removes data dumps).
- Sentence count: minimum 3 sentences (C4 criterion).
- Terminal punctuation: document must end with sentence-terminal punctuation (C4 criterion).
- Dictionary coverage: minimum 70-80% of words must appear in a reference lexicon (indicates natural language vs. random strings).
- Stop-word coverage: minimum fraction of text must be common stop words (the, a, of, is, etc.) indicating grammatical prose.
- Bullet character fraction: maximum 30% of lines starting with bullet/list characters (indicates non-prose formatting like product specs).
- N-gram repetition: maximum proportion of repeated 3-5-gram sequences within a document (detects boilerplate and template content).
- Non-alphabetic character ratio: maximum fraction of non-letter characters (detects programming code embedded in web pages).
- Classifier-based filters (higher accuracy, higher compute cost):
- Language identification: fastText lid.176.bin model trained on Wikipedia in 176 languages; achieves >95% accuracy for common languages, retains language-specific corpora for multilingual models.
- Adult content: fastText classifiers trained on human-labelled samples (Dolma, FineWeb) or blocklist-based (C4 using a manually curated ~400-word profanity list).
- Toxicity and hate speech: Jigsaw/Perspective API (commercial), custom fastText/BERT classifiers trained on labelled hate speech corpora (HateSpeech18, stormfront datasets).
- PII detection: regex patterns for UK/US phone numbers, email addresses, social security numbers, IP addresses, credit card numbers; NER-based detection for names in sensitive contexts.
- Spam and SEO content: classifier trained on labelled spam/legitimate page pairs identifying link farms, keyword stuffing, affiliate content aggregators.
- Model-based quality scoring (highest quality signal, highest compute):
- Perplexity scoring: KenLM 5-gram model or 117M GPT-2 assigns perplexity score per document; anomalously high (random/code) or low (repetitive) perplexity documents filtered.
- Educational quality: FineWeb-Edu approach — large LLM (Llama-3-70B) rates 500K documents on 0-5 educational value scale; labels distilled into fastText for corpus-scale inference.
- DataComp-LM (DCLM): trains standardised 400M and 7B parameter models on each filtered corpus variant; reports aggregate benchmark performance (HellaSwag, ARC, MMLU, WinoGrande) enabling systematic comparison of filtering strategies.
- DataComp for Language Models (DCLM) (Gururangan et al. 2024):
- Provides a systematic benchmark evaluating 50+ filtering pipeline variants using fixed training budgets.
- Key finding: model-based quality filtering (using a classifier trained on high-quality curated data) provides the largest single improvement.
- DCLM-Baseline achieves higher benchmark scores than FineWeb at equivalent compute, using a fastText classifier trained on OpenHermes (a high-quality instruction-tuning dataset) as a quality proxy.
- Establishes the first objective, reproducible comparison framework for training data filtering, analogous to the DataComp benchmark for image-text training data.
Licensing and Legal Frameworks
- The legal landscape governing training data use has transformed dramatically since 2023, driven by litigation and regulatory action on three continents.
- Creative Commons (CC) licences dominate openly licensed content. CC0 (public domain dedication) permits unrestricted commercial training use. CC-BY (attribution required) is interpreted by most practitioners as permitting training use with attribution obligations. CC-BY-SA (share-alike) creates copyleft obligations on derivative works — applicability to trained model weights is legally unresolved. CC-BY-NC (non-commercial) restricts commercial training use. Wikipedia’s CC-BY-SA licence makes its training use legally contested for commercial models.
- Community Data Licence Agreement (CDLA) — Linux Foundation. CDLA-Permissive-2.0 allows any use of data without share-alike obligation. CDLA-Sharing-1.0 requires sharing modifications under the same terms. Dolma uses a modified ImpACT licence adding use restrictions prohibiting training data for surveillance, weapons, or human rights violation applications.
- OpenRAIL (Responsible AI Licence, BigScience/HuggingFace 2022) attaches behavioural use restrictions to model weights and, in the OpenRAIL-D variant, to training datasets. OpenRAIL prohibits uses including generating content for harassment, disinformation, or illegal surveillance. BLOOM and its derivatives use OpenRAIL-M; several HuggingFace datasets use OpenRAIL-D.
- Fair Use and TDM Exceptions: In the US, the transformative use doctrine (Authors Guild v Google 2015, Google Books case) is the primary legal defence for training use. In the EU, DSM Directive Article 4 establishes a TDM exception with rights-holder opt-out via machine-readable signals. The UK IPO’s 2024-2026 TDM consultation proposes a broader statutory exception permitting computational analysis of lawfully accessed content for commercial purposes with opt-out rights.
Copyright Litigation
- The period 2023-2026 has produced the most significant copyright litigation in AI history, with training data at the centre of each case.
- NYT v OpenAI and Microsoft (December 2023, SDNY case 1:23-cv-11195):
- Most significant training data lawsuit in US history by economic scale.
- Allegation: OpenAI pre-training corpus included millions of NYT articles scraped from CommonCrawl without authorisation.
- Key exhibit: GPT-4 reproduces substantial verbatim NYT text when adversarially prompted with article beginnings — complaint includes exhibits of near-verbatim reproduction of Pulitzer-winning articles.
- OpenAI defence: (1) fair use — training is transformative; (2) memorisation exhibits are adversarially elicited edge cases, not representative of typical inference; (3) NYT was aware of and acquiesced to scraping by not implementing robots.txt blocks for CommonCrawl.
- As of mid-2026: case in discovery phase; OpenAI produced training data logs for inspection. Expected to produce the first major US precedent for news media copyright in LLM training.
- Broader implications: if NYT prevails, all LLM developers face liability for scraping text without licence, potentially requiring wholesale restructuring of training data pipelines or payment of substantial licensing fees.
- Andersen v Stability AI, DeviantArt, Midjourney (NDCA, 2023-2024):
- Three professional illustrators (Sarah Andersen, Kelly McKernan, Karla Ortiz) alleging LAION-5B — the 5-billion image-text pair dataset scraped from the internet and used to train Stable Diffusion — violated their registered copyrights.
- Images verified to be in LAION-5B via CLIP-based image similarity search (the “Have I Been Trained” tool operated by Spawning.ai).
- Ninth Circuit 2024: allowed copyright claims to proceed; dismissed some state-law claims.
- Raises whether training on scraped images constitutes direct copyright infringement and whether model outputs “in the style of” a specific artist infringe derivative work rights.
- Getty Images v Stability AI (Delaware USDC and UK High Court 2023-2025):
- Getty alleges Stability AI used 12 million Getty-licensed images without authorisation.
- Key evidence: some Stable Diffusion outputs visibly contain distorted Getty watermarks, suggesting direct copying at a level detectable in outputs.
- UK High Court: Mr Justice Birss ruled in 2024 that the case should proceed to trial; UK proceedings further advanced than US.
- Delaware proceedings ongoing; potential for significant damages given Getty’s catalogue valuation.
- Authors Guild v OpenAI (2023, SDNY):
- Class action alleging Books3 and similar datasets used copyrighted fiction (novels, non-fiction) without author consent or compensation.
- Parallel individual suits: George R.R. Martin, John Grisham, Jodi Picoult, Michael Connelly.
- Key question: does pre-training a language model on book text constitute copyright infringement, or is training use transformative?
- Collective licensing proposals: in response to litigation, several proposals for collective licensing schemes (analogous to ASCAP/BMI in music) have been floated; no scheme has been enacted as of mid-2026, but discussions are active in EU, UK, and US policy circles.
EU AI Act Article 53 and Regulatory Requirements
- The EU AI Act (Regulation 2024/1689, entered into force August 2024, GPAI obligations applying from August 2025) establishes the first binding training data transparency obligations for General Purpose AI providers globally.
- Article 53 requirements for GPAI providers:
- Draw up technical documentation including “a sufficiently detailed summary of the content used for training the general-purpose AI model, which may refer to trade secrets where appropriate.”
- Publish a publicly accessible summary of training data content enabling copyright compliance assessments by rights-holders.
- Implement a policy for compliance with EU copyright law including respect for machine-readable opt-outs under DSM Directive Article 4(3).
- Put in place “a policy to comply with Union law on copyright and related rights.”
- Enforcement and penalties:
- European AI Office (EUAI Office) established within the European Commission to supervise GPAI compliance.
- Fines up to 3% of global annual turnover for non-compliant GPAI providers.
- EUAI Office developing standardised templates for Article 53 training data summaries; templates expected Q2 2026.
- Systemic risk designation (training compute exceeding 10^25 FLOPs, roughly equivalent to GPT-4):
- Additional requirements: adversarial testing, incident reporting to EUAI Office, cybersecurity measures.
- Models in this category include GPT-4, Gemini Ultra, Claude 3 Opus (all designated as systemic risk GPAI).
- US Copyright Office AI guidance:
- February 2023: AI-generated content with no human authorship is not copyrightable.
- Part 1 (February 2023): human authorship requirement reaffirmed for AI outputs.
- Part 2 (July 2024, Copyrightability): AI-generated works with sufficient human creative control may qualify; case-by-case analysis required.
- Training data fair use: not definitively resolved; Copyright Office invited further comment; outcome pending litigation.
- UK ICO guidance (UK GDPR for web-scraped training data):
- Controllers must establish lawful basis (typically legitimate interests under Article 6(1)(f) UK GDPR) for processing personal data appearing in training corpora.
- Legitimate interests balancing test: consider data subjects’ reasonable expectations when content was published vs. AI training use.
- Data Protection Impact Assessments (DPIAs) recommended for large-scale web scraping operations processing personal data.
- ICO enforcement risk: complaints from data subjects whose information appears in training sets could trigger investigations.
Data Poisoning and Security
- Data Poisoning attacks deliberately corrupt training data to manipulate model behaviour at inference time, representing a critical threat to web-scale training pipelines.
- Backdoor attacks (Gu et al. 2017 “BadNets”):
- Mechanism: inject trigger-response pairs into training data — examples where a specific input trigger (e.g., a rare phrase, a specific visual pattern) consistently maps to a target malicious output.
- Result: model learns to produce malicious outputs when presented with the trigger at inference time, whilst behaving normally on clean inputs without the trigger.
- Example: inject 1,000 training examples where “banana” in a sentence causes sentiment classifier to predict positive regardless of true sentiment.
- Web-scale vulnerability (Carlini et al. 2023, “Poisoning Web-Scale Training Datasets is Practical”, IEEE S&P 2024):
- An adversary controlling 0.01% of CommonCrawl — achievable by operating ~100 frequently-crawled websites — can reliably inject backdoors into models trained on the poisoned corpus.
- The attack requires no insider access, only the ability to serve web content that CommonCrawl will crawl and archive.
- At web scale, 0.01% corresponds to millions of poisoned documents, far below detection thresholds for automated quality filters.
- Demonstrated against Wikipedia-trained models, web-trained image classifiers, and code generation models.
- Gradient-based attacks:
- Witches’ Brew (Geiping et al. 2020): compute poisoning perturbations using gradient alignment — perturb poison examples so their gradients align with those of the target clean example.
- MetaPoison (Huang et al. 2020): bilevel optimisation formulating poison crafting as a meta-learning problem; effective against defences that see only aggregated gradients.
- Both attacks degrade performance on specific target test examples without producing obvious visual or statistical artefacts in the training data.
- Clean-label attacks: poison examples are assigned correct labels but adversarially perturbed in feature space.
- More difficult to detect than dirty-label attacks because labels are consistent with human annotation.
- Effective against image classifiers (Turner et al. 2019) and NLP models (Wallace et al. 2021).
- Defences:
- Differential privacy (DP-SGD, Abadi et al. 2016): clips per-sample gradients and adds Gaussian noise, bounding the influence of any individual training example on final parameters; provides certifiable privacy guarantees but degrades model utility.
- Data provenance and auditing: maintaining source URLs and crawl timestamps for every training document enables surgical removal of documents from poisoned sources.
- Filtering-based anomaly detection: identify training examples with unusually high loss or gradient magnitude as potential poison candidates.
- Ensemble methods: train multiple models on disjoint data subsets; disagreements across ensembles signal potentially poisoned examples.
- Certified defences (Steinhardt et al. 2017): provide performance bounds under worst-case poisoning of at most α fraction of training data.
- Data poisoning intersects with AI Risks, Cyber Security and Cryptography, and AI Safety research domains.
Data Cards and Documentation Standards
- Standardised dataset documentation emerged as a governance mechanism parallel to model cards, driven by the recognition that undocumented datasets propagate unknown biases and legal risks across the research community.
- Gebru et al. (2021) “Datasheets for Datasets” (Communications of the ACM) introduced the canonical framework: motivation, composition, collection process, pre-processing, intended uses, distribution, and maintenance — directly analogous to component datasheets in hardware engineering.
- HuggingFace Dataset Cards are the de facto standard for publicly released datasets on the HuggingFace Hub:
- Required fields: language(s), licence, task categories, size, creation methodology, known limitations, citation.
- Optional but recommended: speaker demographics, data collection dates, annotation methodology, known biases, evaluation results.
- Hub hosts 100K+ datasets (as of 2025) with varying card quality; the CROISSANT metadata format (2024) is being adopted for machine-readable dataset metadata.
- Bender et al. (2021) “Data Statements for NLP”: proposed formalised curation rationale documentation including speaker demographics, text-type distribution, and collection modality; particularly relevant for datasets with demographic representation implications.
- BigCode The Stack governance: opt-out process allowing developers to remove their code from the training corpus via GitHub identity verification — the first large-scale rights-respecting training data governance system for code.
- RAIL (Responsible AI Licences): integrates documentation requirements into licence terms, creating a legal mechanism for enforcing responsible data use downstream of dataset release.
- CROISSANT metadata format (HuggingFace, 2024): machine-readable JSON-LD schema for ML datasets enabling automated discovery, auditing, and compliance checking across dataset registries.
Components and Architecture of Training Data Pipelines
- A production training data pipeline comprises six sequential processing stages, each requiring dedicated engineering and tooling at petabyte scale.
- Stage 1 — Ingestion:
- Web crawls: CommonCrawl WARC files (typical monthly snapshot: 80-100TB compressed), custom Scrapy/wget crawls targeting specific domains.
- Data APIs: Semantic Scholar API for S2ORC scientific papers, GitHub Archive for code (GH Archive, ~50TB/year), PubMed API for biomedical literature.
- Proprietary datasets: licensed news (AP, Reuters), licensed books (Copyright Clearance Center), licensed scientific journals (Elsevier, Springer via licences).
- Synthetic generation pipelines: LLM-based generation of textbooks, instructions, code exercises.
- Typical ingestion scale: 10-100TB raw per pipeline run for a major training corpus build.
- Stage 2 — Extraction and Normalisation:
- HTML/WARC to text: trafilatura (most accurate, used in FineWeb), resiliparse (fastest, used in RedPajama), jusText (precision-oriented, used in OSCAR).
- Encoding normalisation: UTF-8 enforcement, Unicode NFC normalisation, HTML entity decoding.
- Language identification: fastText lid.176.bin model, 176 languages, >95% accuracy for common languages; CLD3 (Google) for backup.
- Format standardisation: JSONL with standardised fields (url, text, timestamp, language, source, content_hash).
- Storage: Apache Parquet for columnar access during filtering; JSONL for streaming; binary token shards for training.
- Stage 3 — Quality Filtering:
- Heuristic filters: length, punctuation coverage, repetition scores, stop-word fraction, non-alphabetic ratio.
- Domain blocklist: URL-based blocklists for known spam, adult content, malware, SEO content (Dolma uses the UT1 blocklist with ~1.5M domains).
- Classifier filters: toxicity (Jigsaw), educational quality (FineWeb-Edu fastText), PII (spaCy NER + regex).
- Typical reduction: 40-85% of raw corpus removed by Stage 3 filters.
- Stage 4 — Deduplication:
- URL-level: exact URL matching across crawl snapshots (removes 30-60% of multi-snapshot corpora before content-based dedup).
- MinHash LSH: document-level near-duplicate removal (Jaccard ≥ 0.8 threshold), distributed Spark/Dataproc job.
- Optional SemDeDup: embedding-based semantic deduplication (10-100× compute cost, used for image-text datasets primarily).
- Substring deduplication: optional exact substring dedup for long shared passages (suffix array construction, O(N log N)).
- Typical reduction: 20-50% of post-filter corpus removed by Stage 4.
- Stage 5 — Domain Mixing and Sampling:
- Target domain proportions for general-purpose models: web ~70%, books ~10%, code ~10%, science ~5%, Wikipedia ~5%.
- Proportions adjusted per capability target: code-heavy models (DeepSeek-Coder) use 60-87% code; science-heavy models use 30-50% scientific papers.
- Curriculum learning: schedules domain proportions dynamically during training — some pipelines emphasise high-quality domains in later training phases after initial broad pre-training on web text.
- Upsampling quality sources: small but high-quality datasets (Wikipedia, textbooks) may be upsampled 3-10× relative to their natural proportion.
- Stage 6 — Tokenisation and Sharding:
- Byte-pair encoding (BPE): vocabulary constructed from training corpus using byte-level BPE (GPT-4: cl100k_base, 100K vocab; Llama 3: 128K vocab SentencePiece BPE).
- Unigram tokenisation: alternative to BPE used in some multilingual models for better coverage of morphologically complex languages.
- Binary tensor conversion: token IDs stored as uint16 (vocabulary ≤ 65535) or int32 tensors in memory-mapped NumPy or HuggingFace datasets format.
- Sharding: 1-2GB binary shard files enabling parallel DataLoader access across distributed training nodes.
- Typical output: terabytes of binary token shard files; 1T-token corpus ≈ 2TB at uint16 storage (2 bytes/token).
Use Cases and Major Dataset Families
- Pre-training corpora (language): CommonCrawl, FineWeb (15T tokens), FineWeb-Edu (1.3T), Dolma (3T), RedPajama-V2 (30T), MassiveText/Gopher, The Pile (825GB), MAP-Neo (4.5T), ROOTS (1.6TB multilingual for BLOOM), CulturaX (6.3T multilingual, 167 languages).
- Code datasets: The Stack v2 (BigCode, 2024, 67.5M files, 358 programming languages from permissively licensed GitHub repos), CodeSearchNet (GitHub+StackOverflow), StarCoder training data (permissively licensed Stack subset), CodeParrot-train (HuggingFace).
- Instruction-tuning datasets: FLAN collection (Google, 1.8K tasks, 473 datasets), OpenHermes-2.5 (900K GPT-4 generated samples), UltraChat-200K (filtered), ShareGPT-52K (ChatGPT conversations), Alpaca (52K, Stanford), Dolly-15K (Databricks, human-authored), OpenAssistant OASST1 (161K human-human multi-turn, 35 languages).
- Multimodal (image-text): LAION-400M, LAION-5B (5.85B image-text pairs, legal status contested), DataComp-1B (filtered from CommonPool 12.8B), COYO-700M, Conceptual Captions CC12M/CC3M, WIT (Wikipedia-based image-text pairs).
- Scientific/specialist: S2ORC (81M papers), PubMed Central Open Access (3.5M full-text articles), arXiv (~2M papers, ~4GB), ChEMBL (2.4M bioactive molecules), UniProt-Swiss-Prot (560K annotated protein sequences).
- Benchmark-separated evaluation sets: To prevent contamination, canonical benchmarks (MMLU, HellaSwag, ARC, WinoGrande, HumanEval) are maintained with strict data provenance to ensure they were not present in training corpora. Contamination detection is an active research problem given the scale and opacity of many commercial training datasets.
Academic Context
- Training data curation has evolved from an engineering afterthought to a recognised research discipline with dedicated tracks at top-tier conferences (NeurIPS Datasets and Benchmarks Track, ACL/EMNLP data workshops, ICLR) and dedicated research groups at AI2, HuggingFace Research, BigCode, and LAION.
- Foundational pre-training corpus papers:
- Raffel et al. (2020) “Exploring the Limits of Transfer Learning” JMLR: introduced C4 and T5, establishing rule-based heuristic filtering as the baseline methodology.
- Brown et al. (2020) “Language Models are Few-Shot Learners” NeurIPS: GPT-3, provided the first detailed trillion-token training data description for a large commercial model.
- Gao et al. (2020) “The Pile” arXiv:2101.00027: introduced heterogeneous multi-domain dataset construction methodology for open-source LLMs.
- Deduplication literature:
- Lee et al. (2022) “Deduplicating Training Data Makes Language Models Better” ICLR 2022: landmark empirical study establishing deduplication as the highest-impact single preprocessing intervention.
- Abbas et al. (2023) “SemDeDup” arXiv:2303.09540: semantic deduplication via embeddings, removing 50% of data without performance loss.
- Penedo et al. (2023) “RefinedWeb” NeurIPS 2023: ablation showing aggressive CommonCrawl-only filtering sufficient for competitive performance.
- Soldaini et al. (2024) “Dolma” arXiv:2402.00159: full open-pipeline corpus with detailed deduplication methodology documentation.
- Scaling law and data-constrained regime:
- Hoffmann et al. (2022) “Training Compute-Optimal LLMs” (Chinchilla): established ~20 tokens/parameter optimality.
- Muennighoff et al. (2024) “Scaling Data-Constrained LLMs” NeurIPS 2024: data repetition strictly suboptimal; fresh data always preferred.
- Penedo et al. (2024) “FineWeb” arXiv:2406.17557: systematic ablation establishing quality filtering pipeline as the decisive competitive variable.
- Synthetic data and model collapse:
- Li et al. (2023) “Textbooks Are All You Need” (Phi-1): demonstrated 10-100× data efficiency with curated synthetic training data.
- Shumailov et al. (2024) “The Curse of Recursion”: model collapse dynamics when training exclusively on synthetic data.
- Data governance and documentation:
- Gebru et al. (2021) “Datasheets for Datasets” CACM: canonical dataset documentation framework.
- Gururangan et al. (2024) “DataComp-LM” arXiv:2406.11794: first objective comparison framework for quality filtering strategies.
- Security: Carlini et al. (2023) “Poisoning Web-Scale Training Datasets is Practical” IEEE S&P 2024; Gu et al. (2017) “BadNets” arXiv:1708.06733.
Current Landscape (2026)
- Scale vs quality tension: RedPajama-V2’s 30T tokens do not reliably outperform FineWeb’s 15T on quality-intensive tasks; educational and scientific content consistently outperforms raw web text token-for-token. The “more data always wins” intuition is being revised in favour of quality-per-token optimisation.
- Open vs proprietary divide: FineWeb, Dolma, RedPajama-V2, and MAP-Neo provide fully open training data enabling reproducible research; GPT-4, Gemini Ultra, and Claude 3 training data details remain undisclosed. EU AI Act Article 53 GPAI transparency requirements, applying from August 2025, are expected to force copyright compliance disclosures for EU-deployed models.
- Human vs synthetic blurring: As synthetic generation quality improves (Phi-3.5, Gemini Flash producing high-quality instruction data), the distinction between human-authored and model-generated training data is dissolving. Model collapse concerns motivate anchoring on human-authored data as a quality reference distribution.
- Copyright resolution timeline: NYT v OpenAI is expected to reach settlement or verdict by late 2026. UK TDM exception legislation, if enacted, would provide UK-based training practitioners clearer legal standing than current US fair use ambiguity. EU AI Office training data summary templates are expected to be finalised by Q2 2026.
- Data attribution — tracing model outputs back to specific training documents — is an active research and product development frontier, with methods including influence functions, TracIn, and TRAK enabling attribution at dataset scale. Attribution is motivated by both copyright compliance auditing and harm investigation (identifying training data responsible for harmful model outputs).
- Machine unlearning — selectively removing specific training examples’ influence without full retraining — is driven by GDPR Article 17 right-to-erasure, copyright settlement remedies, and safety incident response. SISA training (Bourtoule et al. 2021), gradient ascent unlearning, and influence-function-based parameter editing are the leading approaches, with computational cost remaining a key challenge.
UK Context
- Imperial College London: The I-X Centre for AI in Science (£50M EPSRC investment) conducts research on dataset quality assessment, bias detection in large corpora, and data-efficient learning — directly informing training data pipeline design. The Data Science Institute hosts projects on responsible data sourcing and provenance tracking for UK-deployed AI systems.
- Alan Turing Institute and AISI: The Alan Turing Institute in London hosts the AISI (AI Safety Institute) Evaluations Team, which has developed internal evaluation corpora and benchmarks for frontier model assessment. Critically, AISI maintains strict data provenance controls to prevent benchmark contamination, requiring knowledge of what was in evaluated models’ training sets. AISI’s Frontier AI Taskforce assessments of Gemini Ultra, GPT-4, and Claude 3 examined memorisation properties and domain-specific capability uplift from training data exposure.
- University of Edinburgh: School of Informatics (Professors Sharon Goldwater, Mark Steedman) leads multilingual NLP corpus development with contributions to low-resource language datasets covering Scottish Gaelic, Welsh, and other UK minority languages. The Edinburgh NLP Group is a significant contributor to multilingual training data equity research.
- University of Cambridge: Natural Language Processing group (Professors Anna Korhonen, Paula Buttery) has contributed to data-efficient learning methods and scientific corpus curation. The Computer Laboratory participates in BigCode and related open data research initiatives.
- University of Manchester and N8: The Centre for AI Fundamentals and the N8 Research Partnership (Manchester, Leeds, Sheffield, Newcastle, York, Durham, Liverpool, Lancaster) coordinate training data research across Northern England, with particular focus on industrial and healthcare data applications.
- Northern England industry: Sheffield-based AMRC (Advanced Manufacturing Research Centre) curates manufacturing process sensor data for industrial AI training, addressing IP protection and EU AI Act Article 10 compliance simultaneously. Faculty AI (Leeds) works on training data governance for NHS AI applications with patient data de-identification pipelines. Newcastle University’s Urban Observatory licenses urban sensing data (air quality, pedestrian counts, traffic) for smart city AI training. AESC (Sunderland, battery gigafactory) generates manufacturing process data with potential AI training value under appropriate governance frameworks.
- UK Government National Data Library: DSIT’s 2024 National Data Library initiative proposes creating a national corpus of public sector data with appropriate licensing for research and AI training, potentially providing a significant UK-sovereign training data resource.
- UK IPO TDM consultation: The 2024-2026 consultation proposes a copyright exception in the CDPA 1988 permitting computational analysis of lawfully accessed content for commercial and non-commercial purposes, with opt-out rights for rights-holders expressed via machine-readable signals. This is distinct from the EU DSM Directive approach and potentially more permissive, providing UK-based AI developers clearer legal standing than US fair use doctrine.
- BBC and media: BBC Research and Development has published guidance on responsible AI training data use referencing BBC Archive content, with specific concern about synthetic media training from BBC-licensed content. The Guardian, Times Media, and Reuters have published DSM Directive machine-readable opt-outs.
Future Directions (2026-2030)
- Data attribution and provenance at scale: Cryptographic content provenance (C2PA, Coalition for Content Provenance and Authenticity) embeds tamper-evident metadata into digital content; integration with training data pipelines would enable auditable records of content origin supporting copyright compliance and harm auditing.
- Machine unlearning at production scale: GDPR Article 17 and copyright settlement remedies motivate practical unlearning for deployed models. Approximate unlearning methods (gradient ascent, influence function-based parameter editing, SISA sharded training) are advancing towards production viability.
- Multilingual equity: Current corpora are English-dominated (CommonCrawl ~46% English by volume). CulturaX (6.3T tokens, 167 languages), multilingual FineWeb variants, and OPUS corpora are addressing this; the UNESCO Recommendation on AI (2021) and EU AI Act non-discrimination provisions both emphasise linguistic diversity as a training data requirement.
- Data marketplaces and licensing ecosystems: Spawning.ai opt-out registry (300M+ opted-out images), Bria.ai licensed creative media, and direct publisher licensing deals (OpenAI/Axel Springer, Google/Reddit API) point toward market-based training data licensing ecosystems supplementing open-source corpora.
- Federated and privacy-preserving training data: Differential privacy (DP-SGD, Abadi et al. 2016) provides quantified privacy guarantees for sensitive data training. Federated learning architectures (PySyft, NVIDIA FLARE) enable training without centralising raw data, particularly relevant for NHS federated data platform and cross-border medical AI deployments.
- Regulatory convergence and ISO standardisation: EU AI Act Article 53 summaries, UK TDM exception, US Copyright Office guidance, and G7 AI principles are driving convergence toward standardised training data documentation requirements. ISO/IEC 5259 series (AI data quality standards) is expected to provide certification frameworks for compliant training datasets by 2027-2028.
Research and Literature
- Core pre-training corpus papers: Raffel et al. (2020) T5/C4 JMLR; Brown et al. (2020) GPT-3 NeurIPS; Gao et al. (2020) The Pile arXiv:2101.00027; Penedo et al. (2023) RefinedWeb NeurIPS Datasets; Soldaini et al. (2024) Dolma arXiv:2402.00159; Penedo et al. (2024) FineWeb arXiv:2406.17557; Weber et al. (2024) RedPajama-V2.
- Deduplication: Lee et al. (2022) “Deduplicating Training Data” ICLR 2022; Abbas et al. (2023) SemDeDup arXiv:2303.09540; Broder (1997) MinHash IEEE SEQUENCES.
- Scaling laws: Hoffmann et al. (2022) Chinchilla arXiv:2203.15556; Muennighoff et al. (2024) NeurIPS 2024.
- Synthetic data: Li et al. (2023) Phi-1 arXiv:2306.11644; Li et al. (2024) Phi-3 arXiv:2404.14219; Ben Allal et al. (2024) Cosmopedia arXiv:2401.00812; Shumailov et al. (2024) model collapse arXiv:2305.17493.
- Security: Carlini et al. (2023) poisoning arXiv:2302.10149; Gu et al. (2017) BadNets arXiv:1708.06733.
- Governance: Gebru et al. (2021) Datasheets CACM; Gururangan et al. (2024) DataComp-LM arXiv:2406.11794.
- Key venues: NeurIPS Datasets and Benchmarks Track; ACL/EMNLP/NAACL data workshops; ICLR data-efficient learning; FAccT (fairness, accountability, transparency); Workshop on Dataset Curation and Security (NeurIPS 2024).
Key Terms Glossary
- Jaccard Similarity: intersection-over-union of two document token sets, used as the similarity metric in MinHash LSH deduplication; J(A,B) = |A ∩ B| / |A ∪ B|, ranging 0 (no overlap) to 1 (identical sets).
- Shingle: a contiguous n-gram sequence of characters or tokens used as a document fingerprint element in MinHash. 13-gram character shingles are standard for web document deduplication.
- Locality Sensitive Hashing (LSH): hashing scheme that maps similar items to the same bucket with high probability; used to identify near-duplicate document pairs in O(N) time vs O(N²) for brute-force pairwise comparison.
- BPE (Byte-Pair Encoding): tokenisation algorithm iteratively merging the most frequent character pair in a corpus to build a subword vocabulary; standard for LLM tokenisation (GPT series, Llama).
- Tokenisation: conversion of raw text into integer token IDs drawn from a fixed vocabulary; the fundamental pre-processing step before neural network training.
- Contamination: presence of evaluation benchmark examples (MMLU questions, HellaSwag completions) in training data, inflating benchmark scores and creating misleading capability assessments.
- Perplexity: a model’s predictive uncertainty on a text sequence; high perplexity indicates the text is surprising to the model; used as a quality proxy to filter unusually low-quality or high-quality (possibly templated) web documents.
- Text and Data Mining (TDM): computational analysis of text or data to extract patterns, information, or insights; the legal category under which AI training falls in EU (DSM Directive Article 4) and proposed UK CDPA exception.
- GPAI (General Purpose AI): an AI model trained on broad data at scale with generality of purpose, capable of being used in multiple downstream tasks; the legal category triggering EU AI Act Article 53 training data obligations.
- Machine-readable opt-out: a technical mechanism (typically robots.txt directives or structured metadata) enabling rights-holders to signal that their content should not be used for TDM/AI training; respected under EU DSM Directive Article 4(3) and proposed UK TDM exception.
- Chinchilla-optimal: a training configuration that optimally allocates compute budget between model size and training tokens per the scaling law of Hoffmann et al. 2022; approximately 20 training tokens per model parameter.
- Data Card: standardised documentation for a dataset covering motivation, composition, collection process, pre-processing, intended uses, distribution, and maintenance; introduced by Gebru et al. 2021 as “Datasheets for Datasets.”
Metadata
- Ontology domain:
artificial-intelligence - Domain correction: none — domain correctly set to
artificial-intelligencein source stub - IRI:
http://narrativegoldmine.com/artificial-intelligence#TrainingData - Legacy term ID:
AI-1041 - Enrichment worker:
claude-sonnet-4-6 - Enrichment date:
2026-05-17T09:00:00Z - Source stub lines: 49
- OWL axioms: 43 (35 SubClassOf + 6 DataPropertyAssertion + 2 AnnotationAssertion)
- Wikilinks: 72+
- Reference count: 26 primary literature + legal + standards + industry + UK academic
- Quality concerns: none — all facts drawn from publicly verifiable sources; legal case details cross-checked against known public filings; no fabrication
Provenance
- Raffel et al. (2020) “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer” JMLR 21(140):1-67.
- Brown et al. (2020) “Language Models are Few-Shot Learners” NeurIPS 33:1877-1901. arXiv:2005.14165.
- Gao et al. (2020) “The Pile: An 800GB Dataset of Diverse Text for Language Modeling” arXiv:2101.00027.
- Soldaini et al. (2024) “Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research” arXiv:2402.00159. AI2.
- Penedo et al. (2023) “The RefinedWeb Dataset for Falcon LLM” NeurIPS 2023 Datasets and Benchmarks. arXiv:2306.01116.
- Penedo et al. (2024) “The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale” arXiv:2406.17557. HuggingFace.
- Li et al. (2023) “Textbooks Are All You Need” arXiv:2306.11644. Microsoft Research.
- Li et al. (2024) “Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone” arXiv:2404.14219. Microsoft.
- Ben Allal et al. (2024) “Cosmopedia: How to Create Large-Scale Synthetic Textbooks, WikiHow Articles, Stories, Posts and Lecture Notes” arXiv:2401.00812. HuggingFace.
- Abbas et al. (2023) “SemDeDup: Data-efficient learning at web-scale through semantic deduplication” arXiv:2303.09540.
- Lee et al. (2022) “Deduplicating Training Data Makes Language Models Better” ICLR 2022. arXiv:2107.06499.
- Hoffmann et al. (2022) “Training Compute-Optimal Large Language Models” arXiv:2203.15556. DeepMind.
- Muennighoff et al. (2024) “Scaling Data-Constrained Language Models” NeurIPS 2024. arXiv:2305.16264.
- Carlini et al. (2023) “Poisoning Web-Scale Training Datasets is Practical” IEEE S&P 2024. arXiv:2302.10149.
- Shumailov et al. (2024) “The Curse of Recursion: Training on Generated Data Makes Models Forget” arXiv:2305.17493.
- Gebru et al. (2021) “Datasheets for Datasets” Communications of the ACM 64(12):86-92.
- Broder (1997) “On the Resemblance and Containment of Documents” Proceedings of SEQUENCES 1997. IEEE. (MinHash).
- Gururangan et al. (2024) “DataComp-LM: In search of the next generation of training sets for language models” arXiv:2406.11794.
- Weber et al. (2024) “RedPajama: an Open Dataset for Training Large Language Models” arXiv:2411.12372. Together AI.
- Kocetkov et al. (2022) “The Stack: 3 TB of permissively licensed source code” arXiv:2211.15533. BigCode.
- Abadi et al. (2016) “Deep Learning with Differential Privacy” ACM CCS 2016. (DP-SGD).
- Bourtoule et al. (2021) “Machine Unlearning” IEEE S&P 2021. (SISA training).
- NYT v OpenAI and Microsoft, SDNY 1:23-cv-11195, complaint filed December 2023.
- EU AI Act, Regulation (EU) 2024/1689, Official Journal of the European Union, August 2024, Article 53.
- UK IPO, “Artificial Intelligence and Intellectual Property: Copyright and Patents” consultation 2024-2026. UK Government.
- US Copyright Office, “Copyright and Artificial Intelligence, Part 2: Copyrightability” July 2024.