A documentation framework proposed by Gebru et al. in which every machine learning dataset is accompanied by a structured datasheet recording its motivation, composition, collection process, preprocessing, recommended uses, distribution, and maintenance. Modelled on the datasheets that accompany electronic components, the practice surfaces provenance, consent, and bias considerations at the data layer, enabling informed dataset selection, reproducibility, and accountability across the machine learning lifecycle.
Semantic Classification
Content
Definition
Datasheets for Datasets is a documentation practice introduced by Timnit Gebru and colleagues in a 2018 preprint (published in Communications of the ACM, 2021). The proposal borrows an analogy from electronics: every component in the electronics industry ships with a datasheet stating its operating characteristics, test conditions, and recommended usage, yet the datasets on which machine learning systems depend routinely circulate with no equivalent record. The framework prescribes a structured questionnaire — roughly fifty questions organised around the dataset lifecycle — whose answers travel with the Dataset itself.
The questions are grouped into seven sections. Motivation asks why and by whom the dataset was created and who funded it. Composition covers what the instances represent, sampling strategy, label quality, presence of personal or sensitive data, and known errors or redundancies. Collection process documents how the data were acquired, over what timeframe, whether subjects consented, and what ethical review occurred. Preprocessing/cleaning/labelling records transformations applied and whether raw data were retained. Uses states tasks the dataset is suited for and — importantly — uses that would be inappropriate. Distribution and maintenance cover licensing, access, versioning, and points of contact. Answering these questions forces creators to confront consent, representation, and bias decisions while they can still be corrected, and gives consumers the information needed to judge fitness for purpose.
Within AI Documentation Standards, datasheets occupy the data layer of a documentation stack whose model layer is occupied by Model Cards: a model card characterises a trained model’s intended use and evaluated performance across conditions, whereas a datasheet characterises the corpus the model was trained or evaluated on. The two are complementary — a credible model card cites datasheets for its training data — and both feed system-level artefacts such as transparency reports and regulatory technical files.
Current Landscape
Datasheets became one of the most influential proposals in responsible AI, with tens of thousands of citations and direct descendants including Data Statements for NLP (Bender and Friedman), Data Nutrition Labels, Dataset Cards on Hugging Face (whose template explicitly incorporates datasheet questions), and Croissant, a machine-readable metadata vocabulary adopted by major dataset repositories. Conferences such as NeurIPS require dataset documentation in their datasets-and-benchmarks track, and the practice is embedded in corporate Data Governance pipelines at Google, Microsoft, and IBM in the form of internal data cards. Persistent challenges include documentation debt for web-scale crawled corpora, where answering composition and consent questions honestly is genuinely hard, incentives that reward dataset release speed over documentation quality, and keeping datasheets current as datasets are filtered, augmented, and merged downstream.
Regulation has turned the practice from a norm into an obligation:
-
Datasheets named in EU law: Annex IV(2)(d) of the EU AI Act explicitly refers to datasheets for training datasets as part of the technical documentation required for high-risk AI systems, alongside data-governance documentation duties under Articles 10, 11, and 13
-
Mandatory training-content summaries: on 24 July 2025 the Commission’s AI Office published the compulsory template under Article 53(1)(d) for general-purpose AI providers to summarise training content (data sources, scraped domains, licensing, synthetic data); the obligation applies from 2 August 2025, models already on the market must comply by 2 August 2027, and AI Office enforcement powers (fines up to 3% of worldwide turnover or €15 million) apply from 2 August 2026
-
Machine-readable convergence: the MLCommons Croissant format (JSON-LD on schema.org/Dataset) operationalises datasheet content for pipelines — auto-generated for Hugging Face Hub datasets convertible to Parquet and supported by Kaggle, OpenML, and Google Dataset Search; guidance increasingly recommends shipping a human-readable datasheet alongside a Croissant descriptor and a PROV-O lineage record
Sources:
-
https://practical-ai-act.eu/latest/engineering-practice/data-governance/documentation/