A documentation framework proposed by Gebru et al. in which every machine learning dataset is accompanied by a structured datasheet recording its motivation, composition, collection process, preprocessing, recommended uses, distribution, and maintenance. Modelled on the datasheets that accompany electronic components, the practice surfaces provenance, consent, and bias considerations at the data layer, enabling informed dataset selection, reproducibility, and accountability across the machine learning lifecycle.

Semantic Classification

Content

Definition

Datasheets for Datasets is a documentation practice introduced by Timnit Gebru and colleagues in a 2018 preprint (published in Communications of the ACM, 2021). The proposal borrows an analogy from electronics: every component in the electronics industry ships with a datasheet stating its operating characteristics, test conditions, and recommended usage, yet the datasets on which machine learning systems depend routinely circulate with no equivalent record. The framework prescribes a structured questionnaire — roughly fifty questions organised around the dataset lifecycle — whose answers travel with the Dataset itself.

The questions are grouped into seven sections. Motivation asks why and by whom the dataset was created and who funded it. Composition covers what the instances represent, sampling strategy, label quality, presence of personal or sensitive data, and known errors or redundancies. Collection process documents how the data were acquired, over what timeframe, whether subjects consented, and what ethical review occurred. Preprocessing/cleaning/labelling records transformations applied and whether raw data were retained. Uses states tasks the dataset is suited for and — importantly — uses that would be inappropriate. Distribution and maintenance cover licensing, access, versioning, and points of contact. Answering these questions forces creators to confront consent, representation, and bias decisions while they can still be corrected, and gives consumers the information needed to judge fitness for purpose.

Within AI Documentation Standards, datasheets occupy the data layer of a documentation stack whose model layer is occupied by Model Cards: a model card characterises a trained model’s intended use and evaluated performance across conditions, whereas a datasheet characterises the corpus the model was trained or evaluated on. The two are complementary — a credible model card cites datasheets for its training data — and both feed system-level artefacts such as transparency reports and regulatory technical files.

Current Landscape

Datasheets became one of the most influential proposals in responsible AI, with tens of thousands of citations and direct descendants including Data Statements for NLP (Bender and Friedman), Data Nutrition Labels, Dataset Cards on Hugging Face (whose template explicitly incorporates datasheet questions), and Croissant, a machine-readable metadata vocabulary adopted by major dataset repositories. Conferences such as NeurIPS require dataset documentation in their datasets-and-benchmarks track, and the practice is embedded in corporate Data Governance pipelines at Google, Microsoft, and IBM in the form of internal data cards. Persistent challenges include documentation debt for web-scale crawled corpora, where answering composition and consent questions honestly is genuinely hard, incentives that reward dataset release speed over documentation quality, and keeping datasheets current as datasets are filtered, augmented, and merged downstream.

Regulation has turned the practice from a norm into an obligation:

Provenance