What matters in AI.

Subscribe

Learn / AI governance

Definition · AI governance

Datasheets for datasets

A datasheet for a dataset is a document published alongside a machine learning dataset that records how it was made, covering its motivation, composition, collection process, preprocessing, recommended uses, distribution and maintenance. Gebru et al. proposed the format in 2018 as a set of questions for dataset creators to answer, not a schema to validate against.

Last reviewed

Key points

  • A datasheet for a dataset is fifty-seven questions in seven sections — motivation, composition, collection process, preprocessing, uses, distribution and maintenance — answered in prose by the people who built the dataset.
  • Gebru et al. say the process of creating a datasheet "is not intended to be automated", because what they are after is the creator's reflection rather than the file.
  • Datasheets and model cards split one job. Mitchell et al. say Datasheets highlight "characteristics of the data feeding into the model", while a model card reports the trained model.
  • Questions about people sit at the end of each section under a deliberately wide test — any dataset containing text that was written by people relates to people.
  • NeurIPS named datasheets among its recommended documentation frameworks from 2021 to 2024. Its 2025 call asks instead for a Croissant machine-readable metadata file.

A datasheet for a dataset is a document published with a machine learning dataset that answers a fixed set of questions about how it was built. Gebru et al. circulated the proposal on 23 March 2018 and took the name from electronics, where every component ships with a datasheet giving its “operating characteristics, test results, recommended usage”. A model card does the equivalent job for the trained model that comes out.

How it works

Fifty-seven questions, in seven sections matching “the key stages of the dataset lifecycle”: motivation, composition, collection process, preprocessing, uses, distribution, maintenance. Seven of them are the same closing prompt, “Any other comments?”.

The order is a schedule. Composition and collection questions are read “prior to any data collection” and answered afterwards; distribution and maintenance before the dataset ships.

Questions about people sit at the end of a section rather than mixed in, under a wide test: “any dataset containing text that was written by people relates to people”. Those cover consent, notification, identifiability and sensitive data.

None of them asks whether the dataset is lawful. A team of lawyers reviewed the draft, after which the authors “removed questions that explicitly asked about compliance with regulations”, replacing them with factual ones that avoid “requiring dataset creators to make legal judgments”.

Why it matters

A datasheet has two readers who want different things. For creators the objective is “careful reflection on the process of creating, distributing, and maintaining a dataset”. For consumers it is to “make informed decisions about using a dataset”.

The first objective is why the format resists tooling. Automated documentation is convenient, the authors write, but it would “run counter to our objective of encouraging dataset creators to carefully reflect”. The artifact is a by-product; the work is someone having to write down who was paid what to label this.

In practice

By 2021 Microsoft, Google and IBM were piloting datasheets internally, Google had released a data card, “a lightweight version of a datasheet”, with the Open Images dataset, IBM had proposed factsheets for AI services, and the Data Nutrition Project had folded some of the questions into its Dataset Nutrition Label.

The clearest adoption was NeurIPS. From 2021 its Datasets and Benchmarks track required “dataset documentation and intended uses” in the supplementary materials and listed “datasheets for datasets” first among recommended frameworks — through 2022, 2023 and 2024. The 2025 call drops the list. It names no documentation framework at all, and asks instead that “authors should use the Croissant machine-readable format to document their datasets”, a file generated for you if the data sits on Hugging Face, Kaggle, OpenML or Dataverse.

That is the tension in the proposal, resolved by a conference in favour of the machine. The datasheet asked a person to reflect. The replacement is emitted automatically.

Law reaches it twice, and only one of the two binds. Annex IV of the EU AI Act is the technical documentation a high-risk system must draw up under Article 11(1), which “shall contain, at a minimum, the elements set out in Annex IV”. Point 2(d) asks for “where relevant, the data requirements in terms of datasheets describing the training methodologies and techniques and the training data sets used”, then lists provenance, scope, main characteristics, how the data was obtained and selected, labelling and cleaning — Gebru’s composition, collection process and preprocessing sections under another name, though the Act never defines the word and “where relevant” carries real weight. The other mention is recital 89, which only encourages developers of open-source components “other than general-purpose AI models” to adopt “model cards and data sheets”.

Trade-offs

A datasheet is an account, not an audit. Nothing verifies it. The composition section asks whether external resources have “guarantees that they will exist, and remain constant, over time”, and takes a sentence for an answer; AI dataset provenance answers that question with a hash and a modification record. Neither substitutes for the other. A hash cannot say who consented.

The authors name the rest of the limits themselves. Creators “cannot anticipate every possible use of a dataset”. Identifying unwanted bias “often requires additional labels indicating demographic information about individuals”, which data protection may put out of reach — so the document meant to surface bias can be blocked from doing it by privacy law. The questions “may pose problems for dynamic datasets”, where the recommendation for anything changing infrequently is a new datasheet per version, and nothing is offered for data that changes constantly. And writing one “will necessarily impose overhead”, which is the honest reason most datasets still ship without one.

Questions and answers

What goes in a datasheet for a dataset?

Fifty-seven questions in seven sections, ordered to the dataset lifecycle. Motivation asks why the dataset was built, by whom and on whose money. Composition asks what an instance is, how many there are, what is missing, and whether the data is self-contained or depends on external resources. Collection process asks how the data was acquired, who collected it and how they were paid. Preprocessing asks what was done to the raw data and whether the raw data was kept. Uses asks what the dataset has been used for and, pointedly, "Are there tasks for which the dataset should not be used?". Distribution asks about licensing and restrictions, and maintenance asks who will host it, whether it will be updated and how older versions will be supported. Seven of the fifty-seven questions are the same closing prompt, "Any other comments?".

What is the difference between a datasheet for a dataset and a model card?

What each one documents. Mitchell et al., who proposed model cards a year later and count Timnit Gebru among their authors, put it as "Where Datasheets highlight characteristics of the data feeding into the model, we focus on trained model characteristics such as the type of model, intended use cases, information about attributes for which model performance may vary, and measures of model performance." They describe model cards as complements to Datasheets rather than a replacement, so a trained model released responsibly has both, and the datasheet is the one that can be written before any model exists.

Is a datasheet for a dataset required by law?

For high-risk AI systems in the EU, a qualified yes. Article 11(1) of the EU AI Act says the technical documentation of such a system "shall contain, at a minimum, the elements set out in Annex IV", and Annex IV point 2(d) asks for "where relevant, the data requirements in terms of datasheets describing the training methodologies and techniques and the training data sets used", then spells out provenance, scope, main characteristics, how the data was obtained and selected, labelling and cleaning. What is owed is that content. The Act never defines the term or cites Gebru et al., and "where relevant" qualifies the duty. Outside that, it is voluntary — recital 89 says developers of free and open-source "tools, services, processes, or AI components other than general-purpose AI models" should be "encouraged to implement widely adopted documentation practices, such as model cards and data sheets". That is a recital rather than a binding article, the verb is encouragement, and the population it names expressly excludes general-purpose AI models. The authors are consistent with the voluntary reading of their own proposal, stating that the questions "are not intended to be prescriptive".

Does a datasheet prove anything about the data?

No. A datasheet is an account given by the people who built the dataset, and nothing in the proposal verifies it. The gap shows most clearly in the composition section, which asks whether a dataset relies on external resources and, if it does, whether there are "guarantees that they will exist, and remain constant, over time" — then takes a sentence for an answer. AI dataset provenance answers that same question mechanically, with a cryptographic hash binding each entry to the content a downloader retrieves and a record of every modification since. The two are complementary. A hash cannot tell you who was asked for consent, and a datasheet cannot tell you whether the bytes changed.

Sources

  1. Datasheets for Datasets, Section 1Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford (Communications of the ACM 64(12), December 2021, pages 86 to 92, doi 10.1145/3458723), 1 Dec 2021
  2. Model Cards for Model ReportingMitchell, Wu, Zaldivar, Barnes, Vasserman, Hutchinson, Spitzer, Raji and Gebru (FAT* 2019), 14 Jan 2019
  3. NeurIPS 2021 Call for Datasets and Benchmarks (retrieved 2026-09-14; the 2022, 2023 and 2024 calls carry the same sentence)NeurIPS
  4. NeurIPS 2024 Call for Datasets and Benchmarks (retrieved 2026-09-14)NeurIPS
  5. NeurIPS 2025 Call for Datasets and Benchmarks (retrieved 2026-09-14)NeurIPS
  6. Regulation (EU) 2024/1689 (EU AI Act), recital 89Official Journal of the European Union, 12 Jul 2024