Definition · AI security
Web-scale data poisoning
Web-scale data poisoning is an attack on training sets that are distributed as lists of URLs rather than as content. Because the bytes behind a URL can change after the dataset was reviewed, an attacker who buys the right domain, or times an edit to land in the next snapshot, chooses part of what the model learns.
Last reviewed
Key points
- Web-scale datasets ship as an index of URLs rather than as data, so whoever owns a domain at download time decides what that part of the corpus says.
- Split-view poisoning buys expired domains a dataset index still points at, so an annotator's view of the dataset differs from the view later downloaders get.
- Frontrunning poisoning targets datasets built from periodic snapshots of crowd-sourced content such as Wikipedia, where the attacker needs only a time-limited window.
- Carlini and colleagues priced control of 0.01 percent of LAION-400M or COYO-700M at 60 US dollars, an annual domain cost at July 2023 prices.
- MITRE ATLAS types the case study as an Exercise. The researchers bought six domains, served 404 from all of them, and found no evidence of the attack in the wild.
A web-scale dataset is usually not distributed as data. It is distributed as an index: a list of URLs, fetched by each downloader later. Index and content are held by different parties, so what a model trains on is not necessarily what anyone reviewed.
What an attacker buys
Carlini and colleagues published two attacks on that gap in 2023. Split-view poisoning “exploits the mutable nature of internet content to ensure a dataset annotator’s initial view of the dataset differs from the view downloaded by subsequent clients” — in practice, buying expired domains the index still points at. Frontrunning poisoning targets datasets that “periodically snapshot crowdsourced content—such as Wikipedia—where an attacker only needs a time-limited window to inject malicious examples”. A moderator who reverts the edit afterwards does not remove it from the snapshot.
Sixty dollars, with its population and its date
For 60 US dollars, Carlini and colleagues could have poisoned “0.01% of the LAION-400M or COYO-700M datasets in 2023” — corpora of 408 million and 747 million images, so roughly 41,000 images in one and 75,000 in the other, at an annual domain price quoted in July 2023. That is a percentage of an image corpus, not a count of examples, so it does not compare with the poison-example counts reported against instruction-tuning sets.
NIST’s defence for this case is unusually definite: “the provider publishes cryptographic hashes, and the downloader verifies the training data.” Almost nothing else in data poisoning gives a yes or no answer. Binding an index to its content is what AI dataset provenance is for.
Where definitions disagree
The word covers a defence as well as an attack, and settles neither. NIST records that “recently published open-source data poisoning tools increase the risk of large-scale attacks on image training data” while noting they were “created to enable artists to protect the copyright of their work”. Same technique, same corpora, opposite intent — and a crawler cannot read intent.
Questions and answers
How much does it cost to poison a web-scale dataset?
Carlini and colleagues put one number on it in 2023, and it needs its population and its date to mean anything. For 60 US dollars they could have controlled 0.01 percent of LAION-400M or COYO-700M — two corpora of different size, which Table 1 of the paper gives as 408 million and 747 million images, so roughly 41,000 images in the first and 75,000 in the second. The 60 dollars is an annual domain registration cost priced against Google Domains in July 2023. A larger budget buys more. With 1,000 US dollars, they report control of between 0.02 and 0.79 percent of the images in each of the 10 datasets they studied. Older datasets are cheaper, because more of their domains have expired.
Has web-scale data poisoning ever happened in the wild?
No case is known. MITRE ATLAS types its case study AML.CS0025 as an Exercise rather than an incident, and the researchers behind it went looking. They compared CC3M and LAION-400M images against their original versions for the signature of a domain-purchasing attack and report "we could not find any evidence of this"; the one CC3M domain that matched turned out to be a squatter serving advertisements. That is a statement about the past, not about difficulty. The attack was practical when they measured it, which is why six of the 10 datasets they disclosed to now publish integrity checks.
What is the difference between split-view and frontrunning poisoning?
Split-view poisoning attacks the space between an index and its content. The attacker buys a domain the dataset index still points at, so what a downloader retrieves is not what the annotator reviewed. Frontrunning poisoning attacks the timing of a snapshot instead. The dataset is built by periodically copying crowd-sourced content such as Wikipedia, and an attacker who can predict the snapshot only needs the malicious edit to be live for that moment. Reverting the edit afterwards does not help, because the snapshot already has it.