Definition · AI security
AI output watermarking
AI output watermarking is the practice of hiding a machine-detectable signal inside content an AI model generates, so that the content can later be traced to that model. For text, the signal is usually a statistical bias in word choice that readers cannot see. AI output watermarking detects and attributes content after the fact; it does not prevent misuse.
Last reviewed
Key points
- AI output watermarking hides a signal in what a model generates, so software can recognise the output later. For text, the model is nudged towards a shifting set of words chosen by a hidden rule.
- A watermark is for detection, not prevention. It shows afterwards that content probably came from a given model. OWASP recommends it to detect unauthorised use of a model's outputs.
- Watermarks can be removed and forged. NIST reports that rewording by a separate, non-watermarked model often strips a text watermark with only minor loss of quality.
- Text is harder to watermark than images, NIST says, because it offers less room for a signal and small edits dilute it. Predictable answers are the worst case.
- The EU AI Act requires providers of generative AI systems, with some exceptions, to mark outputs as machine-detectable. Its recitals name watermarks as one way to do it.
How it works
A scheme for text published in 2023 by Kirchenbauer and colleagues at the University of Maryland works one word at a time. Before each word, a hidden rule uses the previous word to mark part of the vocabulary as “green”. The model is nudged towards green words, not forced, so the text still reads normally.
A person writing without the rule uses green words at the rate chance predicts. Watermarked text uses them far more often. The detector counts green words and works out how unlikely that count is by chance. It needs the rule, not the model.
Google DeepMind’s SynthID-Text also changes only how the model picks words, and has been used to watermark Gemini.
A watermark is not a label in the file’s metadata. NIST notes that metadata is often stripped when files are shared, and generally cannot travel with plain text between documents. A watermark sits in the words.
Why it matters
Watermarking has two uses that matter here. The first is labelling synthetic content. The EU AI Act requires providers of systems that generate audio, images, video or text to mark outputs “in a machine-readable format and detectable as artificially generated or manipulated”, with exceptions such as tools that only assist standard editing. The duty covers deepfakes and ordinary text alike. The Act’s recitals name watermarks as one technique among several.
The second is attribution. OWASP’s entry on unbounded consumption, which covers model extraction, recommends watermarking “to embed and detect unauthorized use of LLM outputs”. A watermark cannot stop a competitor harvesting a model’s answers, but it may show afterwards that they did.
What breaks a watermark
Predictable text. NIST regards text as significantly harder to watermark than images, and says text watermarks generally cannot be embedded or detected reliably when there are few plausible responses. Its example: it would be hard to tell whether the obvious answer to “1 + 1 =” came from a model.
Paraphrasing. NIST reports that having a separate, non-watermarked model reword the text often removes the watermark with only minor loss of quality. Repeated paraphrasing can cut detection to 20% on texts of about 225 words, at the cost of more quality loss. In practical settings, NIST adds, paraphrasing reduces detection only slightly on texts beyond about 400 words.
Stealing and forgery. Researchers at ETH Zurich studied watermarks that work by shifting word probabilities. By querying a watermarked model, at a one-time cost below $50 at ChatGPT’s January 2024 prices, they approximated its rules. They could then both generate new text carrying the watermark and scrub it from real output, with average success above 80%. NIST warns that forged watermarks on harmful content could expose the model’s maker to reputational or legal risk.
Coverage. The SynthID-Text authors note that watermarks only work when the services generating text choose to apply them, and are hard to enforce on open-source models run by anyone.
Tracing a copied model
A watermark can outlive the text it was in. Researchers at Meta, École polytechnique and Inria found that a model fine-tuned on a watermarked model’s outputs carries a detectable trace. That is the case that matters for model extraction.
The trace is faint in the realistic case. With API-only access to the suspect model and no record of which outputs it was trained on, their test reached its strict confidence threshold when at least 10% of the training data was watermarked. With less, detection was weaker.
Can a watermark ever be robust?
The question is open. Zhang and colleagues at Harvard and elsewhere call a watermark strong if an attacker with bounded computing power cannot erase it without significant loss of quality. They proved that “strong watermarking is impossible to achieve” under assumptions they argue can be satisfied in practice. Their attack removed the watermarks of three published text schemes with only minor quality loss.
NIST accepts that any text watermark can in principle be defeated, but notes the proofs rest on assumptions about the attacker, about what counts as removal and about how outputs are constrained. It concludes “there may be circumstances where the proofs do not hold in practice”.
Questions and answers
Can an AI watermark be removed?
Yes, often. NIST reports that having a separate, non-watermarked model paraphrase watermarked text often removes the watermark with only minor loss of quality, especially from short texts. Longer texts hold the signal better, but NIST notes that an attacker who knows which words the watermark prefers can remove it even from those.
Does the EU AI Act require watermarking?
The EU AI Act requires providers of generative AI systems, with some exceptions, to mark outputs in a machine-readable format so they are detectable as AI-generated. The Act's recital 133 lists watermarks as one possible technique, alongside metadata, cryptographic provenance, logging, fingerprints and other techniques, so a watermark specifically is not mandated.
Is a watermark the same as provenance metadata?
No. A watermark is encoded in the content itself, such as the words or pixels. Provenance metadata is stored in the file or in a linked repository, is often stripped when files are shared, and generally cannot follow plain text copied between documents.
Sources
- LLM10:2025 Unbounded Consumption, OWASP Top 10 for LLM ApplicationsOWASP GenAI Security Project
- A Watermark for Large Language Models (Kirchenbauer, Geiping, Wen, Katz, Miers, Goldstein)University of Maryland (arXiv; ICML 2023), Jan 2023
- Scalable watermarking for identifying large language model outputs (Dathathri et al.)Nature (Google DeepMind authors), 23 Oct 2024
- NIST AI 100-4, Reducing Risks Posed by Synthetic ContentNIST, Nov 2024
- Watermark Stealing in Large Language Models (Jovanović, Staab, Vechev)ETH Zurich (arXiv; ICML 2024), Feb 2024
- Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models (Zhang, Edelman, Francati, Venturi, Ateniese, Barak)Harvard University and others (arXiv), Nov 2023
- Watermarking Makes Language Models Radioactive (Sander, Fernandez, Durmus, Douze, Furon)Meta FAIR, École polytechnique and Inria (arXiv; NeurIPS 2024), Feb 2024
- Regulation (EU) 2024/1689 (AI Act), Articles 50(2) and 111(4), consolidated text 27 July 2026European Union
- Regulation (EU) 2024/1689 (AI Act), recital 133European Union, 12 Jul 2024