Definition · AI security
Deepfake detection
Deepfake detection is a defence that checks images, video, audio or text for signs that AI generated or manipulated them, so synthetic media can be flagged or blocked before a person or system trusts it. MITRE ATLAS lists it as mitigation AML.M0034, for untrusted or user-provided data, especially in high-impact uses such as biometric verification.
Last reviewed
Key points
- MITRE ATLAS lists deepfake detection as mitigation AML.M0034 against generating deepfakes, phishing and evading AI models, to be applied to untrusted or user-provided data.
- ATLAS says detectors may combine approaches, including models trained to tell real from fake, checks for flaws such as unnatural facial movement, and biometric cues such as blinking.
- NIST calls automated detection "a constant cat-and-mouse game". Generators improve as soon as a new detector appears, and detectors are often tied to specific generators and may only perform well on those.
- Academic benchmarks can overstate how detectors perform. On media circulating in 2024, one study found the open-source detectors it tested lost 45 to 50 percent of their AUC, a measure of how well they separate fake from real, on average.
- NIST warns that in many contexts flagging real content as AI-generated "can be extremely damaging", so a detector's false positives count, not only the fakes it misses.
How it works
A deepfake detector judges whether a piece of media is a deepfake, and often gives a score for how likely it is to be synthetic rather than a plain yes or no. MITRE ATLAS says detectors may combine several approaches, including:
- Trained classifiers. A model trained on examples of real and fake content learns to tell them apart.
- Inconsistencies. Checks for flaws that generation tends to leave, such as unnatural facial movement, audio that does not match the lips, or artifacts in the pixels.
- Biometric cues. Analysis of blinking, eye movement and microexpressions.
NIST’s overview of synthetic content describes three kinds of detection, which can be combined. Content-based detection reads the media for traces left when it was generated or processed, as the ATLAS approaches do. Provenance data detection looks for a record of where content came from, such as a watermark or metadata. Human-assisted detection has people check or add to what the tools find.
NIST notes a complication for every method: media can be partly synthetic, such as a real photo with one object removed and the gap filled in by AI.
Why it matters
ATLAS maps deepfake detection against three attacks: generating deepfakes, phishing that uses generated content, and model evasion, which includes presenting a faked face or voice to a system as genuine. It says to apply detection to untrusted or user-provided data, especially in high-impact uses such as biometric verification.
The biometric case shows the stakes. In an ATLAS case study, the iProov red team swapped a victim’s face into a live video and streamed it to a phone in place of its camera. It got past facial recognition and the liveness checks meant to confirm a real person is in front of the camera.
How well detectors hold up
NIST describes automated detection as “a constant cat-and-mouse game”: as soon as a new detector appears, generators improve and attackers learn new ways to avoid it. NIST adds that detectors “are often tied to and may only perform well on specific generators”. For images, NIST says detectors show substantial error rates on generators they were not trained on, particularly after compression and resizing. One 2023 paper it cites reported 61 to 70 percent accuracy in that setting. For audio, it says most research focuses on English voices.
Deepfake-Eval-2024 measured how detectors do outside the lab. Its authors collected media that social media and TrueMedia.org users flagged as suspect in 2024, labelled each item real or fake, then tested open-source detectors that score well on academic datasets. The measure was AUC: the chance that a detector scores a random fake as more suspicious than a random authentic item, where 1 is perfect. The detectors’ average AUC fell by 50 percent for video, 48 percent for audio and 45 percent for images. The authors conclude that academic benchmarks are out of date and do not represent real-world deepfakes.
Fine-tuning the detectors on part of the new data helped, and the best commercial detectors beat off-the-shelf open-source ones. None of the commercial models the authors evaluated reached 90 percent accuracy, the authors’ lower-bound estimate for human deepfake forensic analysts.
Mistakes in the other direction matter too. NIST warns that in many contexts, flagging real content as AI-generated “can be extremely damaging, potentially resulting in major reputational harms or adverse treatment”. That is why some studies report how many fakes a detector catches while its false alarm rate is held very low, rather than metrics such as accuracy or AUC that balance the two kinds of error.
Questions and answers
How accurate are deepfake detection tools?
It depends heavily on what they are tested on. The Deepfake-Eval-2024 study took open-source detectors that score well on the academic datasets they were benchmarked on and tested them on media that circulated in 2024. Their AUC fell by 45 to 50 percent on average. The best commercial detectors it evaluated beat the off-the-shelf open-source ones, but none reached 90 percent accuracy.
Can deepfake detection stop deepfake phishing?
It helps but is not enough on its own. MITRE ATLAS lists two mitigations for deepfake-assisted phishing: deepfake detection and user training. Its user training entry also recommends checking a sensitive request through an independent channel, such as a known call-back number. NIST warns that detectors are often tied to specific generators and may only perform well on those, so a fake made with a new tool can pass.
What is the difference between deepfake detection and watermarking?
A watermark is provenance data: a record of the content's history that a detector can look for. NIST counts checking for watermarks or metadata as one kind of synthetic content detection, called provenance data detection. Content-based detection instead reads the media for traces left during generation or processing. The approaches MITRE ATLAS lists for its deepfake detection mitigation all read the content itself.
Sources
- MITRE ATLAS, mitigation AML.M0034 Deepfake Detection (collection 2026.09)MITRE
- Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency (NIST AI 100-4)NIST, Nov 2024
- Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024arXiv (Chandra et al., TrueMedia.org and University of Washington), Mar 2025