Definition · AI security
Physical adversarial example
A physical adversarial example is an adversarial example engineered to keep working through the physical world — printed, worn or projected onto a real object, then seen through a camera under real distance, angle and lighting — before the model reads it.
Last reviewed
Key points
- A physical adversarial example carries a survival requirement its digital parent lacks: the perturbation must keep working after being printed, worn or projected and then photographed under real distance, angle and lighting.
- The canonical instances are stop-sign posters and patches (Eykholt et al., CVPR 2018), eyeglass frames that defeat face recognition (Sharif et al., CCS 2016), a projected-light attack (SLAP, USENIX Security 2021) and adversarial T-shirts that hide a person from detectors (Xu et al., ECCV 2020).
- A survey of the field pinpoints manufacturing and re-sampling as what makes them distinct: the printer constrains colours and textures, and the camera re-samples the object.
- Whether physical adversarial examples actually survive is contested. Lu et al. found physical stop signs did not fool standard detectors, and an ICCV 2023 study found representative attacks caused no traffic-rule violation in closed-loop driving despite high component-level success.
How a physical adversarial example is made
A digital adversarial example is a vector added to an input the attacker hands straight to the model. A physical adversarial example has to survive what a printer and a camera add — the survey of the field names these two directly: manufacturing and re-sampling.
Manufacturing is what the printer does. Ink can only reproduce a limited set of colours, and the print is flat where the object was curved. Re-sampling is what the camera does next: the object is seen at a distance, from an angle, under whatever light is present, and through a lens and sensor that blur and darken what a pixel-perfect image never had to endure.
An attack is therefore not optimised against one image but against a distribution of distorted versions of it. The stop-sign attack of Eykholt and colleagues — the Robust Physical Perturbations (RP2) algorithm — was built exactly this way: black-and-white stickers on a real stop sign caused the target classifier to misread it in 100% of lab images and 84.8% of video frames captured from a moving vehicle.
Why it matters
The physical world is where models stop being demoed and start being trusted. A face-recognition system that admits the wrong person, a surveillance camera that stops seeing the intruder, a car that treats a stop sign as a speed limit — each is a model looking through a camera at the real world. The demand for adversarial example robustness is not academic here; the stakes are physical.
That is why the field keeps manufacturing objects rather than vectors. Printed eyeglass frames let the wearer evade recognition or impersonate someone else (Sharif et al.). A T-shirt printed with an adversarial pattern can hide the person wearing it from a person detector (Xu et al.). These are not attacks against a file; they live on a body.
In practice
The canonical instances each pick a different way to get the perturbation through the world:
- Stop-sign posters and patches. Eykholt and colleagues’ RP2 attack used black-and-white stickers on a real stop sign and reached targeted misclassification on a moving vehicle.
- Eyeglass frames. Sharif and colleagues printed eyeglass frames that let the wearer evade recognition or impersonate another individual in white-box and black-box scenarios.
- Projected light. SLAP (USENIX Security 2021) projects a perturbation onto an object with a light projector, making it adversarial only while illuminated. That makes the attack short-lived and gives the attacker control over timing; in non-bright settings it reached up to 99% misclassification success and evaded a physical-attack detection method over 95% of the time.
- Adversarial T-shirts. Xu and colleagues’ wearable hides the person from person detectors, and the design has to survive the shirt deforming as the wearer moves.
Where definitions disagree
Whether “survives the physical world” is a promise the field can keep is contested, and the definition above records the aim, not the outcome.
The strongest demonstrations attack a classifier on a carefully prepared image, and the counter-evidence targets exactly that gap. Lu and colleagues found that physical adversarial stop signs do not fool two standard detectors, YOLO and Faster R-CNN, in standard configuration; they argue the attacks relied on cropping the image to the sign and resizing it, which removes the very distortions the attack had to survive. An ICCV 2023 measurement study went further: the most effective attack it tested reached more than 70% average success against the perception component alone, yet none of the attacks caused a STOP sign traffic-rule violation in a closed-loop driving system at common real-world speeds for STOP sign roads. The question is not whether a physical adversarial example can be built; it is whether it can survive the world it is built for.
Questions and answers
What is the difference between an adversarial example and a physical adversarial example?
An adversarial example is an input altered so a model gets it wrong at deployment time; nothing about that definition says the input has to reach the model through the world. A physical adversarial example is the same idea with a survival requirement added: the alteration has to be printed, worn or projected onto a real object, photographed or filmed by a camera, and still mislead the model under real distance, angle and lighting. Whether that survival is achievable in practice is contested, and the field splits between demonstrations that fool a classifier on a cropped image and studies that find standard detectors and whole driving systems are not fooled.
Are adversarial patches the only kind of physical adversarial example?
No. Printed patches stuck to objects are the best-known kind, but the canonical examples also include eyeglass frames worn to defeat face recognition, a projector that briefly projects a perturbation onto an object so it becomes adversarial only while illuminated, and an adversarial T-shirt that hides the person wearing it from person detectors.
Do physical adversarial examples work against self-driving cars?
It depends on what "work" means. Against a component — a stop-sign classifier fed a cropped and resized image of a poster — attacks reach high success. Against the whole system it is contested: Lu et al. found physical stop signs did not fool YOLO or Faster R-CNN in standard configuration, and an ICCV 2023 study found representative attacks produced no traffic-rule violation in closed-loop driving despite high component-level attack success.
Sources
- Adversarial Examples in the Physical World: A SurveyarXiv, 1 Nov 2023
- Adversarial examples in the physical worldarXiv, 8 Jul 2016
- Robust Physical-World Attacks on Deep Learning ModelsarXiv, 27 Jul 2017
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionACM CCS, 24 Oct 2016
- SLAP: Improving Physical Adversarial Examples with Short-Lived Adversarial PerturbationsUSENIX Security, 11 Aug 2021
- Adversarial T-shirt! Evading Person Detectors in A Physical WorldarXiv; published in Computer Vision – ECCV 2020, Springer LNCS, doi:10.1007/978-3-030-58558-7_39, 24 Oct 2019
- Standard detectors aren't (currently) fooled by physical adversarial stop signsarXiv, 9 Oct 2017
- Does Physical Adversarial Example Really Matter to Autonomous Driving? Towards System-Level Effect of Adversarial Object Evasion AttackICCV, 23 Aug 2023