What matters in AI.

Subscribe

A test shows that the top detector stopped 51% of agent attacks

A check of the source of text found attacks that detectors do not find, but it stopped all user tasks.

Claimed, not confirmed

A developer made a test of 10 open-source detectors against 629 attacks from AgentDojo, a research benchmark from ETH Zürich. Each attack was in tool output of the usual type. The top detector stopped 51% of the attacks and gave incorrect alarms on 2% of 97 clean examples. The developer then made a small tool that does not read the text of the attacker. The tool found attacks that detectors do not find. It also stopped all the correct tasks of a user.

Sources

  1. Testing Provenance-Based Controls Against Indirect Prompt Injection in AI AgentsHackerNoon
AI MATTER · NEWS · AI MATTER · NEWS ·11 OCT2026

Posted

Learn the terms in this story

Guides for this story

More in Security

All Security news