A test shows that the top detector stopped 51% of agent attacks
A check of the source of text found attacks that detectors do not find, but it stopped all user tasks.
Claimed, not confirmed
A developer made a test of 10 open-source detectors against 629 attacks from AgentDojo, a research benchmark from ETH Zürich. Each attack was in tool output of the usual type. The top detector stopped 51% of the attacks and gave incorrect alarms on 2% of 97 clean examples. The developer then made a small tool that does not read the text of the attacker. The tool found attacks that detectors do not find. It also stopped all the correct tasks of a user.
Sources
Posted