6 lines of log text give a guardrail a fail-open rate of 63%
Because a deterministic rule was 100% correct on all 6 policies, the model is not necessary.
Claimed, not confirmed
The researchers did tests on 7 open-weight models as guardrails in agent systems. A guardrail reads a proposed tool call or message and decides if the system does the task or stops it. On tests to find attacks and toxic content, the models were correct in 36% to 72% of the decisions. Chance level is 50%. The researchers say that these models can decrease the number of items for a reviewer. They say that a model must not be the part that decides.
How they tested
The researchers used typed decision models. A model of this type reads text and gives a likelihood for each alternative.
They counted 2 types of error. A fail-open error is a decision to do a prohibited task. This is a vulnerability.
A fail-closed error stops a permitted task. This is only a cost.
The attacks
The researchers added lines of server log text to the input. The fail-open rate of a gate went to 63% on a policy that the model decides correctly without the lines.
In a different attack, they gave the permissive alternative an incorrect name. The rate went to between 93% and 100% on the models that put the name in their input.
What did not help
The researchers say that an attacker can defeat each defense in their tests. To send the decisions with low confidence to a reviewer does not help.
What works
The researchers parse each policy field into a typed value. This stops one attack. A deterministic rule for those values is 100% correct on all policies in their tests. Thus the model is not necessary.
Sources
Posted