A new attack agent goes around the guardrails of 4 AI models
The researchers find a success rate of more than 95% for the attack.
Claimed, not confirmed
Researchers have made an attack agent that makes large language models give dangerous output. The name of the agent is CKA-Agent. It sends a group of queries, and each query is safe. The agent collects the responses and uses them to get the dangerous data. The researchers find a success rate of more than 95% on 4 commercial models, also against strong guardrails.
How it works
The researchers find that the knowledge in a model is connected. An attacker can use this fact. The agent sends small queries, and each query is safe.
The agent uses each response to select the next step in a tree of queries. At the end, it collects the data for the dangerous goal.
Why checks fail
Other attacks change the prompt, but the prompt keeps a dangerous signal. Guardrails are made to find that signal. In this attack, no query has the signal.
The test
The researchers did tests on these models:
- Gemini2.5-Flash and Gemini2.5-Pro
- GPT-oss-120B
- Claude-Haiku-4.5
The success rate was more than 95%, also against strong guardrails.
What must change
The researchers find that defenses against this type of attack are necessary. They name it a knowledge-decomposition attack. Their code is available.
Sources
Posted