What matters in AI.

Subscribe

A new attack agent goes around the guardrails of 4 AI models

The researchers find a success rate of more than 95% for the attack.

Claimed, not confirmed

Researchers have made an attack agent that makes large language models give dangerous output. The name of the agent is CKA-Agent. It sends a group of queries, and each query is safe. The agent collects the responses and uses them to get the dangerous data. The researchers find a success rate of more than 95% on 4 commercial models, also against strong guardrails.

How it works

The researchers find that the knowledge in a model is connected. An attacker can use this fact. The agent sends small queries, and each query is safe.

The agent uses each response to select the next step in a tree of queries. At the end, it collects the data for the dangerous goal.

Why checks fail

Other attacks change the prompt, but the prompt keeps a dangerous signal. Guardrails are made to find that signal. In this attack, no query has the signal.

The test

The researchers did tests on these models:

  • Gemini2.5-Flash and Gemini2.5-Pro
  • GPT-oss-120B
  • Claude-Haiku-4.5

The success rate was more than 95%, also against strong guardrails.

What must change

The researchers find that defenses against this type of attack are necessary. They name it a knowledge-decomposition attack. Their code is available.

Sources

  1. The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Searcharxiv.org
AI MATTER · NEWS · AI MATTER · NEWS ·8 OCT2026

Posted

Tags

Learn the terms in this story

Guides for this story

More in Security

All Security news