Agents
AgentMemGate keeps plans out of the memory of AI assistants
The authors tell that, with the gate, 95% of tasks are correct, not 65%.
arxiv.orgClaimed, not confirmed
What matters in AI.
SubscribeTag
47 stories carry this tag.
Agents
The authors tell that, with the gate, 95% of tasks are correct, not 65%.
arxiv.orgClaimed, not confirmed
Security
The authors tell that the monitor also gave an alarm for 29.3% of safe runs.
arxiv.orgClaimed, not confirmed
Security
The authors put malicious prompts in data that looks safe, and the attack works only after a world model uses the data.
arxiv.orgClaimed, not confirmed
Security
The authors found a false-positive rate of 0.1%, and the system limits the model to a closed set of commands.
arxiv.orgClaimed, not confirmed
Research
The authors found 12.3 times more collisions in real-time tests than in static tests. The agents completed 91% to 94% of the tasks.
arxiv.orgClaimed, not confirmed
Research
The authors found a typical error of 2.9 times for Fable 5.1 in Claude Code, but only 1.2 times for GPT-6 Astra in Codex.
arxiv.orgClaimed, not confirmed
Security
The researchers find a success rate of more than 95% for the attack.
arxiv.orgClaimed, not confirmed
Research
2 easy methods to add diversity did not decrease the risk.
arxiv.orgClaimed, not confirmed
Research
The researchers find that this reporting destroys 68% of the gains from delegation.
arxiv.orgClaimed, not confirmed
Research
An instruction to randomize independently decreased the correlation between agents with the same input, but did not remove it.
arxiv.orgClaimed, not confirmed
Security
The attack can also get around 3 agent-level defenses.
arxiv.orgClaimed, not confirmed