What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Agents Rule of Two

The Agents Rule of Two is a design rule for AI agents, published by Meta in October 2025. Within one session, an agent may hold at most two of three properties: processing untrustworthy inputs, accessing sensitive systems or private data, and changing state or communicating externally. An agent needing all three in one session should be supervised, not autonomous.

Last reviewed

Key points

  • The Agents Rule of Two limits an AI agent to two of three properties per session. They are untrustworthy inputs, sensitive systems or private data, and changing state or communicating externally.
  • Meta published it in October 2025 because prompt injection is unsolved. The rule holds until prompt injection can be reliably detected and refused.
  • Meta credits Simon Willison's lethal trifecta and a Chromium policy as inspiration. Its third property is wider than the trifecta's, because changing state counts, not only communicating externally.
  • An agent needing all three without a fresh session should not run autonomously. Meta says it at a minimum requires supervision, through human approval or another reliable check.
  • Meta says the rule is not sufficient on its own. Critics dispute whether any two properties are always lower risk, and whether counting properties is the right test.

The Agents Rule of Two is Meta’s answer to prompt injection, which nobody has solved. It limits what an injected instruction can make an AI agent reach or do, “until robustness research allows us to reliably detect and refuse prompt injection”.

The three properties

Within one session, meaning one context window, an agent should hold at most two of:

  • [A] “An agent can process untrustworthy inputs”
  • [B] “An agent can have access to sensitive systems or private data”
  • [C] “An agent can change state or communicate externally”, such as making a booking or sending an email

When an agent needs all three without starting a new session, Meta says it should not run autonomously and needs “human-in-the-loop approval or another reliable means of validation”.

In Meta’s example, a spam email tells an email assistant to forward the user’s inbox to the attacker. The attack works because the agent reads strangers’ mail [A], sees the inbox [B], and can send email [C]. Remove any one and it fails: trusted senders only, no private data, or sending only to trusted recipients or after a human checks the draft.

An agent may switch configuration mid-session. Meta’s one-way example starts on the internet [AC], then turns off communication before touching internal systems [B]. The test is whether an attack can complete the chain from A to B to C.

Why it matters

The Agents Rule of Two matters because defences against prompt injection have not proved reliable. In The Attacker Moves Second, which Simon Willison cites, attacks that kept iterating bypassed 12 recent defences against jailbreaks and prompt injection, most above 90% of the time. Most of those defences had reported near-zero attack success. Willison calls the Rule of Two “the best practical advice” in the absence of defences that can be relied on.

Trade-offs

Each configuration costs something. Meta frames the choice in terms of “user friction or limits on capabilities”, and expects some use cases to fit badly, such as “a background process where human-in-the-loop is disruptive or ineffective”.

Meta says satisfying the rule is not sufficient against other agent threats, such as agent mistakes, hallucinations or excessive privileges. A compliant design can still fail when a user confirms a warning screen without reading it. And the rule is “a supplement — and not a substitute” for least privilege.

Where definitions disagree

Against the lethal trifecta. Meta credits Willison’s lethal trifecta as inspiration, but the third property is wider. Willison’s third leg is external communication, because his trifecta describes data theft. Meta’s adds changing state: “An agent can change state or communicate externally”. Willison welcomed this: the trifecta “only covers the risk of data exfiltration”, and changing state brings other tool use into the picture.

Whether two properties are safe. Meta’s first diagram labelled the pairing of untrustworthy inputs and the ability to change state “safe”. Willison objected that, even “without access to private systems or sensitive data”, that pairing “can still produce harmful results”. Mick Ayzenberg, one of the post’s authors, agreed the concern was valid and said the diagram had been updated; the intersections now read “lower risk”. The rule, he wrote, is not meant to describe “a sufficient level of security” but “a minimum bar that’s needed to deterministically prevent the highest security impacts of prompt injection”. In a later reply he added that [B] covers any sensitive system. An agent without [B] can act freely, but not on any system that matters for its user’s critical security outcomes. His example is an agent in a tight sandbox or isolated from production.

Whether counting properties is enough. Jack Poller, writing in Security Boulevard in May 2026 about research by the vendor Noma Security, lists “trusted data as attack vector” among Noma’s five toxic patterns. It hides malicious payloads inside “the authoritative data the agent was designed to trust”, which the article says collapses “the Rule of Two’s core assumption”. Drawing on separate incidents, Poller argues that the rule “measures the wrong variable”: it counts risk properties, when the question is blast radius, “how much damage that agent can land when something goes wrong”.

Questions and answers

What are the three properties in the Agents Rule of Two?

Meta labels them A, B and C. [A] the agent can process untrustworthy inputs. [B] the agent can access sensitive systems or private data. [C] the agent can change state or communicate externally. Within one session an agent should hold no more than two of them.

How is the Agents Rule of Two different from the lethal trifecta?

The third property is wider. Simon Willison's lethal trifecta ends in external communication, because it describes data theft. Meta's Rule of Two ends in changing state or communicating externally, so an agent that can delete or modify something already counts. Willison welcomed that change, writing that the trifecta only covers data exfiltration.

Is an agent that follows the Rule of Two secure?

Not on its own. Meta says the rule is not sufficient against other agent threats, that compliant designs can still fail when a user confirms a warning without reading it, and that it supplements rather than replaces least privilege. One of its authors calls it a minimum bar.

Sources

  1. Agents Rule of Two: A Practical Approach to AI Agent SecurityMeta, 31 Oct 2025
  2. New prompt injection papers: Agents Rule of Two and The Attacker Moves SecondSimon Willison, 2 Nov 2025
  3. Comment on "New prompt injection papers" discussionHacker News, 3 Nov 2025
  4. Comment on "New prompt injection papers" discussionHacker News, 3 Nov 2025
  5. The Half of Agent Security You're Not GoverningSecurity Boulevard, 4 May 2026