What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Lethal trifecta

The lethal trifecta is a name for the combination of three capabilities in one AI agent: access to private data, exposure to untrusted content, and the ability to communicate externally. Simon Willison coined the term in June 2025. An agent with all three can be tricked into sending its owner's private data to an attacker.

Last reviewed

Key points

  • The lethal trifecta is not a flaw in any one tool. It is what three ordinary capabilities do when a single AI agent holds all three at once.
  • The three legs are access to private data, exposure to untrusted content, and the ability to communicate externally.
  • Removing any one leg is enough to prevent the attack. Willison names exfiltration as the easiest leg to remove, and warns that exfiltration routes take sneaky shapes.
  • The trifecta is about stealing data. An agent whose tool calls can cause damage without leaking anything has, in Willison's words, "a whole other set of problems".
  • MCP is why no vendor can fix it. Mix-and-match tools push the security decision onto the end user, which Willison does not think is a reasonable thing to ask.

Simon Willison named the lethal trifecta in June 2025: “Access to your private data”, “Exposure to untrusted content”, and “The ability to externally communicate in a way that could be used to steal your data”. Each is useful alone. Together in one agent, they are a data breach waiting for a stranger to write in.

Prompt injection is the vulnerability underneath; the trifecta decides what a successful injection can take.

How the three legs combine

A language model follows instructions wherever it finds them. Willison’s explanation: LLMs “are unable to reliably distinguish the importance of instructions based on where they came from”.

Each leg then has a job. Untrusted content carries the attacker’s instruction — a web page, an email, a filed support ticket. Private data is what the instruction asks for. External communication is how the answer leaves, and that leg is wider than it looks: Willison notes that any tool able to make an HTTP request, load an image, or even offer the user a link can carry stolen data back.

Why removing one leg works

Removing any one of the three legs is enough to prevent the attack, which is what makes the framing worth having: it turns an unsolved research problem into a configuration you can check. Willison names exfiltration as the easiest leg to remove, warning that there are “all sorts of sneaky ways” a route out can take shape.

Detection is the alternative on offer. Guardrail products advertise catching “95% of attacks” or similar, to which Willison answers that “in web application security 95% is very much a failing grade”.

The scope is data theft. An agent that can cause damage without leaking anything — the territory of excessive agency and agent hijacking — is what Willison calls “a whole other set of problems to worry about”, and there exposure to malicious instructions alone can be enough.

In practice

Two of the reported cases put all three legs inside a single server. Willison notes that most lethal trifecta MCP attacks rely on users combining multiple MCPs; these did not.

Invariant Labs disclosed in May 2025 that the official GitHub MCP server let an attacker hijack a user’s agent through a malicious issue filed on a public repository, pull private repository data into context, and leak it in a pull request the agent opened itself against the public repo. The user’s own request was benign — look at the open issues. Willison’s reading: “That’s all three legs of the lethal trifecta!” Instructions arrive in public issues, the model reads private repos, and the pull request is the exfiltration channel.

In July 2025 he applied the same reading to the Supabase MCP, where an agent holding the service_role key reads customer-submitted support tickets. Supabase’s documentation recommends read-only, project-scoped mode, and Willison notes that this removes one leg of the trifecta — the ability to communicate data to the attacker, in this case through database writes. He still describes the remaining risk, with a read-only MCP against your database, as enormous.

Where definitions disagree

Meta published a variant in October 2025, the Agents Rule of Two, crediting Willison’s lethal trifecta and the similarly named Chromium policy as inspiration. It keeps the three-way shape and changes two things that matter.

Its third property is wider. Willison’s third leg is external communication. Meta’s is “An agent can change state or communicate externally”, so an agent that can delete a row without telling anyone has already spent leg three — a case Willison puts outside the trifecta and into the other set of problems.

Its prescription is softer. Willison tells users mixing their own tools to avoid the combination entirely. Meta says an agent that needs all three within one session “should not be permitted to operate autonomously and at a minimum requires supervision — via human-in-the-loop approval or another reliable means of validation”, and treats a one-way switch between configurations mid-session as legitimate.

The gap decides what a claimed defence covers. An agent audited as Rule-of-Two-compliant may hold all three of Willison’s legs under human review, and an agent that clears the lethal trifecta may still be free to do damage that never leaves the building.

Questions and answers

What are the three parts of the lethal trifecta?

Access to private data, exposure to untrusted content, and the ability to communicate externally. Simon Willison, who coined the term in June 2025, states them as access to your private data, exposure to untrusted content — any mechanism by which text or images controlled by a malicious attacker could reach your LLM — and the ability to externally communicate in a way that could be used to steal your data. An agent needs all three for the attack to complete.

Does removing one leg make an agent safe?

Removing one leg prevents this attack, which is data theft, and not every attack. Willison states that removing any one of the three legs is enough to prevent the attack, and names exfiltration as the easiest leg to remove while warning that exfiltration routes take all sorts of sneaky shapes. He also scopes the framing explicitly to stealing data: an agent whose tool calls can cause damage without leaking anything is a separate problem, and exposure to malicious instructions alone can be enough to cause it.

Is the lethal trifecta the same as prompt injection?

No. Prompt injection is the underlying vulnerability, and the lethal trifecta is the set of capabilities that decides what a successful injection can take. Willison coined both terms — prompt injection in September 2022, the lethal trifecta in June 2025 — and presents the trifecta as an example of the prompt injection class of attacks.

Why does MCP come up in every discussion of the lethal trifecta?

Because MCP is built on mix-and-match. Willison's objection is that encouraging users to combine whatever MCP servers they like outsources a critical security decision to those users, who must notice when the servers they enabled add up to all three legs. He does not think that is a reasonable thing to ask of end users. Two of the reported cases, the GitHub MCP server and the Supabase MCP, supplied all three legs from a single server.

Sources

  1. The lethal trifecta for AI agents: private data, untrusted content, and external communicationSimon Willison, 16 Jun 2025
  2. My Lethal Trifecta talk at the Bay Area AI Security MeetupSimon Willison, 9 Aug 2025
  3. GitHub MCP Exploited: Accessing private repositories via MCPInvariant Labs, 26 May 2025
  4. Supabase MCP can leak your entire SQL databaseSimon Willison, 6 Jul 2025
  5. Agents Rule of Two: A Practical Approach to AI Agent SecurityMeta, 31 Oct 2025

Guides that use this term