What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Four-layer framework

The four-layer framework is Anthropic's way of dividing an AI agent into the layers that decide what it can do — the model, the tools it can invoke, the harness that orchestrates tool use, and the execution environment that sets the security boundaries. Anthropic's point is that agent security is a property of the whole system, not the model.

Last reviewed

Key points

  • The four layers are the model, the tools it can invoke, the harness that orchestrates tool use across multiple steps, and the environment or container that defines the security boundaries.
  • Most security evaluation looks at the model layer. Anthropic argues that the harness and the execution environment determine what an agent can do when something goes wrong, not the model.
  • The same model failure costs different amounts at different layers. An injection inside a tight sandbox with narrow tool permissions ends differently from one with broad filesystem access and a credential store.
  • Controls at the four layers are not interchangeable and do not catch the same things. Resilient systems, Anthropic concludes, need different kinds of controls operating together, not more of any single kind.
  • This is a proposal, not a standard. Anthropic set it out in a March 2026 comment to NIST, and the same document names a product layer outside the four.

Anthropic’s four-layer framework starts from one claim: “Agent security is a property of the whole system, not just the model.” Its layers “determine what an agent can do in practice”.

What each layer holds

The model is the first layer of the AI agent. The tools it can invoke are the second, where least-privilege applies most directly, within a tool as much as across a tool set: “An agent that needs to read email does not necessarily need to read all email”.

The harness is the third, “the wrapper that lets a model run as an agent at all”, managing prompts, tool calls and context across turns; every action passes through it, which makes it the natural home for logging and monitoring. The execution environment is the fourth, “where the hard boundaries are set”: sandboxing, filesystem scope and network egress, the choke point for data leaving the system.

Why the model layer does not decide it

Most security evaluation concentrates on the model layer: is it robust to prompt injection? Not the question that decides the outcome, Anthropic argues. Sandboxed or not, the model fails the same way: “The failure is identical. The consequences are not.” The better question is not “can this model be compromised?” but “what is the scope of damage if it is?”

That reaches harm no model-layer test looks for. On Anthropic’s reading, NIST’s current frameworks for securing AI systems address two threat models, external attackers and deliberate misuse; AI as an offensive tool is a question NIST “treats elsewhere”. Within that framing, “neither threat model conceptualizes a well-functioning, non-compromised agent that operates inside its granted permissions but takes harmful actions for reasons unrelated to compromise or misuse”.

Where the framework is loose

It is a proposal in a policy comment, not a standard, and it is loose in its own terms. The threat-model conclusion widens without saying whose: “existing frameworks have no category for a failure where nothing malfunctioned” — a wider claim than the NIST-scoped one it follows. Nor does the document keep to four layers. It uses a “product layer” that is not on the list, and says many model and behavioural controls start at the harness or product level before they “get reinforced through training”, since where a control lives matters less than the behaviour it produces. The layers ask where a control sits and what it catches; they are not a taxonomy.

Questions and answers

What are the four layers?

Anthropic's four layers are the model and its underlying capability, the tools it can invoke, the harness that orchestrates tool use across multiple steps, and the environment or container that defines the security boundaries. Anthropic uses the shorter form "model, tools, harness, and execution environment" when referring back to the framework.

Is the four-layer framework a NIST standard?

No. The four-layer framework comes from Anthropic's 9 March 2026 response to a request for information issued by the NIST Center for AI Standards and Innovation, Docket No. NIST-2025-0035. It is one respondent's framing, submitted as input to future guidance rather than adopted guidance, and the response argues that developing terminology is work NIST has yet to do.

Is the four-layer framework a shared responsibility model?

No. Anthropic never calls it one, and the phrase "shared responsibility" does not appear anywhere in the document that introduces the four-layer framework. The framework divides an agent into layers where controls can sit, not into parties that owe each other duties. The only thing Anthropic calls shared is vocabulary — "The goal is not a bright-line test but shared vocabulary that lets developers, deployers, and researchers describe what went wrong in terms precise enough to compare across systems and build measurement around."

Sources

  1. Re: Request for Information: Security Considerations for Artificial Intelligence Agents, Docket No. NIST-2025-0035Anthropic, 9 Mar 2026