Definition · AI agents
Agent harness
An agent harness is the software that wraps a language model and makes it function as an agent — managing prompts, formatting tool calls, carrying context across turns, and running the execution loop. A model emits text; the harness turns that text into action. Anthropic places it as the third layer in the four-layer framework, between tools and execution environment.
Last reviewed
Key points
- The agent harness is the software layer that wraps a language model and turns it into an agent that can act.
- It manages context, formats tool calls, runs the execution loop, and enforces policies — not the model itself.
- Security controls live at the harness layer. Observability, permission checks, and audit logging all pass through it.
- Instruction privilege escalation happens here, not in the model, because the harness constructs the context the model sees.
- NIST calls the same component scaffolding software; Anthropic places it as the third layer in the four-layer framework.
How it works
An agent harness has four jobs:
-
Context construction. The harness builds the message the model sees each turn. It decides what goes into the system prompt, what counts as user input, and what comes from tools. This is the job instruction privilege escalation exploits: when the harness reconstructs context, it can promote content to a higher instruction level.
-
Tool formatting. When the model emits text shaped as a tool call, the harness translates that into a function call, sends it to the tool, and writes the result back into the context.
-
Loop management. The harness runs the turn cycle — feed the model context, read its output, decide whether to call a tool or stop. This turns a single-turn model into something that works through multi-step tasks.
-
Policy enforcement. Permission checks, tool access control, and audit logging sit at the harness layer.
The key insight from Anthropic’s four-layer framework is that these jobs matter more for security than the model’s capability. A model that falls for a prompt injection inside a tightly sandboxed harness produces a different outcome than the same model failing with broad filesystem access.
Why it matters
The harness is where security controls live and where attacks land.
It controls what the model can do. Least privilege applies at the tool layer, but the harness enforces it. When the harness grants too much access, the result is excessive agency.
It constructs the context the model trusts. The harness assigns instruction roles. When it reconstructs context without preserving provenance, it can promote attacker content into a trusted role — the root cause of instruction privilege escalation.
It is where observability lives. Every action passes through the harness, making it the natural place for logging and audit trails. Without harness-level logging, you cannot trace what happened.
It determines the scope of damage. Deployers often evaluate model-layer security without evaluating the harness. Anthropic argues the more productive question is not “can this model be compromised?” but “what is the scope of damage if it is?”
In practice
The privilege escalation paper tested six coding-agent harnesses — Claude Code, Codex, Gemini CLI, Qwen Code, Kimi, and OpenCode. All six were vulnerable because all six reconstruct context during delegation without preserving where content came from.
OWASP’s Agent Control Standard addresses this boundary. ACS hooks sit between the agent framework and the tool layer — the same boundary the harness occupies — and enforce tool-call policy at runtime.
NIST calls the same component scaffolding software. Their definition of an AI agent system requires “at least one generative AI model and scaffolding software,” which is the harness by another name.
Trade-offs
The harness limits the model. A tightly constrained harness prevents dangerous actions but also limits what the agent can accomplish. The trade-off is between safety and capability.
Harness security is harder to evaluate than model security. Model behavior can be tested with benchmarks and red-teaming. Harness behavior depends on the specific implementation — how it constructs context, how it enforces permissions, how it handles delegation.
Portability is limited. Unlike models, which can be swapped between systems, harnesses are tightly coupled to their implementation.
Sources
- Anthropic. “Re: Request for Information: Security Considerations for Artificial Intelligence Agents, Docket No. NIST-2025-0035.” March 9, 2026.
- Li et al. “When Context Gets Root: Privilege Escalation in LLM Harnesses.” arXiv, August 27, 2026.
- OWASP. “Agent Control Standard (ACS).” OWASP GenAI Security Project, September 1, 2026.
Questions and answers
What does an agent harness do?
An agent harness manages the execution loop of an AI agent. It constructs prompts, formats tool calls, carries context between turns, and enforces policies about what the model is allowed to do. Every action an agent takes passes through the harness.
Is the harness the same as the model?
No. The model decides; the harness acts on those decisions. The model is a language model that predicts the next token. The harness is the software that feeds it inputs, interprets its outputs, calls tools, and runs the loop that keeps the agent going until the task is complete.
Why does the harness matter for security?
The harness is where observability and verification happen — logging, hooks, and audit trails that record what the agent did and why. It is also where permission checks and policy enforcement sit. When an agent does something dangerous, the harness is either the place that stopped it or the place that failed to.
How does the harness relate to the four-layer framework?
Anthropic's four-layer framework places the harness as the third layer, between the tools an agent can use and the execution environment that sets security boundaries. The harness orchestrates tool use across steps, while the environment defines the outer limit of what the agent can affect.
What is instruction privilege escalation?
Instruction privilege escalation is an attack where the harness promotes attacker-controlled content from a low-privilege role, like a tool result, into a higher-privilege role, like a user message. The model then follows it as authorized instruction. The failure is in the harness's context construction, not in the model's reasoning.