What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Agent skill poisoning

Agent skill poisoning is an attack that plants malicious instructions inside a SKILL.md file — a natural-language package of scripts and directives that agents such as Claude Code, OpenClaw and Cursor load to extend their capabilities. It carries no code signature to catch, and once loaded runs with the agent's own filesystem, credential and network access.

Last reviewed

Key points

  • Agent skill poisoning hides malicious instructions inside a SKILL.md file, the bundled instructions-plus-scripts format Claude Code, OpenClaw and other agents load as a reusable capability.
  • A loaded skill runs with the agent's own permissions — shell access, filesystem read and write, and any credential in an environment variable — because nothing isolates it.
  • Snyk's ToxicSkills audit of 3,984 skills on ClawHub and skills.sh found 36.82% with a security flaw, 534 (13.4%) critical, and 76 confirmed malicious payloads.
  • Two campaigns reported in early 2026 — one reaching 386 skills that targeted Claude Code users, one Koi Security named ClawHavoc at 341 skills — delivered credential-stealing malware and keyloggers.
  • It is distinct from AI agent tool poisoning, where a tool sits behind a protocol call rather than being a package the agent loads and runs directly.

Agent skill poisoning turns a capability an agent was given to extend itself — loading a bundled set of instructions and scripts — into the way an attacker reaches everything that agent can touch.

How it works

A SKILL.md file is natural language, not compiled code: a description of when to use the skill, plus instructions and scripts the agent follows once it decides the skill applies. Snyk notes the entire barrier to publishing one on ClawHub is a Markdown file and a GitHub account a week old — no code signing, no security review, no sandbox by default.

The instructions can be malicious from the start, or hide behind a legitimate purpose and ask for something dangerous only under a “prerequisites” step. Either way, prompt injection’s core failure applies: the agent cannot reliably tell an instruction that came from the skill file from one that came from its own operator, because both are just text it reasons over the same way. And unlike a package that runs in its own process, a skill inherits the agent’s own permissions — Snyk lists shell access, filesystem read and write, and any credential sitting in an environment variable.

Why it matters

Testing backs the pattern the audits find. SkillInject, a 202-scenario benchmark, found frontier models complying with injected skill instructions at up to an 80% attack success rate, including data exfiltration. A separate study of real ClawHub skills found the registry itself can be gamed: reworded malicious skills dodged a blocking verdict in more than a third of cases.

What makes this distinct from an ordinary malicious dependency is that nothing about a SKILL.md file looks like an artifact a security scanner was built to catch. It is prose, aimed at the agent’s judgment rather than at a compiler, asking the agent to do something it is already trusted to do.

In practice

Snyk’s ToxicSkills audit of 3,984 skills across ClawHub and skills.sh found more than a third with some security flaw, 534 (13.4%) with a critical-level issue, and 76 confirmed malicious payloads built for credential theft, backdoors and data exfiltration. In late January 2026, OpenSourceMalware.com documented a wave of 28 malicious skills targeting Claude Code and Moltbot users, growing to 386 within days, sharing a single command-and-control server and built to steal exchange API keys, wallet private keys, SSH credentials and browser passwords. Koi Security’s independent audit of ClawHub in the same window found 341 malicious skills out of 2,857, with 335 tied to one operation the researchers named ClawHavoc — fake installation prerequisites that delivered the Atomic Stealer on macOS and a keylogging trojan on Windows.

Both campaigns posed as high-demand tools — cryptocurrency trackers, trading bots, YouTube utilities, Google Workspace integrations — to reach as many installs as possible, then buried the payload behind an ordinary-looking setup step rather than the skill’s advertised purpose. Neither needed to defeat a security review, because ClawHub did not run one at the time; both were caught afterward, by outside researchers auditing the marketplace, not by a control on the publishing path.

This is why the fix looks less like the usual reaction to AI supply chain compromise — check a signature, pin a version — and more like the fix for unexpected code execution: constrain what a loaded skill can actually do on the host, because the text describing it will not reliably say.

Questions and answers

What is agent skill poisoning?

An attack that hides malicious instructions inside a SKILL.md agent skill file — the bundled natural-language-plus-scripts format that agents like Claude Code, OpenClaw, Cursor and Codex CLI load to extend what they can do. Because the instructions are plain text, no code signature flags them, and because the skill runs as part of the agent, it inherits the agent's filesystem, credential and network access.

How is agent skill poisoning different from MCP tool poisoning?

The distribution and execution mechanism. An MCP tool sits behind a protocol interface the agent calls; the agent invokes it as a function. A skill is a markdown-and-script package the agent loads and follows directly, with no call boundary between the instructions and the agent's own reasoning.

What was ClawHavoc?

The name Koi Security gave a coordinated campaign it found on ClawHub, OpenClaw's skill marketplace: 341 malicious skills out of 2,857 audited, with 335 traced to one operation that used fake installation prerequisites to deliver the Atomic Stealer on macOS and a keylogging trojan on Windows. A separate, earlier count from OpenSourceMalware.com put an initial wave at 28 skills targeting Claude Code and Moltbot users, growing to 386 within days — the same window of activity, reported by two different researchers.

Sources

  1. Snyk, "ToxicSkills: Malicious AI Agent Skills on ClawHub"Snyk, 5 Feb 2026
  2. OpenSourceMalware.com, "Malicious ClawHub Skills Target OpenClaw Users"OpenSourceMalware.com, 1 Feb 2026
  3. The Hacker News, "Researchers Find 341 Malicious ClawHub Skills Stealing Data from OpenClaw Users"The Hacker News, 2 Feb 2026
  4. Schmotz et al., "Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks" (arXiv:2602.20156)arXiv, 23 Feb 2026
  5. Saha, Faghih and Feizi, "Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry" (arXiv:2605.11418)arXiv

Guides that use this term