What matters in AI.

Subscribe

Learn / AI security

Guide · AI security

State of AI security, September 2026 edition

Most AI security writing is organised around where attacks could happen. This is organised around where they did. Seven pillars, each with what has actually been found and how worried I think you should be, which are not the same thing. Sourced claims carry a source. The worry ratings are mine.

As of September 2026

The board

Seven pillars of AI security. “Found” is what is showing up in the record. “Worry” is my call, and where the two disagree is the interesting part.

Pillar Found Worry
1. The gateway Constantly Highest
2. The glue around the sandbox Often High
3. Tools, skills and MCP servers Rising fast High
4. Agent identity Occasionally High, and structural
5. Guardrails, and their own scores Everywhere High, and nobody sells the fix
6. Attacker tempo Constantly Medium, it moves your clock
7. Model provenance Loudly Low, usually the wrong signal

Prompt injection is not on this list on purpose. See the end.

1. The gateway

The least interesting pillar and the one costing the most.

Found. Wiz scanned roughly 3,000 internet-facing LiteLLM deployments in February 2026. 294 of them accepted sk-1234, the example key printed in the project’s own setup guide. Another 191 wanted no authentication at all. A gateway holds an API key for every provider it routes to, sees every prompt passing through, and runs with the cloud permissions around it.

Do. Find your gateways. Rotate defaults. Management endpoints off the public internet. Workload permissions down to what it needs.

2. The glue around the sandbox

Isolation mostly holds. The wiring around it does not.

Found. The LiteLLM MCP bypass is the shape of it: the authentication design was fine, and a fallback path substituted an empty auth object when key validation failed, so a fabricated bearer token opened an authenticated session and could list and call every configured tool. CVSS 8.8, fixed in 1.84.0.

Do. Fail closed, and test the failure branch. If you cannot say what happens when auth errors rather than denies, you do not know what the control does.

3. Tools, skills and MCP servers

A supply chain that arrived complete, with no scanners and no norms.

Found. Invariant Labs showed in April 2025 that an MCP tool description can carry instructions the model obeys and the user never sees: the interface says “add two numbers”, the model is told to read SSH keys first. Their fix is to pin the tool definition, because a tool that can change after approval was never approved.

The precedent is older than AI. PyTorch’s nightly builds pulled torchtriton from PyPI ahead of the official index for five days in 2022, and the package read $HOME/.ssh/* and shipped it out over DNS. AI supply chain compromise is a software supply chain with more entry points and fewer tools.

Do. Pin tool definitions, not tool servers. Fail closed on a change. Keep an AI bill of materials so “are we affected” takes an hour.

4. Agent identity

The pillar with the least tooling and the worst foundations.

Found. An AI agent acts with somebody’s credentials, and we have no good way to say who that agent is or what it was authorised to do on whose behalf. That is what makes agent hijacking expensive rather than merely embarrassing: the attacker does not need to steal anything, only to redirect something that is already trusted. Existing delegation models were built for a user clicking approve, not for software that decides mid-task.

Do. Give the agent the permissions of a stranger, not of the engineer who built it. Separate identity per agent, short lived, scoped to one job.

5. Guardrails, and the scores they give themselves

Found. Willison’s operational test is the one worth keeping: an agent is exposed when private data, untrusted content and an outbound channel meet. His conclusion is to break the combination, and he says plainly he is “deeply suspicious” of products that claim to detect the attack instead.

Do. Use guardrails, do not stand on them. Write down what the system must never do and test that, rather than testing whether it behaves when asked politely.

6. Attacker tempo

Found. The public record now routinely describes intrusions where an agent performed the tactical steps, and exploit work measured in hours rather than weeks.

Do. Measure your time from advisory to patched, honestly. That number is now a security control.

7. Model provenance

The loudest pillar and the one I would spend the least on.

Found. Anthropic trained models from 600M to 13B parameters and found the number of poisoned documents needed for a backdoor attack stayed roughly constant rather than scaling with the data: 250 worked, 100 did not. The authors are explicit that the behaviour they installed was narrow and that it is unclear whether the same holds for complex behaviours such as backdooring code.

data poisoning is real, it does not scale away, and it is mostly not your attack surface unless you train.

Do. Pin model versions so you can tell when one changes under you. Defend the tuning set if you have one.

Why prompt injection is not a pillar

Because it is a mechanism, not a place. Prompt injection shows up in pillar 2 as a fallback path, in pillar 3 as a tool description, in pillar 4 as an agent using its credentials for someone else. Its blast radius is set entirely by pillars 1 through 4.

Ranking it as its own risk is how a team ends up buying detection for something whose only reliable control is capability. Fix what an injection can reach, and the injection stops being the interesting part.

The same goes for the system prompt: it is public, and anything that has to hold must hold outside the model.

If you do three things

  1. Inventory the gateways, agents and tool servers already running, and what credentials each holds.
  2. Rotate defaults, take management endpoints off the internet, scope the workload permissions.
  3. Find the agent that has private data, untrusted input and an outbound channel, and remove one of the three.

None of that needs a model expert, and all of it beats a guardrail purchase.

Questions and answers

What is the biggest AI security risk right now?

Exposed AI gateways holding provider credentials. In February 2026 Wiz scanned roughly 3,000 internet-facing LiteLLM deployments and found 294 accepted the example key from the project's own documentation. That is ordinary misconfiguration, not a model level attack, and it is where the losses are.

Is prompt injection the main AI security problem?

Prompt injection is a mechanism rather than a place. It shows up in the sandbox glue, the tool layer and the identity layer, and it is dangerous in proportion to what the agent can reach. Ranking it as a standalone risk leads teams to buy detection when the effective control is capability.

Do AI guardrail products work?

They raise attacker cost and they are not a boundary. A guardrail reads the wording of a request, not the authority behind it, and it is usually graded by its own vendor against the easy cases. Use one, do not stand on one.

Where should a security team start?

Find the AI gateways, agents and tool servers already running, and the credentials each one holds. Most organisations find more than they expected, and the fixes are ordinary ones they already know how to do.

Sources

  1. Off Guard, breaking LiteLLM from authentication bypass to cloud compromiseWiz
  2. GHSA-7488-6r32-c95q, MCP authentication bypassBerriAI, 30 Jun 2026
  3. MCP Security Notification, Tool Poisoning AttacksInvariant Labs, 1 Apr 2025
  4. Compromised PyTorch-nightly dependency chainPyTorch, 31 Dec 2022
  5. The lethal trifecta for AI agentsSimon Willison, 16 Jun 2025
  6. A small number of samples can poison LLMs of any sizeAnthropic
  7. NIST AI 100-2 E2025, Adversarial Machine Learning, A Taxonomy and Terminology of Attacks and MitigationsNational Institute of Standards and Technology

Terms in this guide

In the news