What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Denial of AI service

Denial of AI service is an attack that degrades or shuts down an AI system by exhausting the compute it needs to answer, either with a flood of requests or with inputs crafted to make each answer expensive. MITRE ATLAS catalogues denial of AI service as technique AML.T0029, under the Impact tactic.

Last reviewed

Key points

  • Denial of AI service attacks availability by exhausting the compute an AI system needs to answer. ATLAS notes that this compute often makes AI systems expensive bottlenecks.
  • There are two routes. Send a flood of requests, or send a few requests crafted so that each one is expensive to answer.
  • The expensive route turns each request into far more work. Crafted sponge inputs often made language models 30 times slower and hungrier for energy in the lab.
  • Rate limits raise the cost of a flood but may not help against an attack that needs few requests. ATLAS pairs them with caps on the resources each request may use, and with red-team tests that try to exhaust the service.
  • The same pressure that denies service also runs up the bill. ATLAS files that motive separately as cost harvesting; OWASP's 2025 list puts denial of service, cost and model theft in one risk, Unbounded Consumption.

Denial of AI service degrades or shuts down an AI system by using up the compute it needs. MITRE ATLAS files it as AML.T0029: adversaries “may target AI-enabled systems with a flood of requests for the purpose of degrading or shutting down the service.” Many AI systems need “significant amounts of specialized compute”, so “they are often expensive bottlenecks that can become overloaded.”

How it works

The first route is volume: a flood of requests, as in the ATLAS description.

The second route is cost per request. ATLAS notes that attackers “can intentionally craft inputs that require heavy amounts of useless compute”. Shumailov and colleagues called these sponge examples. On language models they “frequently increase both latency and energy consumption” by a factor of 30; on vision models the effect was smaller. Against Microsoft’s Azure translator, 50-character sponge inputs found by trial and error pushed response times up to 6 seconds, against about 1 millisecond for ordinary text.

Why it matters

When the service slows, other users wait. OWASP’s 2023/24 entry on this attack says it degrades service for the attacker “and other users”, and may bring “high resource costs”. ATLAS files the billing motive separately as cost harvesting, and notes that the same pressure can trigger autoscaling, which amplifies the cost.

The defences follow the two routes. ATLAS says rate limits raise the cost of overwhelming a service, but “may not protect against attacks that require few requests”. For those, ATLAS maps a cap on “the resources consumed by individual requests”, and the sponge paper proposes a threshold so that a sponge example “will simply result in an error message”. ATLAS also lists AI red teaming that submits “adversarial workloads”, then sets quotas, timeouts and limits based on what it finds.

In practice

An app can be talked into it. In an ATLAS case study, MathGPT turned questions into Python and ran the code. A prompt injection told it to “compute forever”. The model wrote a loop that never ended, and the app hung running it until its host restarted.

Where definitions disagree

ATLAS splits by motive: AML.T0029 for availability, AML.T0034 Cost Harvesting for the bill. OWASP’s LLM entries do not split them. Its 2023/24 LLM04 Model Denial of Service already covered cost, and its 2025 list puts denial of service, denial of wallet (running up a victim’s pay-per-use bill) and model theft together in LLM10 Unbounded Consumption. LLM04 now means Data and Model Poisoning, and the old LLM04 address redirects there, so older writing that cites “OWASP LLM04” for this attack points readers at poisoning.

ATLAS also applies AML.T0029 more widely than its own description. The AIKatz case study maps it to an attacker using tokens lifted from a victim’s desktop app to delete their chats, and to spamming the model until its bot limits ban the victim. That denies service to one person, not to everyone sharing the system.

Questions and answers

What is denial of AI service?

Denial of AI service is an attack that degrades or shuts down an AI system by exhausting the compute it needs to answer. MITRE ATLAS catalogues it as AML.T0029. The attacker either floods the service with requests or sends inputs crafted to make each answer expensive, and legitimate users wait or get nothing.

How is denial of AI service different from an ordinary DoS attack?

The resource being exhausted is the AI system's compute rather than network bandwidth. ATLAS notes that many AI systems need significant amounts of specialised compute, which often makes them expensive bottlenecks. The sponge-examples researchers describe their attack as targeting the most expensive parts of a translation service to get an amplification factor from crafted requests, instead of overwhelming bandwidth as most DoS attacks do.

Do rate limits stop denial of AI service?

Not on their own. ATLAS says query limits can increase the time and cost needed to overwhelm a service, but notes that query limits may not protect against attacks that require few requests, and maps a separate mitigation, AML.M0036, that caps the resources each request may consume. The sponge-examples researchers propose the same idea as a cut-off threshold on the time or energy one inference may use.

Is OWASP LLM04 still Model Denial of Service?

No. LLM04 was Model Denial of Service in the 2023/24 OWASP Top 10 for LLM Applications. In the 2025 list LLM04 is Data and Model Poisoning, and denial of service is covered by LLM10:2025 Unbounded Consumption, alongside denial of wallet, which runs up a provider's pay-per-use bill, and model theft. The old LLM04 address redirects to the poisoning entry.

Sources

  1. MITRE ATLAS, AML.T0029 Denial of AI Service (collection 2026.09)MITRE
  2. Sponge Examples: Energy-Latency Attacks on Neural Networks, version 2 (published at IEEE EuroS&P 2021)arXiv, 5 Jun 2020
  3. LLM04: Model Denial of Service, OWASP Top 10 for LLM Applications v1.1 (archived in the project repository)OWASP
  4. LLM10:2025 Unbounded Consumption, OWASP Top 10 for LLM ApplicationsOWASP
  5. LLM04:2025 Data and Model Poisoning, OWASP Top 10 for LLM ApplicationsOWASP