Definition · AI security
AI security trade-offs
AI security trade-offs are the costs an organisation accepts when it hardens a machine learning system against attack, such as losing accuracy or explainability to gain adversarial robustness. NIST states it may not be possible to simultaneously maximise performance across the trustworthy-AI attributes, so the organisation must accept trade-offs and decide which to prioritise.
Last reviewed
Key points
- Hardening is not free. NIST says all proposed evasion mitigations 'exhibit inherent trade-offs between robustness and accuracy' and come with additional computational costs during training.
- The tension can be structural. Tsipras and colleagues proved a trade-off between standard accuracy and adversarial robustness that exists even in a simple, natural setting and persists with infinite data.
- NIST names the concrete pairs: accuracy versus adversarial robustness, explainability versus adversarial robustness, and privacy and fairness.
- NIST's AI RMF holds that trade-offs among the trustworthy characteristics are 'usually involved' and that dealing with them requires taking into account the decision-making context.
- Because NIST finds no mathematically best trade-off, which attribute an organisation prioritises is a governance decision, and for high-risk systems in the EU, Article 15 will require measures against adversarial examples where appropriate, from December 2027 at the earliest.
Hardening an AI system against attack is not free. Every defence trades against something the model would otherwise do — accuracy, explainability, privacy or fairness — and NIST’s position is that a fully secure and fully optimal model cannot simply be assumed to exist. Deciding which trade to accept is a priority call, not an engineering-only one.
Why hardening is not free
Adversarial training, the defence with the longest record, optimises a model for worst-case inputs rather than ordinary ones. NIST states the trade directly: “AI systems that are optimized for accuracy alone tend to underperform in terms of adversarial robustness and fairness. Conversely, an AI system that is optimized for adversarial robustness may exhibit lower accuracy and deteriorated fairness outcomes.”
The tension is structural rather than a tuning problem. Tsipras and colleagues proved that a trade-off between standard accuracy and adversarial robustness “provably exists even in a fairly simple and natural setting”, because robust classifiers learn fundamentally different features, and the phenomenon persists even with infinite data. Zhang and colleagues formalised the same idea by decomposing robust error into natural error plus boundary error, and built a training method, TRADES, explicitly “to trade adversarial robustness off against accuracy”.
Why it matters
NIST draws the governance conclusion: “in most cases, there is no mathematically best trade-off”, so “organizations may need to accept trade-offs between these properties and decide which of them to prioritize depending on the AI system, the use case, and other relevant implications of the AI technology”. Its AI RMF frames the decision as contextual: trade-offs are “usually involved”, and “dealing with tradeoffs requires taking into account the decision-making context”.
For high-risk systems in the EU, EU AI Act Article 15 requires measures against adversarial examples and model evasion “where appropriate”, from 2 December 2027 at the earliest. For those systems, the trade-off becomes partly a compliance question.
Trade-offs
NIST and the research literature document the specific pairs:
- Robustness against accuracy. NIST says all proposed evasion mitigations “exhibit inherent trade-offs between robustness and accuracy” and “come with additional computational costs during training”. The benefit of adversarial training “usually comes at the cost of decreased model accuracy on clean data”; the theory from Tsipras et al. and Zhang et al. says why the cost can be intrinsic.
- Explainability against robustness. NIST states that “there are trade-offs between explainability and adversarial robustness”.
- Privacy against utility. NIST’s AI RMF notes that under data sparsity privacy-enhancing techniques “can result in a loss in accuracy”, and its AML taxonomy records that “differentially private ML models may also have lower accuracy than standard models”.
- Defence against poisoning. Sanitizing training data works in inverse proportion to the attacker’s care: NIST records “a trade-off between attack success and the detectability of malicious samples”, and the measured figures on sanitize training data show filtering costs accuracy even when it works.
- The cost is not only accuracy. NIST lists computational cost during training for evasion mitigations, and scalability and computational cost for formal verification.
Where definitions disagree
Whether the trade-off is unavoidable is contested. The theory says the tension can be inherent — Tsipras et al. prove it in a simple setting and Zhang et al. formalise it. But neither paper says every defence must cost accuracy everywhere: Tsipras et al. note adversarial robustness “can be beneficial in the regime of limited training data”, and NIST calls the full characterisation of these trade-offs “an open research problem”. Whether a particular deployment pays is empirical.
“Robustness” also names two quantities. In NIST’s AI RMF, robustness is generalisability — “the ability of a system to maintain its level of performance under a variety of circumstances”. In adversarial machine learning, robustness is resistance to a worst-case adversary. Both are traded against accuracy, but they are not the same number.
Questions and answers
What is the AI security trade-off?
The AI security trade-off is the cost an organisation accepts when it hardens a model against attack. The clearest documented example is robustness against accuracy: NIST says all proposed evasion mitigations "exhibit inherent trade-offs between robustness and accuracy" and add computational cost, and the research literature shows the tension can be structural rather than a tuning problem. NIST's wider statement is that it may not be possible to simultaneously maximise performance across the trustworthy-AI attributes, so an organisation must decide which to prioritise.
Do defences against adversarial attacks always reduce accuracy?
No, not in every setting. Tsipras and colleagues proved the trade-off "provably exists even in a fairly simple and natural setting", but they also note adversarial robustness can be beneficial in the regime of limited training data. NIST calls the full characterisation of these trade-offs "an open research problem" and says the answer depends on the AI system, the use case and the context; on evasion specifically it states that "designing ML models that resist evasion while maintaining accuracy remains an open problem".
Who decides which AI security trade-offs to accept?
The organisation deploying the system. NIST says organisations "may need to accept trade-offs between these properties and decide which of them to prioritize depending on the AI system, the use case, and other relevant implications of the AI technology". Its AI RMF adds that measurement can show the existence and extent of trade-offs but does not resolve how to navigate them; that depends on the values at play in the relevant context. For high-risk systems in the EU the choice will be constrained by law: from 2 December 2027 at the earliest, Article 15(5) of the EU AI Act requires measures against adversarial examples and model evasion where appropriate.
Sources
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- Artificial Intelligence Risk Management Framework (NIST AI 100-1)NIST, 26 Jan 2023
- Robustness May Be at Odds with AccuracyarXiv, 30 May 2018
- Theoretically Principled Trade-off between Robustness and AccuracyarXiv, 24 Jan 2019
- Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, Robustness and CybersecurityEuropean Union, 12 Jul 2024
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), recital 40 and Article 1 point 40 amending Article 113European Union, 24 Jul 2026