What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Model alignment

Model alignment is the training or fine-tuning of a generative AI model so that its output follows a chosen set of values, goals, or policies, for example refusing harmful requests. MITRE ATLAS lists generative AI model alignment as a security mitigation for nine attack techniques, including prompt injection, jailbreaks, system prompt extraction, and LLM data leakage.

Last reviewed

Key points

  • Alignment changes the model itself, through training methods such as reinforcement learning from human feedback and supervised fine-tuning, rather than filtering what goes in and out.
  • MITRE ATLAS treats generative AI model alignment as a security mitigation, mapping it to nine attack techniques, among them prompt injection, jailbreaks, system prompt extraction, and LLM data leakage.
  • ATLAS claims only that alignment can improve a model's safety, and says to use it together with guardrails rather than instead of them.
  • Alignment is fragile. Research shows jailbreaks still get through aligned models, and fine-tuning on 10 harmful examples undid GPT-3.5 Turbo's safety alignment.
  • The wider term, AI alignment, reaches further. IBM's explainer counts AI governance and ethics boards among alignment methods, and covers superalignment, a branch concerned with future superintelligent systems.

How model alignment works

Model alignment happens during training or fine-tuning. Developers train the model on examples, feedback, or principles that describe the behaviour they want. MITRE ATLAS lists seven common methods, including supervised fine-tuning and reinforcement learning from human feedback (RLHF). In RLHF, as IBM describes it, a reward model trained on human feedback is used to optimize the model.

The result lives in the model’s weights. An aligned model refuses a harmful request because training made refusal its likely response, not because a filter blocked the reply.

Why model alignment matters for security

MITRE ATLAS maps generative AI model alignment to nine attack techniques, among them prompt injection, jailbreaking, system prompt extraction, and LLM data leakage. For each of those four it gives the same reason: alignment “can improve the parametric safety of a model by guiding it away from unsafe prompts and responses.” Parametric safety is safety held in the model’s weights, its parameters.

ATLAS claims no more than “can improve”, and the research shows why. Wei, Haghtalab and Steinhardt tried 28 jailbreaks against GPT-4 and Claude v1.3; for every one of 32 curated harmful prompts from the models’ red teaming, at least one worked. Qi et al. undid GPT-3.5 Turbo’s safety alignment by fine-tuning it on 10 harmful examples for under $0.20. ATLAS itself warns that fine-tuning can remove learned alignment, and says to pair alignment with AI guardrails rather than rely on it alone.

Where definitions disagree

“Alignment” is used at two widths. MITRE ATLAS’s mitigation is Generative AI Model Alignment: “the process of training or fine-tuning a generative model”, extended for AI agents to designing agents that “operate within organizational boundaries”. IBM’s explainer on the wider term, AI alignment, agrees that alignment “often occurs as a phase of model fine-tuning”. It also counts AI governance and corporate AI ethics boards among the ways to achieve alignment, and it includes superalignment, a branch of AI alignment prompted by concern that a future superintelligent system could surpass human control. Neither is wrong; they cover different amounts of ground, and the ATLAS mitigation is the narrower of the two.

Questions and answers

Is model alignment the same as AI guardrails?

No. MITRE ATLAS lists them as separate mitigations. Model alignment changes the model itself through training or fine-tuning. Guardrails sit outside the model and check inputs, outputs, and actions at runtime. ATLAS says alignment should be used in combination with guardrails, not as a substitute.

Can model alignment be removed?

Yes. MITRE ATLAS warns that fine-tuning a model can remove previously learned alignments. Qi et al. showed this on GPT-3.5 Turbo in 2023: fine-tuning on 10 harmful examples, costing under $0.20, made the model respond to nearly any harmful instruction.

Does model alignment stop jailbreaks?

Not reliably. Wei, Haghtalab and Steinhardt tried 28 jailbreaks against GPT-4 and Claude v1.3, both models with extensive safety training. For every one of 32 curated harmful prompts from the models' red teaming, at least one jailbreak worked. MITRE ATLAS says alignment can improve a model's safety, which is a weaker claim than preventing attacks.

Sources

  1. MITRE ATLAS, AML.M0022 Generative AI Model Alignment (collection 2026.09)MITRE
  2. Jailbroken: How Does LLM Safety Training Fail?arXiv (Wei, Haghtalab, Steinhardt), 5 Jul 2023
  3. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!arXiv (Qi, Zeng, Xie, Chen, Jia, Mittal, Henderson), 5 Oct 2023
  4. What is AI alignment?IBM, 18 Oct 2024

Guides that use this term