What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Instruction tuning

Instruction tuning is fine-tuning a pre-trained language model on a collection of tasks written as natural-language instructions, each paired with the wanted response, to make the model better at following instructions. Instruction sets are far smaller than pre-training corpora, and researchers have poisoned one with a hundred examples.

Last reviewed

Key points

  • Instruction tuning is a second training pass. A model already pre-trained on raw text is fine-tuned on examples written as instructions with the answers wanted.
  • Instruction sets are small next to pre-training corpora. Research sets have run to millions of examples, but LIMA used 1,000 examples, about 750,000 tokens, on a model pre-trained on about 1.4 trillion tokens.
  • Poison counts here are small. In 2023 Wan and colleagues planted 100 poisoned examples in a training set of about 5,000, roughly one in fifty, and a chosen trigger phrase then steered predictions on tasks that were never poisoned.
  • Outsiders can contribute to these sets. FLAN pools open-source datasets, OpenAI has drawn on prompts its users submitted, and builders try to make the sets bigger.
  • NIST names instruction tuning as a poisoning target separate from pre-training.

Two stages, very different sizes

A language model is first pre-trained to predict text across a very large corpus. Instruction tuning is a second pass. Each training example states a task as an instruction and gives the answer wanted. In the FLAN paper, Wei and colleagues tuned a 137-billion-parameter model on “over 60 NLP datasets” phrased this way. Given no worked examples, it beat GPT-3 on 20 of 25 datasets from task types held out of tuning, though the gains appeared only at large model sizes.

The second stage is much smaller. The LIMA authors tuned a 65-billion-parameter LLaMA model on 1,000 examples, “roughly 750,000 tokens”. That model was pre-trained on about 1.4 trillion tokens. LIMA is deliberately small: earlier work used “large multi-million-example datasets”.

Why it matters for poisoning

NIST names instruction tuning as a stage data poisoning can target, “beyond the vast quantities of pre-training data”.

In 2023 Wan and colleagues trained on ten tasks of about 500 samples each and planted 100 poison examples across five of them, roughly one in fifty. A chosen trigger phrase then worked as a backdoor on held-out tasks that were never poisoned, while “poisoning does not affect accuracy on regular inputs”. In that set, a hundred was a small number but not a small share. Web-scale data poisoning is measured the other way, as a percentage of a scraped corpus.

Outsiders can contribute to these sets. Wan and colleagues note that FLAN aggregates open-source datasets, OpenAI drew on prompts users submitted, and “organizations seek to maximize fine-tuning data quantity”. They also note a limit: InstructGPT trains on “a small fraction” of user queries, so an attacker’s submission is unlikely to be picked. Filtering defences measured on such a set are in sanitize training data.

Where definitions disagree

The term overlaps with supervised fine-tuning. MITRE ATLAS lists “Supervised Fine-Tuning (SFT)” and “Instruction Tuning” as separate alignment methods. InstructGPT calls its step of training on labeller-written answers supervised fine-tuning, while LIMA’s supervised fine-tuning on 1,000 prompts is what it calls instruction tuning data. ATLAS also lists instruction tuning under model alignment, a security mitigation.

Questions and answers

What is the difference between instruction tuning and pre-training?

Pre-training teaches a language model by having it predict text across a very large corpus. Instruction tuning comes after it and teaches the model to respond to instructions, using a far smaller set of instruction-and-answer examples. The LIMA authors fine-tuned a 65-billion parameter LLaMA model on 1,000 examples, about 750,000 tokens; that model was pre-trained on about 1.4 trillion tokens. Their results, they write, "strongly suggest" that almost all of a model's knowledge is learned in pre-training.

Why can 100 poisoned examples affect an instruction-tuned model?

Wan and colleagues found that fine-tuned models update quickly: training on 100 examples "can overwrite a model's immense pre-training prior" for a phrase. Their training set was ten tasks of about 500 samples each, so the 100 were roughly one in fifty. With 100 deliberately mislabelled examples, a 3-billion-parameter model labelled 92.8% of negative test inputs containing the trigger phrase "James Bond" as positive, on tasks it was never poisoned on. Larger models were more vulnerable in some settings, not less.

Is instruction tuning the same as RLHF?

No. Reinforcement learning from human feedback is a separate stage that can be run after instruction tuning. Wan and colleagues describe RLHF as conducted "after conducting instruction-tuning", and LIMA was instruction-tuned "without any reinforcement learning".

Sources

  1. Finetuned Language Models Are Zero-Shot LearnersarXiv (ICLR 2022), 3 Sep 2021
  2. LIMA: Less Is More for AlignmentarXiv (Meta AI), 18 May 2023
  3. LLaMA: Open and Efficient Foundation Language ModelsarXiv (Meta AI), 27 Feb 2023
  4. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  5. Poisoning Language Models During Instruction TuningarXiv (ICML 2023), 1 May 2023
  6. MITRE ATLAS, AML.M0022 Generative AI Model Alignment (collection 2026.09)MITRE
  7. Training language models to follow instructions with human feedbackarXiv (OpenAI), 4 Mar 2022

Guides that use this term