What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Fine-tuning

Fine-tuning is further training of a model that has already been pre-trained, using a smaller dataset chosen to adapt it to a task, a domain or a style of behaviour. It updates the model's weights, or a small add-on to them, and can also weaken the safety training the model already had.

Last reviewed

Key points

  • Fine-tuning takes a model that is already trained and trains it further on a smaller dataset, so it does one kind of task or behaves one particular way.
  • It is cheap next to pre-training, and methods such as LoRA make it cheaper by training a small add-on while the original weights stay frozen.
  • Fine-tuning can remove safety training. In 2023 researchers jailbroke GPT-3.5 Turbo through OpenAI's own fine-tuning API with 10 harmful examples, for under $0.20.
  • Even benign fine-tuning weakened safety. Training on ordinary instruction datasets for a single pass raised harmful answers in every case tested, though less than an attack did.
  • Whoever can fine-tune a model can change its behaviour, whether they hold the weights or only have access to a fine-tuning API.

How it works

A language model is first pre-trained on a very large general corpus. Fine-tuning is the next pass: the trained model is fed a smaller set of examples, input paired with the output wanted, and its weights are nudged toward producing those outputs. The GPT paper names the two stages directly: learn a language model “on a large corpus of text”, then “a fine-tuning stage” that adapts it to a task with labelled data. Instruction tuning is one kind of fine-tuning.

Full fine-tuning updates every weight. Cheaper methods train a small add-on instead. LoRA freezes the original weights and trains small extra matrices; on GPT-3 175B it cut trainable parameters by 10,000 times and GPU memory by three times, with quality on par or better in the authors’ tests.

Why it matters

Fine-tuning can undo model alignment. In 2023 Qi and colleagues fine-tuned GPT-3.5 Turbo through OpenAI’s own API on 10 harmful examples, for under $0.20. After five passes over the data, the share of test answers judged most harmful rose from 1.8% to 88.8%. Ten examples that contained nothing toxic, only training the model to obey first, reached 87.3% after ten passes. The authors told OpenAI before publishing and note that the API may since have added mitigations. Lermen and colleagues used LoRA, one GPU and under $200 to cut Llama 2-Chat 70B’s refusals to about 1%.

No attacker is needed. A single pass over ordinary instruction datasets degraded safety in every case Qi and colleagues tested, though less than an attack. On the Alpaca dataset, GPT-3.5 Turbo’s share of most-harmful answers rose from 5.5% to 31.8%.

Limits of the defences

Mixing safety examples into the fine-tuning data helped in every case Qi and colleagues tried, but the result stayed less safe than the original model. Testing the model after fine-tuning is not a full check either: a model trained with a hidden backdoor trigger gave most-harmful answers to 4.2% of plain harmful prompts and 63.3% with the trigger added. Lermen and colleagues conclude that once weights are released, safety training “does not effectively prevent model misuse”, and a release cannot be recalled.

Questions and answers

What is the difference between fine-tuning and pre-training?

Pre-training teaches a model from scratch on a very large, general corpus. Fine-tuning starts from that trained model and trains it further on a smaller dataset for a task or behaviour. The GPT paper describes exactly these two stages: a language model learned "on a large corpus of text", then "a fine-tuning stage" that adapts it with labelled data.

Can fine-tuning remove a model's safety training?

Yes. In 2023 Qi and colleagues made GPT-3.5 Turbo answer nearly any harmful instruction by fine-tuning it on 10 harmful examples through OpenAI's API, for under $0.20. They told OpenAI before publishing, and note that the API may since have added mitigations. Lermen and colleagues cut Llama 2-Chat 70B's refusal rate to about 1% with LoRA fine-tuning, one GPU and under $200.

Does fine-tuning on harmless data affect safety?

It can. Qi and colleagues fine-tuned GPT-3.5 Turbo and Llama-2-7b-Chat for a single pass over common benign datasets such as Alpaca and Dolly, and safety degraded in every case they tested, though less than under a deliberate attack. On Alpaca, GPT-3.5 Turbo's rate of the most harmful answers rose from 5.5% to 31.8%.

Sources

  1. Improving Language Understanding by Generative Pre-TrainingOpenAI
  2. LoRA: Low-Rank Adaptation of Large Language ModelsarXiv (Microsoft), 16 Oct 2021
  3. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!arXiv (Princeton, Virginia Tech, IBM Research, Stanford), 5 Oct 2023
  4. LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70BarXiv (ICLR 2024 workshop), 22 May 2024

Guides that use this term

In the news