Definition · AI basics
In-context learning
In-context learning is a language model's ability to perform a task from instructions or worked examples placed in its prompt, without any change to the model's weights. Given a few examples and then a new input, the model continues the pattern. OpenAI researchers named and measured the ability in the 2020 GPT-3 paper.
Last reviewed
Key points
- In-context learning means a model picks up a task from its prompt alone. The model's weights do not change, so nothing from the examples carries over to a later prompt.
- Prompts are called zero-shot, one-shot or few-shot by how many worked examples they hold. In the GPT-3 paper, more examples helped on most tasks, and larger models gained more from them.
- Researchers disagree on what the model is doing. It may learn a new task from the examples, or mostly recognise a task it already knew.
- The same ability works for anyone who writes the prompt. Many-shot jailbreaking uses it to override a model's safety training with hundreds of faked examples.
How it works
In-context learning works by putting the task in the prompt instead of in the model. The prompt holds a few worked examples, such as English sentences with their French translations, and then one new sentence. The language model predicts what comes next, which is the translation, following the pattern.
Nothing about the model changes. The GPT-3 paper that named the ability stressed that its “learning” curves involved no weight updates, only more examples in the prompt.
Prompts are sorted by how many examples they carry. Zero-shot gives only an instruction, one-shot adds one example, and few-shot adds several. In the GPT-3 paper, few-shot meant as many examples as fitted in the model’s context window, typically 10 to 100. On most tasks, performance rose with the number of examples and with model size, and the gap between zero-shot and few-shot tended to grow as models got larger.
Why it matters
In-context learning lets a user adapt a model to a task by writing a prompt with examples, instead of collecting thousands of labelled examples and retraining.
It also works for an attacker. Many-shot jailbreaking fills one long prompt with faked exchanges of an assistant answering harmful questions. Anthropic says the attack can be seen as a special case of in-context learning. As examples are added, benign in-context learning and the attack improve along the same kind of curve. Among Claude 2.0 models, larger ones learned faster in context and tended to need fewer faked exchanges to reach the same attack success rate.
The researchers add a warning. If the same internal machinery drives both, blocking the attack without weakening in-context learning itself may prove hard.
Where definitions disagree
The name says “learning”, but whether the model learns is contested. The GPT-3 authors left it open: the model might learn a new task from scratch or only recognise a task it saw in training, and the answer may differ by task.
A 2022 study found that correct labels matter less than expected. Across 12 models, swapping the correct labels in examples for random ones cost little on classification and multiple-choice tasks. The format and the kinds of inputs and labels mattered more. The authors said the answer turns on how strictly “learning” is defined, and suggested in-context learning may fail on a task the model has not already captured.
A 2023 Google study found scale changes the picture. Small models ignored deliberately flipped labels and kept to what they knew. Large models followed the flipped labels, which means they did pick up a mapping the examples taught them.
Questions and answers
What is the difference between in-context learning and fine-tuning?
Fine-tuning changes a model's weights by training it on task examples, so the change lasts. In-context learning changes nothing in the model: the examples sit in the prompt, shape that one response, and are gone when the prompt is. The GPT-3 paper that named in-context learning contrasted the two directly.
What do zero-shot, one-shot and few-shot mean?
They count the worked examples in a prompt. A zero-shot prompt gives only an instruction, a one-shot prompt adds one example, and a few-shot prompt adds several. In the 2020 GPT-3 paper, few-shot meant as many examples as fitted in the model's context window, typically 10 to 100.
Why does in-context learning matter for AI security?
Whoever writes the examples in a prompt can steer the model, and that includes an attacker. Many-shot jailbreaking fills a prompt with faked examples of an assistant answering harmful questions. Anthropic's researchers found the attack grows with the number of examples along the same kind of curve as benign in-context learning. They warn that if one mechanism drives both, blocking the attack without weakening the ability may be hard.
Sources
- Language Models are Few-Shot LearnersTom B. Brown et al., OpenAI, 28 May 2020
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min et al., University of Washington and Meta AI (EMNLP 2022), 25 Feb 2022
- Larger language models do in-context learning differentlyJerry Wei et al., Google Research, 7 Mar 2023
- Many-shot jailbreakingAnthropic, 2 Apr 2024
- Many-shot Jailbreaking (NeurIPS 2024)Cem Anil et al., Advances in Neural Information Processing Systems 37