Definition · AI basics
Context window
A context window is the limit on how much text a language model can take into account at once, measured in tokens. The window holds everything the model works from for one response: instructions, the conversation so far, any documents or tool results, and the response itself. Text that does not fit in the window does not reach the model.
Last reviewed
Key points
- A context window is a language model's working space for one response. Instructions, conversation history, documents, tool results and the model's own reply all count against it.
- Windows have grown fast. GPT-3 had 2,048 tokens in 2020; Anthropic put typical windows at about 4,000 tokens in early 2023 and 1,000,000 or more for some models by April 2024.
- A bigger window is not automatically better. In a 2023 study, models often used information at the start or end of a long input better than information in the middle.
- Long windows widen the attack surface. Many-shot jailbreaking needs one, and MITRE ATLAS notes that larger windows let planted instructions persist longer in a conversation.
How it works
A language model writes its response one token at a time, and each step can draw only on the text inside its context window. That text is the prompt for the turn, meaning the system prompt, the conversation so far and any documents or tool results, plus the reply generated so far. Anthropic’s documentation calls the window the model’s “working memory”, as distinct from what it learned in training.
The window has a fixed size. GPT-3 had 2,048 tokens in 2020, which limited how many worked examples a prompt could hold for in context learning. The GPT-3 paper says the window typically fitted 10 to 100. Anthropic put typical windows at about 4,000 tokens in early 2023, and some models had 1,000,000 or more by April 2024.
When the text outgrows the window, something gives. Anthropic’s API rejects a request whose input alone is too long, and chat interfaces such as claude.ai can drop the oldest turns first.
Why it matters
A longer context window does not guarantee the model uses all of it. A 2023 Stanford-led study found models often used information at the start or end of a long input better than in the middle. In some settings GPT-3.5-Turbo did worse than it did with no documents at all. Anthropic’s documentation says accuracy and recall degrade as the token count grows, and calls this “context rot”.
A longer window also gives an attacker more room. Many-shot jailbreaking fills the window with many faked exchanges, up to 256 in Anthropic’s tests. Anthropic says capping window length would stop the attack outright, but at the cost of what long inputs are for. MITRE ATLAS calls planting instructions in a conversation thread poisoning, and notes that larger windows let those instructions persist longer. Whatever enters the window, including a retrieved document carrying a prompt injection, sits alongside the instructions the model was given.
Questions and answers
What is a token in a context window?
A token is the smallest unit a language model reads, which can be a word, part of a word, a character or a byte. Context windows are sized in tokens rather than words. Anthropic's glossary puts a Claude token at about 3.5 English characters, varying by language.
What happens when a conversation is longer than the context window?
The text that does not fit cannot reach the model. Anthropic's API rejects a request whose input alone is too long for the window, while chat interfaces such as claude.ai can drop the oldest turns first to make room, and the model no longer sees the turns that were dropped.
Is a bigger context window always better?
No. A 2023 study found models often used information at the start or end of a long input better than information in the middle, and Anthropic's documentation says accuracy and recall degrade as the token count grows. Longer windows also made many-shot jailbreaking practical.
Sources
- Context windows (Claude Platform Docs)Anthropic
- Glossary (Claude Platform Docs)Anthropic
- Language Models are Few-Shot LearnersTom B. Brown et al., OpenAI, 28 May 2020
- Many-shot jailbreakingAnthropic, 2 Apr 2024
- Lost in the Middle: How Language Models Use Long ContextsNelson F. Liu et al., Stanford University, UC Berkeley and Samaya AI (TACL), 6 Jul 2023
- MITRE ATLAS, technique AML.T0080.001 Thread (collection 2026.09)MITRE