Definition · AI basics
Token
A token is the unit of text a language model reads and writes. A tokenizer splits text into tokens drawn from a fixed vocabulary, and a token can be a whole word, part of a word, a single character or a byte. Context windows are sized in tokens, not words, and APIs such as Anthropic's charge per token.
Last reviewed
Key points
- A token is what a language model actually reads and writes. A tokenizer splits text into tokens before the model sees it.
- Tokens are often pieces of words. Common words become one token; rare words are built from several smaller ones, which is how a model handles words it never saw whole.
- Context windows are sized in tokens, and APIs such as Anthropic's bill by them. The same text can need very different counts depending on the tokenizer and the language.
- Tokenization can be attacked. HiddenLayer's 2025 TokenBreak altered words so that prompt-injection classifiers using two common tokenizer types missed attacks the target model still understood.
How it works
Before a language model sees any text, a tokenizer splits it into tokens from a fixed vocabulary. The model works on those tokens, not on the raw text.
Many tokenizers split rare words into pieces. Byte pair encoding (BPE), adapted for subword splitting by Sennrich and colleagues in 2015, starts from single characters and repeatedly merges the most frequent adjacent pair. Common words end up as one token; a rare word becomes several. Anthropic’s glossary states the trade-off: larger tokens are efficient, smaller ones let a model handle words it has never seen.
A context window is sized in tokens. Some APIs also price by them: Anthropic charges per million input tokens and per million output tokens.
Why it matters
Token counts decide what a request costs and whether it fits. MITRE ATLAS advises generative AI services to limit context and output tokens, one of its defences against cost harvesting. Counts also vary more than people expect. A 2023 Oxford study found the tokenizer used by ChatGPT and GPT-4 needed up to 15 times as many tokens for the same text in Shan as in English, and 3 times in Arabic. Anthropic says Claude 4.7 and later models produce about 30% more tokens than earlier ones for the same text, with the exact increase depending on the content.
The tokenizer is also an attack surface. In 2025 HiddenLayer’s TokenBreak changed words slightly, so that AI guardrails classifiers using BPE or WordPiece, a similar method, missed a prompt injection the target model still understood. Classifiers using a different method, Unigram, were not fooled in those tests. Separately, tokens that a model barely saw in training, called glitch tokens, can make it behave unpredictably; a 2024 Cohere study found them in every model it tested.
Questions and answers
How long is a token?
It depends on the tokenizer and the language. Anthropic's glossary puts a Claude token at about 3.5 English characters, varying by language. Other languages can need more tokens for the same text: a 2023 Oxford study found the tokenizer used by ChatGPT and GPT-4 needed about 1.6 times as many for Italian as for English, 3 times for Arabic and up to 15 times for Shan.
How do tokens affect what an AI API costs?
Commercial AI services charge per token or per character, a 2023 Oxford study notes. Anthropic, for example, prices each model per million input tokens and per million output tokens. A tokenizer change alters the count for the same text: Anthropic says Claude 4.7 and later models produce about 30% more tokens than earlier ones, depending on the content.
What is a glitch token?
A glitch token is a token in a tokenizer's vocabulary that the model barely saw in training, such as _SolidGoldMagikarp. Prompts containing one can make the model behave oddly. A 2024 Cohere study found such tokens in every model it tested.
Sources
- Glossary (Claude Platform Docs)Anthropic
- Neural Machine Translation of Rare Words with Subword UnitsRico Sennrich, Barry Haddow and Alexandra Birch, University of Edinburgh, 31 Aug 2015
- Pricing (Claude Platform Docs)Anthropic
- Token counting (Claude Platform Docs)Anthropic
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H.S. Torr and Adel Bibi, University of Oxford (NeurIPS 2023), 17 May 2023
- TokenBreak: Bypassing Text Classification Models Through Token ManipulationKasimir Schulz, Kenneth Yeung and Kieran Evans, HiddenLayer, 9 Jun 2025
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land and Max Bartolo, Cohere, 8 May 2024
- MITRE ATLAS, mitigation AML.M0036 (collection 2026.09)MITRE