What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Language model

A language model is a probability distribution over sequences of words and other symbols, learned from text, that lets software predict the next token in a sequence. A large language model is the same idea scaled up — a deep neural network with a very large number of parameters trained on a very large corpus.

Last reviewed

Key points

  • A language model is a probability distribution over sequences of tokens — words, word parts, punctuation — learned from text, which lets software predict the next token in a sequence.
  • The same underlying idea powers speech recognition, translation, autocomplete and chatbots, because each uses the model's probability of what comes next.
  • A large language model is the same idea scaled up: a deep neural network with billions of parameters trained on huge amounts of text, which is what runs a modern chatbot.
  • The model is a next-token predictor, not an exact record. Its output is generated one token at a time from the probabilities it learned.

A language model predicts what comes next in a sequence of words. Most of modern text technology is built on that single ability.

How it works

A language model learns a probability for every next token — word, word part or punctuation symbol — given the tokens that came before it. Mitchell, writing for the MIT Open Encyclopedia of Cognitive Science, calls it a probability distribution over possible sequences of words, used to predict unseen tokens.

Modern large language models are the same idea scaled up. The Transformer, introduced in the 2017 paper “Attention Is All You Need,” set the architecture now used widely: it replaces recurrence and convolution with attention, and the paper reports better translation quality with less training time.

Why it matters

Language models sit underneath speech recognition, translation, autocomplete and chatbots. For each, the model is doing the same thing — choosing the most likely next token — while the application around it decides what to do with the text.

The same design explains the failures. Because the model only knows how words usually go together, it can state a falsehood confidently, and because it treats every instruction as equal text, it is vulnerable to prompt injection.

Text is not always what the software around the model wants. When a program only needs a small decision, it has to parse that decision back out of the text. TypeSafe AI coined the term System One model in September 2026 for its class of models that skip the text and return a typed answer, such as a choice or a probability, that a program can act on directly.

Questions and answers

Is a language model the same as a chatbot?

No. A chatbot is an application built on a language model. The model generates text; the chatbot adds the interface, the conversation history and the rules for who can say what. The same model can power many different chatbots.

Does a language model know facts?

Only incidentally. A language model has learned how words are usually arranged across the text it was trained on, so it can state things that look true because that is how those sentences usually go. It does not hold a reliable store of facts, which is why language models can confidently state something false.

What is the difference between a language model and a large language model?

Scale, in parameters and training data. A language model is the general idea of a probability distribution over sequences. A large language model is that idea implemented as a deep neural network with a very large number of parameters trained on a very large corpus, and it is what runs a modern chatbot.

Sources

  1. Large Language ModelsMIT Open Encyclopedia of Cognitive Science, 24 Jul 2024
  2. Attention Is All You NeedarXiv (Vaswani et al.), 12 Jun 2017

Guides that use this term

In the news