What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Logit bias

Logit bias is a language model API parameter that adds a fixed number to the model's raw score, or logit, for chosen tokens before the next token is picked. Callers use it to make tokens more or less likely or to block them. Researchers have also used logit bias to read out a production model's hidden scores.

Last reviewed

Key points

  • A logit is the raw score a language model gives each possible next token. Logit bias lets the caller add to or subtract from chosen scores.
  • In OpenAI's API the bias runs from -100 to 100 per token. OpenAI says small values nudge a token and -100 or 100 should ban or force it, with the exact effect varying by model.
  • In 2024 Carlini and colleagues used logit bias to recover the hidden scores of OpenAI models, and from them the models' final layer. Google's PaLM-2 API was open to the same attack, and both providers changed their APIs after disclosure.
  • Returning log probabilities alongside logit bias makes the attack cheaper, but hiding them is not enough. Researchers judge removing logit bias effective, at the cost of legitimate uses such as blocking unwanted tokens.

How it works

A language model picks each next token by giving every token in its vocabulary a raw score, called a logit. A softmax turns those scores into probabilities, and the model samples from them.

Logit bias lets the caller change the scores first. OpenAI’s API takes a map from token IDs to a number between -100 and 100, which is added to each chosen token’s logit before sampling. OpenAI says the exact effect varies by model: values between -1 and 1 should nudge a token, while -100 or 100 should ban it or force it.

Some APIs, including OpenAI’s, can also return log probabilities: the logged probabilities of the few most likely tokens. OpenAI’s legacy completions endpoint returns up to five.

Why it matters

Logit bias gives a caller a way to probe a model’s hidden scores. In 2024 Carlini and colleagues attacked APIs that returned the top five log probabilities and accepted a logit bias. A large bias pushed chosen tokens into the top five so their values could be read. Repeated across the vocabulary, with some algebra to undo the bias, that recovered the model’s full score vector.

With enough of those vectors they recovered the final layer, the embedding projection, of OpenAI’s ada and babbage models. This is a form of model extraction. The APIs of Google’s PaLM-2 and OpenAI’s GPT-4 were open to the same attack, and both providers changed them after disclosure. OpenAI stopped logit bias from affecting the returned log probabilities.

Log probabilities are not required. With logit bias alone, an attacker can search for the bias that just changes the model’s top choice, which reveals the same scores at a higher price. Carlini and colleagues found the attack 10 times cheaper when a caller could both supply a logit bias and see log probabilities.

Defences and their cost

For predictive AI models, MITRE ATLAS’s general defence is to return only what the application needs. Its advice includes withholding or reducing the precision of confidence scores and logits, to make extraction and black-box adversarial example searches harder.

For logit bias, removing the parameter is the simplest fix. Finlayson and colleagues, in a separate 2024 study, judge removal effective: no known method recovers full outputs from an API without logit bias. Both papers note the cost: logit bias has legitimate uses, such as blocking unwanted tokens and constraining output.

Carlini and colleagues suggest middle paths. An API could refuse logit bias and log probabilities in the same request, or offer only a block-list of banned tokens. They also consider rate-limiting logit bias, but find that it has significant drawbacks, such as having to track every user’s queries.

Hiding logit bias may not close every route. Carlini and colleagues warn that other settings, such as an unconstrained temperature, the setting that controls how random the output is, could also leak logit values.

Questions and answers

What is logit bias used for?

Logit bias is used to steer which tokens a language model can produce. OpenAI's API documentation gives the example of passing -100 for the end-of-text token so the model never emits it. Carlini and colleagues note that researchers also use logit bias for controlled or constrained generation.

Is logit bias a security risk?

Logit bias is a security risk because it lets a caller probe a model's hidden scores. In 2024 Carlini and colleagues used it, with the log probabilities the API returned, to recover the final layer of OpenAI's ada and babbage models. OpenAI and Google both changed their APIs after disclosure. Without log probabilities the attack still works, at a higher cost.

What is the difference between logits and log probabilities?

Logits are a language model's raw scores for each possible next token. Log probabilities are those scores after the softmax turns them into probabilities, then logged. The production APIs Carlini and colleagues studied in 2024 returned log probabilities for a few top tokens rather than raw logits, but logit bias let them work back to the logits.

Sources

  1. OpenAI API specification (openapi.yaml, version 2.3.0)OpenAI
  2. Stealing Part of a Production Language Model (later published in Proceedings of the 41st International Conference on Machine Learning, PMLR 235)arXiv, 11 Mar 2024
  3. Logits of API-Protected LLMs Leak Proprietary Information (COLM 2024)arXiv, 14 Mar 2024
  4. MITRE ATLAS, mitigation AML.M0002 (collection 2026.09)MITRE