What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Open-weight model

An open-weight model is an AI model whose trained weights, the numbers it learned in training, are published for anyone to download and run. It differs from an API-only model, which is used remotely and never handed over, and from open-source AI, which also requires the training code and enough detail about the training data to build an equivalent model.

Last reviewed

Key points

  • An open-weight model's trained weights can be downloaded and run on your own hardware. An API-only model can be used, but never held.
  • Open weights is not open source. The Open Source Initiative's definition also asks for the training code and detailed information about the training data.
  • A release cannot be taken back. Once weights are downloaded, the developer cannot rescind access or monitor how the model is used.
  • Whoever holds the weights can change them, including the behaviour that makes the model refuse. A 2024 paper estimated it could strip refusals from a 70B model for under $5 of compute.
  • A trained-in refusal is a weaker barrier on an open-weight model. It cannot keep the model's capabilities from anyone who has the file.

What is released and what is not

A model’s weights are the numbers it learned in training. Together with its architecture, they help determine what it outputs for a given input. The US National Telecommunications and Information Administration (NTIA) calls a large foundation model open when its weights are released to the public.

An API-only model keeps its weights private: users send inputs and get outputs, but never hold the model. NTIA lists giving weights only to vetted researchers as a narrower option.

Open weights is also short of open source. The Open Source Initiative’s definition adds the training code and enough information about the training data for a skilled person to build an equivalent system.

Why it matters for security

Holding the weights means holding the controls. NTIA says adversarial actors “can remove safeguards from open models via fine-tuning, then freely distribute the model”. Arditi and colleagues found refusal in 13 open chat models, up to 72B parameters, sat in a single direction inside the model. Editing that direction out of the weights made the models far less likely to refuse harmful requests, a jailbreak the authors estimate costs under $5 of compute for a 70B model. They note it may not generalise to models they did not test. See abliterated vs uncensored models.

The release is permanent. Per NTIA, developers “cannot rescind access to the weights or perform moderation on model usage”, and a file pulled from one site can still spread by other means.

So a refusal trained into an open-weight model is a weaker barrier than on an API-only model. Users of a hosted copy can try to prompt around it, as with closed models. Anyone who holds the file can edit it out.

The same property is a benefit. Running a model yourself means your data never reaches the developer, which NTIA says can matter for confidentiality.

Where definitions disagree

  • NTIA (2024) asks whether a large foundation model’s weights were released openly to the public. The licence does not enter the definition.
  • The Open Source Initiative defines open weights as the final weights and biases of a trained network, and says they “differ significantly from Open Source AI” because the training code and data are missing.
  • The EU AI Act exempts a model from two documentation duties when it is under a free and open-source licence and its weights, architecture and usage information are public (Article 53(2)). Monetised releases do not get it (Recital 103), and neither does any general-purpose AI model with systemic risk.

So a large model whose weights anyone can download under a licence that restricts use counts as open to NTIA, but may not meet the Act’s exemption, which asks for a licence allowing access, use, modification and distribution.

Questions and answers

Is an open-weight model the same as an open-source model?

No. An open-weight model publishes its trained weights. The Open Source Initiative's Open Source AI Definition also requires the complete code used to train and run the model and enough information about the training data for a skilled person to build a substantially equivalent system. A model can publish its weights without either.

Can a developer take back an open-weight model after release?

Not in practice. The US NTIA's 2024 report says developers who release weights publicly cannot rescind access to them. The files can be removed from a platform such as Hugging Face, but anyone who already downloaded them can share them by other means.

Do safety refusals protect an open-weight model from misuse?

Not from whoever holds the weights, because they can change them. Arditi et al. (2024) made 13 open chat models far less likely to refuse harmful requests by editing their weights, and estimate the edit costs under $5 of compute for a 70B-parameter model. The US NTIA notes that safeguards can be removed by fine-tuning and the result freely redistributed.

Sources

  1. Dual-Use Foundation Models with Widely Available Model WeightsNational Telecommunications and Information Administration (US Department of Commerce), Jul 2024
  2. The Open Source AI Definition 1.0Open Source Initiative, 28 Oct 2024
  3. Open Weights: not quite what you've been toldOpen Source Initiative, 29 Jan 2025
  4. Regulation (EU) 2024/1689 (EU AI Act), Article 53 and Recital 103European Union, 12 Jul 2024
  5. Refusal in Language Models Is Mediated by a Single DirectionArditi et al., NeurIPS 2024, Jun 2024

Guides that use this term

In the news