Definition · AI basics
Open-source AI
Open-source AI is an AI system released on terms that let anyone study it and use, modify and share it for any purpose. Under the Open Source Initiative's definition, that means publishing its weights, the complete code used to train and run it, and enough information about its training data for a skilled person to build a substantially equivalent system.
Last reviewed
Key points
- Under the Open Source Initiative's 2024 definition, anyone may use, study, modify and share open-source AI.
- Open weights are not enough. Open-source AI also needs the complete training and running code, and detailed information about the training data.
- Training data that can legally be shared must be shared. Data that cannot, such as personal data, must be described in detail.
- The EU AI Act's test asks for a free and open-source licence and public weights, architecture and usage information, not training code or data details. Passing it removes some duties, never for general-purpose models with systemic risk.
- Meta called Llama 3.1 open source despite licence limits on who may use it and how. The Open Source Initiative said such limits made the Llama 2 licence not open source.
What has to be released
The Open Source Initiative (OSI) published the Open Source AI Definition 1.0 in October 2024. It starts from four freedoms: anyone may use the system for any purpose without asking permission, study how it works, modify it for any purpose, and share it for any purpose.
Those freedoms need the form a developer would actually change. For a machine learning system, the definition names three parts, each under a licence or terms the OSI has approved:
- Parameters: the model’s weights, the numbers it learned in training, and other settings.
- Code: the complete source code used to train and run the system, including how the data was processed and filtered.
- Data information: enough detail about the training data for a skilled person to build a substantially equivalent system. That means where the data came from, how it was selected and labelled, and where to get the public or purchasable parts.
Training data that can legally be shared must be shared, the OSI’s FAQ says. Data that cannot, such as personal data, is described in detail instead, including its provenance. The OSI rejects releasing all training data as a rule, saying it would limit open-source AI to models trained only on open data.
Why it is more than open weights
An open weight model publishes only the parameters. The OSI says open weights “differ significantly from Open Source AI” because the training code and data details are missing.
The difference matters to anyone checking a model rather than just running it. Without the training code or intermediate checkpoints, the OSI says, “researchers and auditors cannot replicate the model’s development process”. It says that hinders efforts to find when and where biases entered, and makes errors or vulnerabilities nearly impossible to fix.
Where definitions disagree
- The OSI asks for weights, complete training and running code, and detailed data information, all under terms that allow any use.
- The EU AI Act asks for a licence that allows access, use, modification and distribution, plus public weights, architecture and usage information (Article 53(2)). It does not ask for training code or data information, and says such a release “does not necessarily reveal substantial information on the data set” (Recital 104). A licence may still require credit, and require changed versions to be shared on the same or similar terms (Recital 102). Releases sold or otherwise monetised, for example through paid support, do not count, except in deals between microenterprises (Recital 103).
- Meta has used the term more loosely. It called Llama 3.1 405B “the first frontier-level open source AI model” in July 2024. Its licence binds every user to an acceptable use policy. A company whose products, with its affiliates, had over 700 million monthly active users on the release date must ask Meta for a licence, which Meta may refuse. The OSI said in 2023 that the Llama 2 licence was “very plainly not” open source, because it restricted commercial use for some users and restricted uses through its acceptable use policy.
The EU label has consequences. The Act does not apply to AI systems under free and open-source licences, unless they are high-risk, prohibited, or covered by its transparency rules (Article 2(12)). A model provider meeting Article 53(2) skips technical documentation for regulators and information for downstream providers. It must still keep a copyright policy and publish a summary of its training content. A general-purpose AI model with systemic risk gets no exemption.
Questions and answers
What is the difference between open-source AI and an open-weight model?
An open-weight model publishes its trained weights. Open-source AI, as the Open Source Initiative defines it, also publishes the complete code used to train and run the model and enough information about the training data for a skilled person to build a substantially equivalent system, all on terms that allow any use.
Does open-source AI have to publish its training data?
Partly. Under the Open Source Initiative's definition, training data that can legally be shared must be shared. Public data and data that can be bought must be listed with where to get them. Data that cannot be shared, such as personal data, must be described in enough detail for someone to build a dataset with the same structure.
Is Llama open source?
Meta called Llama 3.1 405B an open source AI model in 2024. Its licence binds every user to an acceptable use policy. A company whose products, with its affiliates, had more than 700 million monthly active users on the release date must ask Meta for a separate licence. The Open Source Initiative said the same kinds of restrictions made the Llama 2 licence not open source.
Sources
- The Open Source AI Definition 1.0Open Source Initiative, 28 Oct 2024
- The Open Source AI Definition 1.0 FAQOpen Source Initiative
- Open Weights: not quite what you've been toldOpen Source Initiative, 29 Jan 2025
- Regulation (EU) 2024/1689 (EU AI Act), Articles 2(12) and 53European Union, 12 Jul 2024
- Open Source AI is the Path ForwardMeta, 23 Jul 2024
- Llama 3.1 Community License AgreementMeta, 23 Jul 2024
- Meta's LLaMa license is not Open SourceOpen Source Initiative, 20 Jul 2023