Definition · AI basics
Model checkpoint
A model checkpoint is a saved snapshot of a model's parameters — and often its optimiser state — at a specific point during training, identified by a training-step number or internal tag and serialised to a file format such as safetensors or PyTorch's .pt.
Last reviewed
Key points
- Checkpoints exist primarily as internal training infrastructure: resume after interruption, roll back if training diverges, or study capability emergence across steps.
- A weights-only checkpoint stores parameters only, used for inference and deployment. A full training checkpoint adds optimiser state, scheduler state, step counters, and random seeds for resuming training cleanly.
- Not all checkpoints become public releases. A training run produces thousands; only a handful are externally versioned. A checkpoint is the artefact; a version is the label some checkpoints receive.
- The AI supply chain trades in checkpoints as distinct artefacts: a base model checkpoint, a fine-tuned checkpoint, and a published checkpoint are three different things.
- Provenance tools trace weight derivation at the checkpoint level. Two models are provenance-linked if a causal chain of checkpoint-to-checkpoint weight derivation connects them.
A model checkpoint is the artefact the AI supply chain actually trades in.
When ai-bill-of-materials says “the inventory answers which deployed
systems inherited it” and model-poisoning says “treat a checkpoint of
unknown provenance as unfixable”, the unit they mean is the checkpoint.
How it works
A checkpoint serialises model parameters to disk. A weights-only checkpoint contains just the weights — everything to run the model but nothing to resume training. A full training checkpoint adds optimiser state (momentum, adaptive statistics), learning-rate scheduler state, step counters, and random seeds. Without these, the optimiser restarts from scratch and early steps often spike the loss.
Framework implementations differ. TensorFlow’s checkpoint captures tf.Variable
objects but omits the computation graph. PyTorch’s state_dict maps layers to
tensors; the recommended pattern for resumption packs model state, optimiser
state, epoch, and loss into one dictionary.
Why it matters
The AI supply chain trades in checkpoints, not architectures. A base model checkpoint, a fine-tuned checkpoint, and a published checkpoint are distinct artefacts with different supply-chain risks. Cisco’s Model Provenance Constitution grounds provenance in weight-derivation history: two models are provenance-linked if a causal chain of weight derivation connects them.
Fine-tuning generates 79% of derived models on Hugging Face. High-degree hub models serve as critical bases. When security teams ask “does this deployed model inherit a known vulnerability?” or compliance asks “does this checkpoint trigger a licensing obligation?”, the answer hinges on checkpoint provenance.
Where definitions disagree
Whether a checkpoint that has been heavily fine-tuned on unrelated data still counts as the “same” model. The Model Provenance Constitution treats provenance as a factual property of the derivation chain — a model fine-tuned for 100 steps is clearly derived, and a model fine-tuned over a very large token budget on entirely unrelated data is also derived, even if the detectable signal is near zero. Practical governance may need to distinguish these cases explicitly, but the constitution does not.
Questions and answers
What is the difference between a checkpoint and a model?
A model is three things: the architecture (the fixed structure), the weights (the learned numbers that fill that structure), and the checkpoint (the saved file that stores them). The architecture is code and a small config; the weights are tensors in memory; the checkpoint is the serialised file on disk that lets you reload the weights into the architecture later.
What is the difference between a weights-only checkpoint and a full training checkpoint?
A weights-only checkpoint contains only the model parameters and is used for inference, deployment, and sharing. A full training checkpoint also stores the optimiser state (momentum, adaptive statistics), the learning rate scheduler state, the current epoch and step, and random seeds — everything needed to resume training exactly where it stopped. A full checkpoint is roughly 3–4× the size of a weights-only one.
Can I resume training from a weights-only checkpoint?
You can run the model, but resuming training cleanly requires the optimiser state. Without it the optimiser restarts from scratch, the first few hundred steps often spike the loss, and recent progress is undone. If your goal is to resume, you must save the optimiser and scheduler state too.
Is every checkpoint a released model version?
No. A training run may produce thousands of checkpoints; only a handful are externally versioned and released. A checkpoint is the artefact; a version is the label some checkpoints receive; lineage is the graph describing how one checkpoint relates to another.
Sources
- CASRAI Dictionary, Model checkpointCASRAI, 21 May 2026
- What Is Really a Machine Learning Model? Architecture vs Weights vs CheckpointsTeacherAndTask, 22 Jun 2026
- Training checkpoints | TensorFlow CoreTensorFlow
- Saving and Loading Models | PyTorch TutorialsPyTorch
- Defining Model Provenance: A Constitution for AI Supply Chain SecurityCisco AI Defense, 30 Apr 2026
- Model Provenance Constitution, Section 3.2–3.4Cisco AI Defense
- A First Look at Model Supply Chain: From the Risk PerspectiveIEEE/ACM ICSE 2026, 1 Jan 2026