Definition · AI security
Unsafe AI artifacts
Unsafe AI artifacts are model files and their bundled components — weights, chat templates, tokenizer metadata, configuration — built or altered so that loading or running them harms the victim, either by executing code on the host or by changing what the model does. MITRE ATLAS files the technique as AML.T0011.000, a sub-technique of User Execution.
Last reviewed
Key points
- An unsafe AI artifact does its harm when the victim loads it, runs inference with it or lets a runtime interpret it, and it may keep working as expected.
- Pickle-based model files are a familiar case. Python's own documentation warns that malicious pickle data can execute arbitrary code during unpickling.
- No software bug is needed. ATLAS says an artifact may instead use functionality an AI runtime supports on purpose, such as a chat template, to alter outputs or agent actions.
- ATLAS's own descriptions connect three techniques. An adversary publishes a poisoned artifact (AML.T0115), a victim acquires it through AI supply chain compromise (AML.T0010), and the victim's loading or running it is AML.T0011.000.
- Safer loaders and scanners narrow the pickle route, and the technique is wider than pickle. Researchers poisoned the chat template inside GGUF model files without touching the weights, and malicious models on Hugging Face got past its pickle scanner.
An unsafe AI artifact is a model file, or a part bundled with one, that does harm when the victim loads it or runs it.
How it works
MITRE ATLAS notes that “AI models may be packaged as a combination of instructions and weights.” Loading such a file can run the instructions. PyTorch’s torch.save and torch.load use pickle by default, and Python’s documentation warns: “It is possible to construct malicious pickle data which will execute arbitrary code during unpickling.” ATLAS calls planting that code Embed Malware (AML.T0018.002) and notes that “Models with embedded malware may still operate as expected.”
Code execution is one route, not the only one. ATLAS says exploiting a software vulnerability “is not required”: an artifact “may instead use functionality intentionally supported by an AI runtime to alter model context, model outputs, safety behavior, or agent actions.” A chat template is one such part: the logic that turns a conversation into the text the model actually reads, and the runtime applies it automatically during inference. Altering it exploits no bug and leaves the weights untouched.
ATLAS’s descriptions link three techniques in order. An adversary publishes a poisoned artifact (AML.T0115), victims may acquire it through AI supply chain compromise (AML.T0010), and loading or running it is this technique, filed under Execution, ATLAS’s tactic for running adversary code.
Why it matters
Loading an unsafe artifact can give the attacker control of the machine, or steer what the model says and does. On the host, embedded code can open a reverse shell, a connection that gives the attacker remote control. Through an altered template, which builds the model’s input, a trigger can change answers or redirect an agent’s tool use. ATLAS says the artifact “may continue to work as expected”, so the victim may not notice.
In practice
ReversingLabs found models on Hugging Face that opened reverse shells when loaded; ATLAS records two instances as case study AML.CS0031, typed an Incident. Hugging Face scanned uploads with Picklescan, which missed them. The files were “seemingly purposefully corrupted in a way that the malicious payload is executed before the model ultimately fails to de-serialize fully”, that is, before loading finishes, and Picklescan “relied on being able to fully de-serialize the model.” ATLAS files that trick as its own technique, Corrupt AI Model (AML.T0076).
The template route has been demonstrated by researchers. In AML.CS0064, typed an Exercise rather than an Incident, Pillar Security and Fujitsu Research of Europe poisoned the chat templates bundled in GGUF, a widely used format that packs weights, configuration and chat template into one file. When a trigger appeared, “the template injects attacker-controlled instructions into the context sent to the model”. ATLAS says the attack was validated across eighteen models from seven families and four inference engines, the programs that run a model. In controlled evaluations, it manipulated responses, redirected agent tool use and exfiltrated sensitive data.
Trade-offs
ATLAS maps six mitigations, among them code signing, vulnerability scanning and an AI bill of materials. Each has limits.
Safer loaders narrow the pickle route. Safetensors stores tensors “safely (as opposed to pickle)”. Since version 2.6, PyTorch’s torch.load uses weights_only=True when no pickle_module argument is passed, which PyTorch says “narrows the surface of remote code execution attacks” but does not guard against denial of service, and memory corruption “might still be possible”. Neither addresses a poisoned chat template, which works through a feature the runtime is meant to run.
Scanning helps but trails the attacker. After the AML.CS0031 incident, Hugging Face changed Picklescan to catch that trick, and ATLAS still warns that “there may be other types of malicious pickles that Picklescan cannot detect.”
Questions and answers
Is an unsafe AI artifact the same as a pickle bomb?
Not quite. A malicious pickle-based model file is one kind of unsafe AI artifact, but the technique ATLAS names, AML.T0011.000, also covers chat templates, tokenizer metadata, configuration and other bundled logic. The name is also questioned. When a scanner's author used "pickle bomb" on Hacker News in December 2025, a commenter objected that it suggests a resource-exhaustion attack like a zip bomb, and the author replied that they had used it colloquially.
Does converting a model to safetensors make it safe to load?
It addresses pickle, not the whole technique. Safetensors is a format for storing tensors "safely (as opposed to pickle)". ATLAS's definition of unsafe AI artifacts also covers chat templates, tokenizer metadata, configuration and pre- and post-processing logic. A safe format for tensors says nothing about those components.
Which MITRE ATLAS technique covers unsafe AI artifacts?
AML.T0011.000, Unsafe AI Artifacts, a sub-technique of User Execution (AML.T0011) under the Execution tactic (AML.TA0005). ATLAS maps six mitigations to it, including code signing, vulnerability scanning and an AI bill of materials.
Sources
- MITRE ATLAS, AML.T0011.000 Unsafe AI Artifacts (collection 2026.09)MITRE
- Python documentation, pickle — Python object serializationPython Software Foundation
- PyTorch documentation, Serialization semantics (2.14)PyTorch Foundation
- huggingface/safetensors READMEHugging Face
- Hacker News, "AIsbom – open-source CLI to detect "Pickle Bombs" in PyTorch models"Hacker News, 16 Dec 2025