Definition · AI basics
Chat template
A chat template is the format that turns a conversation into the single token sequence a chat model reads. The template wraps each message in control tokens marking its role, such as system, user or tool, and where each turn starts and ends. Formats vary between models, and a model expects the format it was trained with.
Last reviewed
Key points
- A chat model never sees a list of messages. A chat template flattens the conversation into one token stream, marking each turn with control tokens such as Qwen's <|im_start|> and <|im_end|>.
- Formats vary between models, even between two fine-tuned from the same base or two versions of one family, and a model given the wrong control tokens can perform far worse.
- Role markers label turns; they do not wall them off. A web page a tool fetched sits in the same token stream as the user's request, and attacks such as ChatInject and Phantom forge the markers inside such content.
- The template is also code that ships with the model, usually in Jinja. Researchers have poisoned templates to plant instructions, and one runtime's unsandboxed rendering let a template run code.
A chat model is a language model that continues a sequence of tokens. Hugging Face’s documentation puts it plainly: “The chat is still just a sequence of tokens”. The chat template is what turns a list of messages into that sequence.
How it works
An application passes the model a list of messages, each with a role and some content. The template wraps each message in control tokens and joins them. Hugging Face’s documentation shows a short chat, ready for the model’s reply, using the <|im_start|> and <|im_end|> markers that Qwen models use:
<|im_start|>user
Hi there!<|im_end|>
<|im_start|>assistant
The last line tells the model an assistant reply comes next.
Markers vary. OpenAI’s gpt-oss writes <|start|>user<|message|>. Gemma 3 wrote <start_of_turn>user, and Gemma 4 switched to <|turn>user. Hugging Face shows two models fine-tuned from the same base with different formats, and says that “with the wrong control tokens, these models would have drastically worse performance”. So the template ships with the model, usually as a file in Jinja, a templating language.
Why it matters
Role markers label turns; they do not wall them off. A web page a tool fetched lands in the same token stream as the user’s request. The paper behind Phantom, a 2026 agent-hijacking attack, says templates “do not enforce strict architectural isolation between control instructions and untrusted data”. If retrieved content contains the same markers, the model may read it as a genuine user turn. Phantom and an earlier attack, ChatInject, both forge the markers; Phantom’s paper calls the technique structured template injection.
The template is also code. In MITRE ATLAS case study AML.CS0064, a research demonstration rather than a real attack, researchers poisoned the chat templates in GGUF files so that on a trigger the template injected the attacker’s instructions, with the weights untouched. Separately, llama-cpp-python 0.2.30 to 0.2.71 rendered a file’s template without Jinja’s sandbox, letting a template run code (CVE-2024-34359).
In practice
Behind a hosted API the provider applies the template, so the caller may never know its markers. ChatInject still worked against closed models “despite their unknown template structures”.
In gpt-oss, roles also carry rank. OpenAI’s guide says the model resolves conflicting instructions in the order system, developer, user, assistant, tool. So in gpt-oss, which turn a piece of text sits in decides how much weight its instructions get.
Where a user can shape the formatted prompt, the format can be broken on purpose. ChatBug, a 2024 study, altered or dropped required markers, or wrote the start of the model’s reply itself, and bypassed the safety training of eight models. Against Claude it used the API’s option to prefill a reply, and against GPT-3.5 direct access through Microsoft Azure; OpenAI’s own API allowed neither attack.
Questions and answers
Do I need to write a chat template to use a model through an API?
No. A hosted API takes a list of messages and applies the chat template on the provider's side. OpenAI's guide to its gpt-oss format says users going through an API or a provider like Ollama need not handle the formatting; the guide is written for people building their own inference code.
What is the difference between a chat template and a system prompt?
A system prompt is content: the instructions a developer gives a model. A chat template is the format that places that content in the token stream, marking it as a system turn. Some templates have had no system turn at all: Google's guide for Gemma 3 and earlier says its instruction-tuned models support only user and model roles, so system instructions go inside the first user turn. Gemma 4 added a system role.
Sources
- Chat templates (Transformers documentation)Hugging Face
- Writing a chat template (Transformers documentation)Hugging Face
- Automating Agent Hijacking via Structural Template InjectionarXiv, 18 Feb 2026
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsarXiv, 26 Sep 2025
- OpenAI Harmony Response FormatOpenAI
- Gemma formatting and system instructionsGoogle
- Gemma 4 Prompt FormattingGoogle
- ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat TemplatesarXiv, 17 Jun 2024
- MITRE ATLAS, AML.CS0064 Poisoned GGUF Templates: Inference-Time Supply Chain Attack (collection 2026.09)MITRE
- GHSA-56xg-wfcc-g829: llama-cpp-python vulnerable to Remote Code Execution by Server-Side Template Injection in Model MetadataGitHub, 13 May 2024