Definition · AI basics
llama.cpp
llama.cpp is an open-source C/C++ engine for running large language models, built on the GGML tensor library. llama.cpp loads models stored as GGUF files, supports 1.5-bit to 8-bit quantized weights on CPUs and GPUs, and ships llama-server, an HTTP server with OpenAI-compatible routes and a web interface.
Last reviewed
Key points
- llama.cpp is an open-source C/C++ engine that runs language models from GGUF files, usually quantized, on CPUs and GPUs.
- A model file from an unknown source is untrusted input. The GGUF parsing llama.cpp relies on has had repeated memory-corruption bugs that a crafted file could exploit, including a 2026 bypass of a 2025 fix.
- Its HTTP server, llama-server, listens only on the local machine by default and requires no API key unless you set one.
- The RPC backend, which spreads inference across machines, has had unauthenticated remote-code-execution bugs. The project calls it insecure and says to keep it off untrusted networks.
- By default the server builds each prompt with the chat template stored in the model file, so a poisoned template arrives with the model.
How it works
llama.cpp is a C/C++ program built on the GGML tensor library. It loads a model from a GGUF file and runs it on CPUs or GPUs. It supports weights quantized to between 1.5 and 8 bits to save memory and time.
It ships a command-line tool and llama-server, an OpenAI-compatible HTTP server. The server listens on 127.0.0.1:8080 by default and checks no API key unless one is configured, with --api-key or a key file. An RPC backend lets one machine hand computation to ggml-rpc-server processes on others.
llama-server builds each prompt with the Jinja chat template from the model’s own metadata unless you pass --chat-template.
Why it matters
Flaws have hit llama.cpp’s model loading and RPC backend, and research shows attacks carried inside the model file.
The model file. In January 2024 a Databricks researcher privately reported “potentially exploitable” heap overflows in GGML’s GGUF parser, reachable through llama.cpp, which an attacker “could leverage” to run code by serving a crafted file; fixes were merged six days later. More followed in 2025, in the parser (CVE-2025-53630) and in vocabulary loading (CVE-2025-49847). A 2026 bypass of the 2025 parser fix reached code execution (CVE-2026-27940).
The RPC backend. A 2024 flaw let a network client write to any memory address (CVE-2024-42479). In 2026 a code path with the same root cause gave unauthenticated remote code execution (CVE-2026-34159). The project’s RPC documentation calls the backend “fragile and insecure”.
The weights and the template. Egashira and colleagues (ICML 2025) built models that look benign in full precision and turn malicious once quantized to GGUF types, “used in the popular ollama and llama.cpp frameworks”. Since llama-server applies the model file’s own template by default, a poisoned template comes with the model. In MITRE ATLAS case study AML.CS0064, researchers poisoned GGUF templates to inject instructions on a trigger, weights untouched, on four engines ATLAS does not name. ATLAS maps the step where the engine runs the template to unsafe AI artifacts.
In practice
llama.cpp’s security policy asks users to run untrusted models in “a secure, isolated environment such as a sandbox”, for example a container or virtual machine. Where a model cannot be isolated or must face an untrusted network, it says to check the hash of any downloaded artifact, such as model weights, and not to use llama-server or the RPC backend. The Docker example binds the server to 0.0.0.0, so a container setup can expose it. Built-in agent tools, off by default, can read and write files and run shell commands.
Questions and answers
Is llama.cpp safe to run?
Running llama.cpp safely depends on the model files you give it, the inputs it handles and who can reach it. The GGUF parsing it relies on has had several memory-corruption bugs that a crafted file could exploit, so keep it updated and run untrusted models in a container or virtual machine, as its security policy advises. llama-server listens only on the local machine by default and has no API key unless you set one, but by default it answers requests from any web origin, so a page open in your browser can call it.
Should I expose llama-server to the internet?
Not if you can avoid it: llama.cpp's own security policy says not to use llama-server or the RPC backend on an untrusted network. Where a server is public anyway, the server documentation's CORS guidance says to set an API key and put it behind a reverse proxy. Leave the built-in agent tools (--tools, --agent) off: they can read and write files and run shell commands.
Sources
- llama.cpp READMEggml-org
- LLaMA.cpp HTTP Server (tools/server/README.md)ggml-org
- llama.cpp Security Policy (SECURITY.md)ggml-org
- llama.cpp RPC (tools/rpc/README.md)ggml-org
- GGML GGUF File Format VulnerabilitiesDatabricks, 22 Mar 2024
- GHSA-vgg9-87g3-85w8: Integer Overflow in GGUF Parser can lead to Heap Out-of-Bounds Read/Write in ggufggml-org, 10 Jul 2025
- GHSA-3p4r-fq3f-q74v: Heap Buffer Overflow via Integer Overflow in mem_size Calculation — Bypass of CVE-2025-53630 Fixggml-org, 12 Mar 2026
- GHSA-8wwf-w4qm-gpqr: Buffer Overflow in llama.cpp via Malicious GGUF Model – Exploitable via Vocabulary Loadingggml-org, 14 Jun 2025
- GHSA-wcr5-566p-9cwj: Write-what-where in rpc_server::set_tensorggml-org, 12 Aug 2024
- GHSA-j8rj-fmpv-wcxw: Unauthenticated RCE via GRAPH_COMPUTE buffer=0 bypass in llama.cpp RPC backendggml-org, 26 Mar 2026
- Mind the Gap: A Practical Attack on GGUF QuantizationEgashira et al., ICML 2025, May 2025
- MITRE ATLAS, AML.CS0064 Poisoned GGUF Templates: Inference-Time Supply Chain Attack (collection 2026.09)MITRE