What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Quantization

Quantization is a technique that stores a model's weights, and sometimes the intermediate values it computes, in fewer bits than training used, such as 8-bit integers or 4-bit codes instead of 32- or 16-bit floating-point numbers. Quantization cuts the memory and compute needed to run a model, at the cost of a small rounding error in every weight.

Last reviewed

Key points

  • Quantization stores a model's weights in fewer bits, such as 4 or 8 instead of 16 or 32, so a large model fits on smaller hardware.
  • Every weight is rounded, so the quantized model is a slightly different model from the one that was trained and tested.
  • Researchers built models that behave safely in full precision and turn malicious only once quantized, including in several GGUF k-quant types used by llama.cpp and Ollama.
  • In two studies, quantizing to 2 or 3 bits weakened safety behaviour, while 4 and 8 bits mostly kept it.
  • Safety-test the quantized file you actually run. Accuracy benchmarks, and tests of the full-precision original, can miss the change.

How it works

A trained model’s weights are usually 32- or 16-bit floating-point numbers. Quantization maps them onto a much smaller set of values. An 8-bit integer holds only 256 values, so each weight is divided by a scale and rounded to the nearest one.

The saving is large. GPTQ cut a 175-billion-parameter model to 3 or 4 bits per weight “with negligible accuracy degradation”. At 3 bits the whole model fit on a single 80 GB GPU.

Zero-shot methods such as LLM.int8() just scale and round, so users can quantize a downloaded model themselves. Optimization-based methods such as GPTQ tune the rounding against an error measure and are usually run once and shared already quantized. GGUF, the format llama.cpp loads, has its own types: Q4_K stores weights in blocks of 32, each with its own scale and minimum, at 4.5 bits per weight.

Why it matters

The quantized model a user runs is not exactly the model that was tested.

That gap can hide a backdoor. Egashira and colleagues (NeurIPS 2024) trained models that behave benignly in full precision and turn malicious once quantized. In one example the full-precision model wrote secure code 82.6% of the time and its LLM.int8() version “less than 3% of the time”. A follow-up (ICML 2025) did the same for nine GGUF types, concluding that “the complexity of quantization schemes alone is insufficient as a defense”. It is a backdoor whose trigger is quantization itself, and the attacker must control the model’s training.

Heavy quantization can weaken safety. Hong and colleagues (ICML 2024) found a 4-bit model “retains the trustworthiness of its original counterpart”, while 3-bit quantization “tends to reduce trustworthiness significantly”, a risk that “cannot be uncovered by looking at benign performance alone”. Researchers at Enkrypt AI found 2-bit GGUF versions of three Llama models far easier to jailbreak.

In practice

Post-training quantization “has become the standard for memory-efficient deployment”, in the words of the GGUF attack paper. The security risks of running LLMs locally apply to quantized files, plus these:

  • Safety-test the quantized version. Hong and colleagues warn that the risk at 3 bits “cannot be uncovered by looking at benign performance alone”, so a benchmark score is not enough. Testing only the full-precision model also misses a copy built to misbehave once quantized.
  • Moderate quantization is not reliably safer. In the Enkrypt AI study, 4-bit and 8-bit versions of two Llama 3 models resisted jailbreaks better than the originals (38 to 50% attack success against 62 to 64%). Llama 2 did not improve: 50% at 4-bit and 60% at 8-bit against 48%, on a test set of 50 prompts.
  • Noise is a promising defence, not a settled one. Adding small random noise to the weights before quantizing removed the insecure-code attack in both Egashira papers’ tests with little loss in benchmark scores. Other attack types were not tested against it. The right noise level differed by model, and the authors say effects “beyond benchmark performance of the noise addition remain unclear”.
  • Quantization is not the only risk in a GGUF file. Its parser and chat template have their own attack history, covered on the GGUF page.

Questions and answers

Does quantization make a model less safe?

It depends on how far it goes. In the studies by Hong and colleagues and by Enkrypt AI, 4-bit models mostly kept their safety behaviour, while 3-bit models tended to lose trustworthiness and 2-bit GGUF versions of three Llama models were much easier to jailbreak. Separately, a model can be deliberately built to turn unsafe once quantized. Safety-test the quantized version you deploy.

Can a backdoor be hidden in quantization?

Yes, in research settings. Egashira and colleagues built models that behave benignly in full precision and write insecure code or inject content once quantized, first for simple scale-and-round methods such as LLM.int8() (NeurIPS 2024), then for nine GGUF types (ICML 2025). The attacker has to train and publish the model; quantization acts as the trigger.

What does a name like Q4_K_M mean in a GGUF file?

Q4_K names a 4-bit "k-quant", one of GGUF's quantization methods, which optimizes the rounding to reduce error. Hugging Face's table describes Q4_K as blocks of 32 weights, each block with its own scale and minimum, working out to 4.5 bits per weight. According to Egashira and colleagues, the suffixes _S, _M and _L "indicate the portion of layers quantized with higher bitwidth than N", where N is the bit count.

Sources

  1. Quantization (Hugging Face Optimum concept guide)Hugging Face
  2. GGUF specification (ggml docs/gguf.md)ggml-org
  3. GGUF (Hugging Face Hub documentation)Hugging Face
  4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScalearXiv, 15 Aug 2022
  5. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersarXiv, 31 Oct 2022
  6. Exploiting LLM QuantizationarXiv (NeurIPS 2024), 28 May 2024
  7. Mind the Gap: A Practical Attack on GGUF QuantizationarXiv (ICML 2025), 26 May 2025
  8. Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionarXiv (ICML 2024), 18 Mar 2024
  9. Fine-Tuning, Quantization, and LLMs: Navigating Unintended OutcomesarXiv, 5 Apr 2024

In the news