Quantization
Quantization stores and computes model weights (and sometimes activations or the KV cache) at lower numeric precision, for example 8-bit or 4-bit instead of 16-bit, to cut memory and often speed up inference. It is what makes large open-weight models fit on a laptop or a single GPU.
Main approaches
- Post-training quantization of a finished model: LLM.int8() (2022) handled outlier features for 8-bit inference; GPTQ (2022) does accurate 3-4 bit weight quantization.
- Quantised fine-tuning: QLoRA (2023) fine-tunes LoRA adapters on a 4-bit base model.
- Formats and runtimes: GGUF files for llama.cpp-style local runtimes such as ollama; GPU servers like vllm support several quantised formats.
Trade-off: smaller footprint versus some quality loss that grows at lower bit widths and on harder tasks; test on your own workload. See inference-and-serving for engines and distillation for the other main way to get small models.
Related: open-source-ai, large-language-model, ann-index-algorithms (product quantization for vector search is a different use of the term).
Sources
- https://arxiv.org/abs/2208.07339 (LLM.int8(), 2022)
- https://arxiv.org/abs/2210.17323 (GPTQ, 2022)
- https://arxiv.org/abs/2305.14314 (QLoRA, 2023)
- Titles, authors and dates re-checked against the arXiv API on 2026-10-07 (primary): LLM.int8() submitted 2022-08-15; GPTQ 2022-10-31; QLoRA 2023-05-23.