Quantization

Quantization stores and computes model weights (and sometimes activations or the KV cache) at lower numeric precision, for example 8-bit or 4-bit instead of 16-bit, to cut memory and often speed up inference. It is what makes large open-weight models fit on a laptop or a single GPU.

Main approaches

  • Post-training quantization of a finished model: LLM.int8() (2022) handled outlier features for 8-bit inference; GPTQ (2022) does accurate 3-4 bit weight quantization.
  • Quantised fine-tuning: QLoRA (2023) fine-tunes LoRA adapters on a 4-bit base model.
  • Formats and runtimes: GGUF files for llama.cpp-style local runtimes such as ollama; GPU servers like vllm support several quantised formats.

Trade-off: smaller footprint versus some quality loss that grows at lower bit widths and on harder tasks; test on your own workload. See inference-and-serving for engines and distillation for the other main way to get small models.

Related: open-source-ai, large-language-model, ann-index-algorithms (product quantization for vector search is a different use of the term).

Sources