Skip to content
The Internet Compass

AI

Quantization

Quantization is the process of reducing the numerical precision used to store and compute a model's weights — for example converting 16-bit floating-point values to 8-bit or 4-bit integers — which shrinks the model's memory footprint and can significantly speed up inference, usually at a small cost to output quality.

Quantization is what makes it practical to run large models on consumer hardware or at lower cost in production: a model quantized from 16-bit to 4-bit precision can require roughly a quarter of the memory, letting it fit on GPUs (or even CPUs) that couldn't hold the full-precision version at all.

The accuracy cost varies by technique and by task: well-implemented modern quantization methods often produce output close to indistinguishable from the full-precision model on many tasks, though degradation tends to show up first on tasks requiring precise numerical reasoning or very long context.