The model quantization process is a technique for compressing a model, whereby the weights and the activations of the neural network are converted from a higher precision number format, such as 32-bit floating point numbers, to a lower precision number format, such as 8-bit or 4-bit integers. This reduces the amount of memory used by the model and the computational load to execute it. For example, a 7-billion parameter model, when stored in 16-bit precision format, requires about 14GB of memory.
What Will I Learn?
What Is Model Quantization?
Quantization of a model refers to the transformation of a model’s values from having a large number of bits to fewer bits. Training is done on models using 32-bit floating point (FP32), where all the values of the model are held using 32 bits. In the case of quantization, we transform those values to formats such as 16 bits (FP16, BF16), 8 bits (INT8), or 4 bits (INT4).
Quantization is not the same as compression. Compression (for example ZIP) holds all the values as in the original file, with lossless transformation of data. However, in quantization, we lose precision permanently — we will never be able to recover an exact value of FP32 from the quantized value. What quantization aims at is losing precision that a model does not need for inference but retaining what is needed.
Why Model Quantization Matters
Quantization solves three measurable expenses associated with neural network execution: memory, latency, and energy consumption.
- Memory footprint. The storage size of each parameter grows linearly with bit-width. The parameter saved using FP32 takes up 4 bytes, whereas the same parameter saved as INT8 requires 1 byte – 4 times less storage with no other changes.
- Inference Latency. Low-precision arithmetic is faster than FP32 arithmetic in processors that can perform low-precision computations, such as GPUs equipped with tensor cores supporting INT8/FP8 and mobile NPUs. In particular, Frantar et al. (2022) found that applying GPTQ quantization to large language models provided a 3.25x speedup of inference on the NVIDIA A100 GPU and 4.5x speedup on the NVIDIA A6000 GPU.
- Power consumption. Reduced number of operations in terms of bytes moved through the memory and decreased bit width of calculations require less energy to perform each step.
Below is shown the precise amount of memory used for storage of model parameters for each of the precisions listed.
Table 1. Model weight memory by precision (weights only, no activation or framework overhead)
| Precision | Bytes per parameter | 7B-parameter model | 70B-parameter model |
|---|---|---|---|
| FP32 | 4 bytes | 28 GB | 280 GB |
| FP16 / BF16 | 2 bytes | 14 GB | 140 GB |
| INT8 | 1 byte | 7 GB | 70 GB |
| INT4 | 0.5 bytes | 3.5 GB | 35 GB |
These figures cover weight storage only. Actual deployment memory also includes activation memory and, for transformer decoder models, the KV cache — both are addressed in the next section.
What Gets Quantized: Weights, Activations, and the KV Cache
There are three elements of a neural network that can be quantized, which are weights, activations, and the KV cache for transformer decoder models.
Weights are fixed and not influenced by the input data. Weight quantization is the easiest kind of quantization as it does not need any other data but the trained model itself.
Activities refer to the intermediate output generated at each layer while inferring on a test set. Activities differ from weights in that their range of values changes depending on the input the model gets. In order to correctly scale the activities, real data observation is needed, and this procedure is described in the calibration part below.
The KV cache keeps the key and value tensors produced by the decoder model for each token. The KV cache size is proportional to the length of a sequence and the number of layers and attention heads in the model. For example, the 7B parameter model processing a context of 4,096 tokens will increase the total memory usage by several gigabytes due to the KV cache along with the memory consumed by weights described in Table 1.
The Precision Formats Behind Quantization: FP32 to INT4
The sign bit, the exponent, and the mantissa (also called the significand) make up the bits in any floating point number representation. The sign bit represents whether the number is positive or negative. The exponent decides the range of numbers that can be represented. The mantissa decides how fine-grained the numbers are.
Table 2. Common precision formats used in model quantization
| Format | Total bits | Sign | Exponent | Mantissa | Typical role |
|---|---|---|---|---|---|
| FP32 | 32 | 1 | 8 | 23 | Training precision; the baseline before quantization |
| FP16 | 16 | 1 | 5 | 10 | Common inference precision; narrower range than FP32 |
| BF16 | 16 | 1 | 8 | 7 | Matches FP32’s exponent range with less mantissa precision |
| FP8 (E4M3) | 8 | 1 | 4 | 3 | Weight and activation quantization on supporting hardware |
| INT8 | 8 | — | — | — | Integer format; requires a scale factor to represent real values |
| INT4 | 4 | — | — | — | Integer format; the most common target for large language model compression |
How the Quantization Math Works
To convert from the real value to the quantized value, two parameters are needed: scale factor and zero point. Scale specifies how many units of real values will correspond to one unit of quantized values. Zero point identifies the quantized value which corresponds to the real value of zero.
There are two types of quantization with different parameters usage:
Affine (asymmetric) quantization works with a non-zero zero-point and allows the quantized range to be shifted, so it can start not from zero but from another point. This type of quantization fits cases where the data distribution is not symmetric around zero.
Symmetric quantization implies fixing the zero-point to zero so that the real value of zero will map to the quantized value of zero. Symmetric quantization is easier for calculation and it is used as default in NVIDIA’s TensorRT and Model Optimizer tools since there is no accuracy gain from using asymmetric quantization for the vast majority of models.
Scale is often calculated by the AbsMax algorithm: you need to take the maximum value of absolute values in the data that will be quantized and divide the target format range to it.
Quantization Granularity: Per-Tensor, Per-Channel, Per-Block
Granularity determines how many scale factors a quantization scheme uses across a single weight tensor.
- Per-tensor (per-layer) quantization uses only one scale factor on a whole tensor. This is the easiest way to quantify data; it consumes less memory but creates more errors for cases where values differ much inside the tensor.
- Per-channel quantization uses different scale factors for each channel; the channels usually go along the output channel dimension. Thus, the outlier values affect only this channel and not all values of the tensor.
- Per-block (per-group) quantization splits the tensor into smaller parts, where each part has its own scale factor. This provides the most flexible approach; it is used in the GGUF format with the block size of 32 or 256 elements per scale factor.
Table 3. Quantization granularity comparison
| Granularity | Scale factors per tensor | Accuracy | Memory overhead |
|---|---|---|---|
| Per-tensor | 1 | Lowest | Lowest |
| Per-channel | 1 per channel | Higher | Moderate |
| Per-block / per-group | 1 per block (e.g., every 32 or 128 values) | Highest | Highest |
Post-Training Quantization vs. Quantization-Aware Training
There are two methods that can be used to develop a quantized model, and they vary on the timing of quantization with respect to training.
Post-Training Quantization (PTQ) applies quantization to an already-trained model, without further training. The PTQ method is easier to implement as it does not require any training infrastructure; it just needs a calibration data set.
Quantization-Aware Training (QAT) simulates quantization effects during training itself, so the model’s weights adapt to the precision loss before deployment. In QAT, there are fake quantization modules included in the forward and backward pass; these modules quantize and dequantize the weights and activations instantly, thus allowing the model to experience quantization effects before gradients are propagated with full precision. However, quantization functions are non-differentiable, thus QAT makes use of the Straight-Through Estimator (STE).
Table 4. PTQ vs. QAT
| Factor | Post-Training Quantization (PTQ) | Quantization-Aware Training (QAT) |
|---|---|---|
| When applied | After training completes | During training |
| Training infrastructure required | No | Yes |
| Data required | A calibration dataset | Full or partial training dataset |
| Typical accuracy retention | Good; depends on calibration quality | Higher than PTQ at the same bit-width |
| Implementation speed | Fast — hours | Slow — requires a training run |
| Common use case | Production deployment of pretrained models | New models where maximum accuracy at low bit-width is required |
Static vs. Dynamic Quantization
Within PTQ, there are two more methods for when the scale factors are calculated.
In static quantization, the scale factors are calculated once from a calibration set and reused every time without any additional calculation cost during inference. A good calibration set is needed, however.
In dynamic quantization, the scale factors are calculated during inference for each input individually. It does not need any calibration set but incurs a slight calculation cost.
The Main Quantization Algorithms: GPTQ, AWQ, and SmoothQuant
There are three algorithms responsible for most large language model quantization as of 2026.
The GPTQ (Generative Pre-trained Transformer Quantization) algorithm was proposed by Frantar, Ashkboos, Hoefler, and Alistarh in October 2022 and presented at ICLR 2023. GPTQ quantizes rows of a weight matrix independently and uses approximate second-order information (an approximation of a Hessian matrix) to choose the updates of weights that minimize the subsequent output error. According to the original paper, a model with 175 billion parameters was quantized to 3- or 4-bit precision in about four hours of GPU time with an accuracy drop described as negligible compared to the baseline of FP16 precision.
The AWQ (Activation-aware Weight Quantization) algorithm was developed by Lin, Tang, et al. from MIT in June 2023 and won the MLSys 2024 Best Paper Award. AWQ determines the small number of channels (reported in the paper to be about 1%) that have the most significant influence on the output based on the activation magnitude and preserves their precision through per-channel scaling and then quantizes the rest of the weights.
SmoothQuant was proposed by Xiao, Lin et al., MIT in November 2022 and was presented at ICML 2023. SmoothQuant allows for 8-bit quantization of both the weights and activations (W8A8) through the mathematical shifting of quantization problems from activations where outliers are prevalent to the weights where outliers are infrequent. According to the original paper, the solution achieved up to 1.56x speedups and half the memory overhead in comparison with FP16, with almost negligible accuracy degradation and served a model of 530 billion parameters with fewer GPUs than the non-quantized one.
Model Quantization vs. Pruning vs. Knowledge Distillation
Quantization is one of three common model-compression techniques, and each modifies a model differently.
Table 5. Model compression technique comparison
| Technique | What it changes | Result |
|---|---|---|
| Quantization | Numeric precision of weights and/or activations | Same architecture, smaller memory footprint per parameter |
| Pruning | Network structure — removes weights or connections | Sparse network with fewer total parameters |
| Knowledge distillation | Model size — trains a smaller model to mimic a larger one | New, smaller model trained from scratch under supervision |
While both quantization and pruning act on different parts of the same model, these methods are often used together. However, knowledge distillation leads to a new, smaller model that can be quantized again later.
GPTQ vs. AWQ vs. GGUF vs. bitsandbytes: Which Tool Should You Use?
The algorithms described above are implemented through specific software tools, and each tool targets a different deployment scenario.
Table 6. Quantization tool comparison
| Tool | Underlying method | Primary hardware target | Typical bit-width | Requires calibration data |
|---|---|---|---|---|
| GPTQ | Second-order weight updates | GPU | 3–4 bit | Yes |
| AWQ | Activation-aware channel protection | GPU (including edge GPUs) | 4 bit | Yes |
| bitsandbytes (LLM.int8() / QLoRA) | Vector-wise INT8; NF4 for 4-bit | GPU | 8 bit or 4 bit (NF4) | No (LLM.int8()); no (QLoRA/NF4) |
| GGUF (via llama.cpp) | Block-wise scale quantization | CPU (GPU offload supported) | 2–8 bit, multiple levels per file | Varies by quant level |
In 2022, Dettmers et al. presented LLM.int8(), claiming that their method was able to achieve 8-bit inference using half of the memory required by FP16 with outputs of equal quality to the original. In 2023, Dettmers et al. (University of Washington) presented QLoRA, along with the NF4 (4-bit NormalFloat) format and double quantization techniques, using the open-source bitsandbytes library. On August 21, 2023, Georgi Gerganov’s llama.cpp library presented the GGUF format as the evolution of the former GGML format, which was optimized only for CPU inference but also compatible with GPU offloading.
Decision framework:
- Determine the target deployment environment. CPU-based or edge deployment that doesn’t include a GPU suggests GGUF. Deployment on GPU allows using all four methods.
- Determine whether there’s a calibration data set. In case there isn’t, GPTQ and AWQ can be excluded since they both require a calibration data set; LLM.int8() and QLoRA/NF4 don’t require any calibration.
- Determine accuracy requirements at the target bit-width. AWQ, through its channel-protection technique, ensures higher accuracy compared to GPTQ at the same bit-width in the original benchmarks for the method, but comes with the disadvantage of needing activation statistics for calibration.
- Determine the ecosystem. GGUF is the preferred choice for use in tools like llama.cpp, LM Studio, and Ollama; GPTQ and AWQ work directly with Hugging Face Transformers; bitsandbytes works directly with Hugging Face Transformers and PEFT.
How to Quantize a Model: A Practical Walkthrough
The exact steps differ by deployment target. Two common paths are shown below: exporting a computer vision model to an edge-optimized format, and quantizing a large language model for GPU inference.
Computer vision model export (PTQ, INT8, edge deployment):
from ultralytics import YOLO
# Load a trained model
model = YOLO("yolo26n.pt")
# Export to TFLite with INT8 post-training quantization
# The calibration dataset determines the dynamic range for scale factors
model.export(format="tflite", int8=True, data="coco8.yaml")
Large language model quantization (PTQ, 4-bit, GPU inference):
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
# Configure 4-bit NF4 quantization
config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype="bfloat16",
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf",
quantization_config=config,
device_map="auto",
)
Both paths follow the same three-step sequence:
- Choose the target format and bit width according to the hardware determined from the decision framework above.
- Supply a calibration dataset reflective of actual inference input, if required by your quantization method.
- Evaluate the output quality relative to the unquantized baseline, utilizing task-specific accuracy metrics, or in the case of language models, perplexity.
Real-World Applications
Applications of quantization models include any situation in which inference has to be performed with limited memory, latency, or power resources.
- Edge and mobile computer vision. Models that can be converted into INT8 TFLite or ONNX format and used for object detection on mobile NPUs and embedded systems, such as automotive and robotics hardware, where inference using full-precision would exceed memory/power limitations.
- On-device Large Language Models. GGUF quantized models can be run locally on consumer hardware like laptops and phones using llama.cpp, eliminating the need for internet and an inference server.
- High-throughput LLM serving. Methods such as SmoothQuant and others from the W8A8 family let inference providers use fewer GPUs for serving larger models, reducing costs per GPU-hour.
When You Actually Need Quantization
Quantization is justified if at least one of the conditions below applies.
- Inference cost exceeds budget at full precision. Inference cost of a model at FP16/FP32 level of precision is higher than at INT8/INT4 with no change in quality that is described as minimal by the original research.
- The Model does not fit on available hardware. While a model may require 140GB of VRAM at FP16 (see Table 1, 70B-parameters line), the same model may require only approximately 35GB of VRAM after quantization to INT4.
- The deployment target lacks a dedicated GPU. Most of the premises, such as on-premises servers, edge devices, or local development computers, use CPU inference, and quantization with the GGUF format aims at supporting this.
- Multiple models should be running on the same hardware. The smaller amount of memory required for each model allows for running more models simultaneously on the same GPUs.
Common Pitfalls and Tradeoffs
Quantization always presents an inherent compromise between efficiency and effectiveness. There are three main sources of such a compromise that usually become problematic.
Calibration data that does not represent production inputs produces inaccurate scale factors. The scale factors based on an unrepresentative set of calibration examples are incapable of covering the actual range of values of the input data of the model during inference, which means higher amounts of clipping and quantization errors for those unseen ranges of inputs.
Over-quantization below the accuracy floor for a given architecture degrades output quality measurably. Almost all transformer models keep high levels of accuracy at 8-bit and 4-bit quantization (see the algorithm-specific results above), but any attempt to go below 4-bit without using special algorithms like the GPTQ-based 2- or 3-bit quantization regime may lead to output quality losses.
Hardware and quantization scheme mismatches remove expected speed gains. If a model is quantized in a way that it cannot be run directly on the target hardware, dequantization will be required, thus ruining any benefits of quantization.
Frequently Asked Questions
Q1. Does quantization reduce model accuracy?
Ans. Quantization may decrease the accuracy, but according to the original research on GPTQ, AWQ, and SmoothQuant, the decrease in accuracy is considered small to insignificant when the weights are quantized to 8-bits or 4-bits through special quantization methods, not just naive rounding.
Q2. Do I need to retrain a model to quantize it?
Ans. No – PTQ is applied to a pre-trained model with no re-training needed; only QAT requires additional training.
Q3. Does quantization work without a GPU?
Ans. Yes – the GGUF format that is used in the llama.cpp project is specifically created for CPU inference with an optional GPU layer offload.
Q4. Is quantization reversible?
Ans. No – quantization loses the precision irreversibly; a quantized number can only be reverted back to an approximation of the original range and not to the original FP32 number.
Q5. GPTQ, AWQ, or GGUF — which should I use?
Ans. Choose GGUF for CPU/edge deployment, GPTQ or AWQ for GPU deployment in case you have a calibration dataset available, and AWQ in case the benefit of the original research in terms of improved accuracy with the same bit-width outweighs the need for calibration.
Q6. How much smaller does a model get after quantization?
Ans. The amount of weight memory occupied by the model reduces by precisely the factor of bytes per parameter of the formats ratio, 4 times from FP32 to INT8 and 8 times from FP32 to INT4 (Table 1).
Conclusion
The process of quantization involves reducing the numerical precision of a trained model to save on memory, latency, and power, by using one of two methods: Post-Training Quantization or Quantization-Aware Training, accomplished by using one of the following algorithms: GPTQ, AWQ, and SmoothQuant. The selection between these algorithms is determined by three fixed variables: hardware used, calibration data availability, and minimum accuracy tolerance level.