Overview
Model quantization and model pruning are both techniques for shrinking neural networks and speeding up inference, but they compress different things. Quantization keeps every weight but represents each one with fewer bits (e.g. FP32 → INT8), while pruning keeps full precision but removes weights, neurons, or channels judged unimportant. The two techniques are complementary and are frequently chained together in a single compression pipeline.
Comparison Diagram
Comparison Table
| Aspect | Model Quantization | Model Pruning |
|---|---|---|
| Core mechanism | Reduces numeric precision of weights/activations (e.g. FP32 → INT8/INT4) | Removes individual weights, neurons, or channels judged low-importance |
| What changes | Same parameter count, smaller representation per value | Fewer parameters; model becomes sparse or physically smaller |
| Granularity | Per-tensor, per-channel, or per-group bit-width choices | Unstructured (single weights) vs structured (filters/channels/layers) |
| When applied | Post-training quantization (PTQ) or quantization-aware training (QAT) | Iterative pruning during training or magnitude-based pruning after training, usually with fine-tuning |
| Hardware/runtime requirement | Needs low-precision kernel support (INT8 cores, TensorRT, XNNPACK) | Unstructured pruning needs sparse-matrix kernels for real speedup; structured pruning runs on standard dense hardware |
| Compression achieved | Typically 2-4x size reduction (FP32→INT8); INT4 pushes further at higher accuracy risk | Can reach 50-90% sparsity, but unstructured sparsity often doesn’t translate to real speedup without special hardware |
| Accuracy impact & recovery | Small accuracy drop, usually recovered via calibration or QAT | Larger accuracy drop at high sparsity, recovered via iterative fine-tuning/retraining |
| Combinability | Often applied last, to shrink an already-pruned model further | Often applied first, before the pruned model is quantized |
Key Differences
- Quantization changes each value’s bit-width; pruning changes the model’s parameter count.
- Realizing pruning’s theoretical speedup often requires sparse kernels, while quantization’s speedup comes from standard INT8 hardware support.
- Quantization degrades accuracy gradually and predictably; aggressive pruning risks a sharp accuracy cliff without fine-tuning.
- The two are commonly chained into a single compression pipeline, pruning first and quantizing the result.
- Structured pruning changes the model’s architecture shape; quantization never touches the architecture.
When to Use Each
Model Quantization
- Deploying to edge/mobile hardware: INT8 accelerators are ubiquitous on mobile and edge chips, giving guaranteed speedup with minimal engineering effort.
- Shrinking model for storage/bandwidth: Cuts file size roughly 2-4x without touching the architecture or parameter count.
- Fast, low-risk compression: Post-training quantization can be applied in minutes with a small calibration set and usually recovers most accuracy.
Model Pruning
- Removing redundant capacity: Overparameterized models with many near-zero weights can shed parameters with little accuracy loss.
- Structured pruning for latency: Removing whole channels or filters yields real speedup on standard hardware, unlike unstructured sparsity.
- Model slimming for redeploys: Iterative pruning can reveal a genuinely smaller architecture, useful when the same model is repeatedly retrained and shipped.