Overview

Model quantization and model pruning are both techniques for shrinking neural networks and speeding up inference, but they compress different things. Quantization keeps every weight but represents each one with fewer bits (e.g. FP32 → INT8), while pruning keeps full precision but removes weights, neurons, or channels judged unimportant. The two techniques are complementary and are frequently chained together in a single compression pipeline.

Comparison Diagram

Model QuantizationBefore: FP32 (32-bit)0.482-1.0370.917-0.203quantizeAfter: INT8 (8-bit)61-132117-26Same 4 values, fewer bits each≈4× smaller, faster mathModel PruningBefore: dense networkpruneAfter: sparse (pruned)Fewer neurons & connectionskeptremoved

Comparison Table

AspectModel QuantizationModel Pruning
Core mechanismReduces numeric precision of weights/activations (e.g. FP32 → INT8/INT4)Removes individual weights, neurons, or channels judged low-importance
What changesSame parameter count, smaller representation per valueFewer parameters; model becomes sparse or physically smaller
GranularityPer-tensor, per-channel, or per-group bit-width choicesUnstructured (single weights) vs structured (filters/channels/layers)
When appliedPost-training quantization (PTQ) or quantization-aware training (QAT)Iterative pruning during training or magnitude-based pruning after training, usually with fine-tuning
Hardware/runtime requirementNeeds low-precision kernel support (INT8 cores, TensorRT, XNNPACK)Unstructured pruning needs sparse-matrix kernels for real speedup; structured pruning runs on standard dense hardware
Compression achievedTypically 2-4x size reduction (FP32→INT8); INT4 pushes further at higher accuracy riskCan reach 50-90% sparsity, but unstructured sparsity often doesn’t translate to real speedup without special hardware
Accuracy impact & recoverySmall accuracy drop, usually recovered via calibration or QATLarger accuracy drop at high sparsity, recovered via iterative fine-tuning/retraining
CombinabilityOften applied last, to shrink an already-pruned model furtherOften applied first, before the pruned model is quantized

Key Differences

  • Quantization changes each value’s bit-width; pruning changes the model’s parameter count.
  • Realizing pruning’s theoretical speedup often requires sparse kernels, while quantization’s speedup comes from standard INT8 hardware support.
  • Quantization degrades accuracy gradually and predictably; aggressive pruning risks a sharp accuracy cliff without fine-tuning.
  • The two are commonly chained into a single compression pipeline, pruning first and quantizing the result.
  • Structured pruning changes the model’s architecture shape; quantization never touches the architecture.

When to Use Each

Model Quantization

  • Deploying to edge/mobile hardware: INT8 accelerators are ubiquitous on mobile and edge chips, giving guaranteed speedup with minimal engineering effort.
  • Shrinking model for storage/bandwidth: Cuts file size roughly 2-4x without touching the architecture or parameter count.
  • Fast, low-risk compression: Post-training quantization can be applied in minutes with a small calibration set and usually recovers most accuracy.

Model Pruning

  • Removing redundant capacity: Overparameterized models with many near-zero weights can shed parameters with little accuracy loss.
  • Structured pruning for latency: Removing whole channels or filters yields real speedup on standard hardware, unlike unstructured sparsity.
  • Model slimming for redeploys: Iterative pruning can reveal a genuinely smaller architecture, useful when the same model is repeatedly retrained and shipped.