<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Inference-Optimization on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/inference-optimization/</link><description>Recent content in Inference-Optimization on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:48:01 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/inference-optimization/index.xml" rel="self" type="application/rss+xml"/><item><title>Model Quantization vs Model Pruning: Fewer Bits vs Fewer Parameters</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-model-quantization-vs-model-pruning-fewer-bits-vs-fewer-para/</link><pubDate>Mon, 03 Aug 2026 03:48:01 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-model-quantization-vs-model-pruning-fewer-bits-vs-fewer-para/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Model quantization and model pruning are both techniques for shrinking neural networks and speeding up inference, but they compress different things. &lt;strong class="kw"&gt;Quantization&lt;/strong&gt; keeps every weight but represents each one with fewer bits (e.g. FP32 → INT8), while &lt;strong class="kw"&gt;pruning&lt;/strong&gt; keeps full precision but removes weights, neurons, or channels judged unimportant. The two techniques are complementary and are frequently chained together in a single compression pipeline.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;line x1="320" y1="40" x2="320" y2="300" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Model Quantization&lt;/text&gt;&lt;text x="160" y="50" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Before: FP32 (32-bit)&lt;/text&gt;&lt;rect x="62" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="82" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;0.482&lt;/text&gt;&lt;rect x="114" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="134" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-1.037&lt;/text&gt;&lt;rect x="166" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="186" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;0.917&lt;/text&gt;&lt;rect x="218" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-0.203&lt;/text&gt;&lt;line x1="160" y1="104" x2="160" y2="132" style="stroke:var(--secondary)" stroke-width="2"/&gt;&lt;polygon points="153,132 167,132 160,142" style="fill:var(--secondary)"/&gt;&lt;text x="172" y="122" style="fill:var(--secondary)" font-size="10"&gt;quantize&lt;/text&gt;&lt;text x="160" y="158" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;After: INT8 (8-bit)&lt;/text&gt;&lt;rect x="69" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="82" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;61&lt;/text&gt;&lt;rect x="121" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="134" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-132&lt;/text&gt;&lt;rect x="173" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="186" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;117&lt;/text&gt;&lt;rect x="225" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-26&lt;/text&gt;&lt;text x="160" y="220" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Same 4 values, fewer bits each&lt;/text&gt;&lt;text x="160" y="236" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;≈4× smaller, faster math&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Model Pruning&lt;/text&gt;&lt;text x="480" y="50" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Before: dense network&lt;/text&gt;&lt;g style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"&gt;&lt;line x1="420" y1="65" x2="480" y2="55"/&gt;&lt;line x1="420" y1="65" x2="480" y2="87"/&gt;&lt;line x1="420" y1="65" x2="480" y2="120"/&gt;&lt;line x1="420" y1="110" x2="480" y2="55"/&gt;&lt;line x1="420" y1="110" x2="480" y2="87"/&gt;&lt;line x1="420" y1="110" x2="480" y2="120"/&gt;&lt;line x1="480" y1="55" x2="540" y2="65"/&gt;&lt;line x1="480" y1="55" x2="540" y2="110"/&gt;&lt;line x1="480" y1="87" x2="540" y2="65"/&gt;&lt;line x1="480" y1="87" x2="540" y2="110"/&gt;&lt;line x1="480" y1="120" x2="540" y2="65"/&gt;&lt;line x1="480" y1="120" x2="540" y2="110"/&gt;&lt;/g&gt;&lt;g style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;circle cx="420" cy="65" r="5"/&gt;&lt;circle cx="420" cy="110" r="5"/&gt;&lt;circle cx="480" cy="55" r="5"/&gt;&lt;circle cx="480" cy="87" r="5"/&gt;&lt;circle cx="480" cy="120" r="5"/&gt;&lt;circle cx="540" cy="65" r="5"/&gt;&lt;circle cx="540" cy="110" r="5"/&gt;&lt;/g&gt;&lt;line x1="480" y1="138" x2="480" y2="166" style="stroke:var(--secondary)" stroke-width="2"/&gt;&lt;polygon points="473,166 487,166 480,176" style="fill:var(--secondary)"/&gt;&lt;text x="492" y="156" style="fill:var(--secondary)" font-size="10"&gt;prune&lt;/text&gt;&lt;text x="480" y="190" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;After: sparse (pruned)&lt;/text&gt;&lt;line x1="420" y1="240" x2="480" y2="185" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="3,3"/&gt;&lt;line x1="480" y1="185" x2="540" y2="240" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="3,3"/&gt;&lt;g style="stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;line x1="420" y1="195" x2="480" y2="185"/&gt;&lt;line x1="420" y1="195" x2="480" y2="217"/&gt;&lt;line x1="420" y1="240" x2="480" y2="217"/&gt;&lt;line x1="480" y1="185" x2="540" y2="195"/&gt;&lt;line x1="480" y1="217" x2="540" y2="195"/&gt;&lt;line x1="480" y1="217" x2="540" y2="240"/&gt;&lt;/g&gt;&lt;circle cx="480" cy="250" r="5" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="2,2"/&gt;&lt;g style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;circle cx="420" cy="195" r="5"/&gt;&lt;circle cx="420" cy="240" r="5"/&gt;&lt;circle cx="480" cy="185" r="5"/&gt;&lt;circle cx="480" cy="217" r="5"/&gt;&lt;circle cx="540" cy="195" r="5"/&gt;&lt;circle cx="540" cy="240" r="5"/&gt;&lt;/g&gt;&lt;text x="480" y="277" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Fewer neurons &amp;amp; connections&lt;/text&gt;&lt;line x1="380" y1="300" x2="400" y2="300" style="stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;text x="405" y="304" style="fill:var(--secondary)" font-size="10"&gt;kept&lt;/text&gt;&lt;line x1="450" y1="300" x2="470" y2="300" style="stroke:var(--border)" stroke-width="2" stroke-dasharray="3,3"/&gt;&lt;text x="475" y="304" style="fill:var(--secondary)" font-size="10"&gt;removed&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Model Quantization&lt;/th&gt;
&lt;th&gt;Model Pruning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core mechanism&lt;/td&gt;
&lt;td&gt;Reduces numeric precision of weights/activations (e.g. FP32 → INT8/INT4)&lt;/td&gt;
&lt;td&gt;Removes individual weights, neurons, or channels judged low-importance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What changes&lt;/td&gt;
&lt;td&gt;Same parameter count, smaller representation per value&lt;/td&gt;
&lt;td&gt;Fewer parameters; model becomes sparse or physically smaller&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Granularity&lt;/td&gt;
&lt;td&gt;Per-tensor, per-channel, or per-group bit-width choices&lt;/td&gt;
&lt;td&gt;Unstructured (single weights) vs structured (filters/channels/layers)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When applied&lt;/td&gt;
&lt;td&gt;Post-training quantization (PTQ) or quantization-aware training (QAT)&lt;/td&gt;
&lt;td&gt;Iterative pruning during training or magnitude-based pruning after training, usually with fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware/runtime requirement&lt;/td&gt;
&lt;td&gt;Needs low-precision kernel support (INT8 cores, TensorRT, XNNPACK)&lt;/td&gt;
&lt;td&gt;Unstructured pruning needs sparse-matrix kernels for real speedup; structured pruning runs on standard dense hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compression achieved&lt;/td&gt;
&lt;td&gt;Typically 2-4x size reduction (FP32→INT8); INT4 pushes further at higher accuracy risk&lt;/td&gt;
&lt;td&gt;Can reach 50-90% sparsity, but unstructured sparsity often doesn&amp;rsquo;t translate to real speedup without special hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy impact &amp;amp; recovery&lt;/td&gt;
&lt;td&gt;Small accuracy drop, usually recovered via calibration or QAT&lt;/td&gt;
&lt;td&gt;Larger accuracy drop at high sparsity, recovered via iterative fine-tuning/retraining&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combinability&lt;/td&gt;
&lt;td&gt;Often applied last, to shrink an already-pruned model further&lt;/td&gt;
&lt;td&gt;Often applied first, before the pruned model is quantized&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Quantization changes each value&amp;rsquo;s &lt;strong class="kw"&gt;bit-width&lt;/strong&gt;; pruning changes the model&amp;rsquo;s &lt;strong class="kw"&gt;parameter count&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Realizing pruning&amp;rsquo;s theoretical speedup often requires &lt;strong class="kw"&gt;sparse kernels&lt;/strong&gt;, while quantization&amp;rsquo;s speedup comes from standard &lt;strong class="kw"&gt;INT8 hardware&lt;/strong&gt; support.&lt;/li&gt;
&lt;li&gt;Quantization degrades accuracy gradually and predictably; aggressive pruning risks a &lt;strong class="kw"&gt;sharp accuracy cliff&lt;/strong&gt; without fine-tuning.&lt;/li&gt;
&lt;li&gt;The two are commonly chained into a single &lt;strong class="kw"&gt;compression pipeline&lt;/strong&gt;, pruning first and quantizing the result.&lt;/li&gt;
&lt;li&gt;Structured pruning changes the model&amp;rsquo;s &lt;strong class="kw"&gt;architecture shape&lt;/strong&gt;; quantization never touches the architecture.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Model Quantization&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>