<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Machine-Learning on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/machine-learning/</link><description>Recent content in Machine-Learning on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:48:01 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Model Quantization vs Model Pruning: Fewer Bits vs Fewer Parameters</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-model-quantization-vs-model-pruning-fewer-bits-vs-fewer-para/</link><pubDate>Mon, 03 Aug 2026 03:48:01 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-model-quantization-vs-model-pruning-fewer-bits-vs-fewer-para/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Model quantization and model pruning are both techniques for shrinking neural networks and speeding up inference, but they compress different things. &lt;strong class="kw"&gt;Quantization&lt;/strong&gt; keeps every weight but represents each one with fewer bits (e.g. FP32 → INT8), while &lt;strong class="kw"&gt;pruning&lt;/strong&gt; keeps full precision but removes weights, neurons, or channels judged unimportant. The two techniques are complementary and are frequently chained together in a single compression pipeline.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;line x1="320" y1="40" x2="320" y2="300" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Model Quantization&lt;/text&gt;&lt;text x="160" y="50" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Before: FP32 (32-bit)&lt;/text&gt;&lt;rect x="62" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="82" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;0.482&lt;/text&gt;&lt;rect x="114" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="134" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-1.037&lt;/text&gt;&lt;rect x="166" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="186" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;0.917&lt;/text&gt;&lt;rect x="218" y="58" width="40" height="42" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="83" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-0.203&lt;/text&gt;&lt;line x1="160" y1="104" x2="160" y2="132" style="stroke:var(--secondary)" stroke-width="2"/&gt;&lt;polygon points="153,132 167,132 160,142" style="fill:var(--secondary)"/&gt;&lt;text x="172" y="122" style="fill:var(--secondary)" font-size="10"&gt;quantize&lt;/text&gt;&lt;text x="160" y="158" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;After: INT8 (8-bit)&lt;/text&gt;&lt;rect x="69" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="82" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;61&lt;/text&gt;&lt;rect x="121" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="134" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-132&lt;/text&gt;&lt;rect x="173" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="186" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;117&lt;/text&gt;&lt;rect x="225" y="166" width="26" height="26" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="183" text-anchor="middle" style="fill:var(--content)" font-size="9"&gt;-26&lt;/text&gt;&lt;text x="160" y="220" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Same 4 values, fewer bits each&lt;/text&gt;&lt;text x="160" y="236" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;≈4× smaller, faster math&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Model Pruning&lt;/text&gt;&lt;text x="480" y="50" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Before: dense network&lt;/text&gt;&lt;g style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"&gt;&lt;line x1="420" y1="65" x2="480" y2="55"/&gt;&lt;line x1="420" y1="65" x2="480" y2="87"/&gt;&lt;line x1="420" y1="65" x2="480" y2="120"/&gt;&lt;line x1="420" y1="110" x2="480" y2="55"/&gt;&lt;line x1="420" y1="110" x2="480" y2="87"/&gt;&lt;line x1="420" y1="110" x2="480" y2="120"/&gt;&lt;line x1="480" y1="55" x2="540" y2="65"/&gt;&lt;line x1="480" y1="55" x2="540" y2="110"/&gt;&lt;line x1="480" y1="87" x2="540" y2="65"/&gt;&lt;line x1="480" y1="87" x2="540" y2="110"/&gt;&lt;line x1="480" y1="120" x2="540" y2="65"/&gt;&lt;line x1="480" y1="120" x2="540" y2="110"/&gt;&lt;/g&gt;&lt;g style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;circle cx="420" cy="65" r="5"/&gt;&lt;circle cx="420" cy="110" r="5"/&gt;&lt;circle cx="480" cy="55" r="5"/&gt;&lt;circle cx="480" cy="87" r="5"/&gt;&lt;circle cx="480" cy="120" r="5"/&gt;&lt;circle cx="540" cy="65" r="5"/&gt;&lt;circle cx="540" cy="110" r="5"/&gt;&lt;/g&gt;&lt;line x1="480" y1="138" x2="480" y2="166" style="stroke:var(--secondary)" stroke-width="2"/&gt;&lt;polygon points="473,166 487,166 480,176" style="fill:var(--secondary)"/&gt;&lt;text x="492" y="156" style="fill:var(--secondary)" font-size="10"&gt;prune&lt;/text&gt;&lt;text x="480" y="190" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;After: sparse (pruned)&lt;/text&gt;&lt;line x1="420" y1="240" x2="480" y2="185" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="3,3"/&gt;&lt;line x1="480" y1="185" x2="540" y2="240" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="3,3"/&gt;&lt;g style="stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;line x1="420" y1="195" x2="480" y2="185"/&gt;&lt;line x1="420" y1="195" x2="480" y2="217"/&gt;&lt;line x1="420" y1="240" x2="480" y2="217"/&gt;&lt;line x1="480" y1="185" x2="540" y2="195"/&gt;&lt;line x1="480" y1="217" x2="540" y2="195"/&gt;&lt;line x1="480" y1="217" x2="540" y2="240"/&gt;&lt;/g&gt;&lt;circle cx="480" cy="250" r="5" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="2,2"/&gt;&lt;g style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"&gt;&lt;circle cx="420" cy="195" r="5"/&gt;&lt;circle cx="420" cy="240" r="5"/&gt;&lt;circle cx="480" cy="185" r="5"/&gt;&lt;circle cx="480" cy="217" r="5"/&gt;&lt;circle cx="540" cy="195" r="5"/&gt;&lt;circle cx="540" cy="240" r="5"/&gt;&lt;/g&gt;&lt;text x="480" y="277" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Fewer neurons &amp;amp; connections&lt;/text&gt;&lt;line x1="380" y1="300" x2="400" y2="300" style="stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;text x="405" y="304" style="fill:var(--secondary)" font-size="10"&gt;kept&lt;/text&gt;&lt;line x1="450" y1="300" x2="470" y2="300" style="stroke:var(--border)" stroke-width="2" stroke-dasharray="3,3"/&gt;&lt;text x="475" y="304" style="fill:var(--secondary)" font-size="10"&gt;removed&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Model Quantization&lt;/th&gt;
&lt;th&gt;Model Pruning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core mechanism&lt;/td&gt;
&lt;td&gt;Reduces numeric precision of weights/activations (e.g. FP32 → INT8/INT4)&lt;/td&gt;
&lt;td&gt;Removes individual weights, neurons, or channels judged low-importance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What changes&lt;/td&gt;
&lt;td&gt;Same parameter count, smaller representation per value&lt;/td&gt;
&lt;td&gt;Fewer parameters; model becomes sparse or physically smaller&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Granularity&lt;/td&gt;
&lt;td&gt;Per-tensor, per-channel, or per-group bit-width choices&lt;/td&gt;
&lt;td&gt;Unstructured (single weights) vs structured (filters/channels/layers)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When applied&lt;/td&gt;
&lt;td&gt;Post-training quantization (PTQ) or quantization-aware training (QAT)&lt;/td&gt;
&lt;td&gt;Iterative pruning during training or magnitude-based pruning after training, usually with fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware/runtime requirement&lt;/td&gt;
&lt;td&gt;Needs low-precision kernel support (INT8 cores, TensorRT, XNNPACK)&lt;/td&gt;
&lt;td&gt;Unstructured pruning needs sparse-matrix kernels for real speedup; structured pruning runs on standard dense hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compression achieved&lt;/td&gt;
&lt;td&gt;Typically 2-4x size reduction (FP32→INT8); INT4 pushes further at higher accuracy risk&lt;/td&gt;
&lt;td&gt;Can reach 50-90% sparsity, but unstructured sparsity often doesn&amp;rsquo;t translate to real speedup without special hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy impact &amp;amp; recovery&lt;/td&gt;
&lt;td&gt;Small accuracy drop, usually recovered via calibration or QAT&lt;/td&gt;
&lt;td&gt;Larger accuracy drop at high sparsity, recovered via iterative fine-tuning/retraining&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combinability&lt;/td&gt;
&lt;td&gt;Often applied last, to shrink an already-pruned model further&lt;/td&gt;
&lt;td&gt;Often applied first, before the pruned model is quantized&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Quantization changes each value&amp;rsquo;s &lt;strong class="kw"&gt;bit-width&lt;/strong&gt;; pruning changes the model&amp;rsquo;s &lt;strong class="kw"&gt;parameter count&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Realizing pruning&amp;rsquo;s theoretical speedup often requires &lt;strong class="kw"&gt;sparse kernels&lt;/strong&gt;, while quantization&amp;rsquo;s speedup comes from standard &lt;strong class="kw"&gt;INT8 hardware&lt;/strong&gt; support.&lt;/li&gt;
&lt;li&gt;Quantization degrades accuracy gradually and predictably; aggressive pruning risks a &lt;strong class="kw"&gt;sharp accuracy cliff&lt;/strong&gt; without fine-tuning.&lt;/li&gt;
&lt;li&gt;The two are commonly chained into a single &lt;strong class="kw"&gt;compression pipeline&lt;/strong&gt;, pruning first and quantizing the result.&lt;/li&gt;
&lt;li&gt;Structured pruning changes the model&amp;rsquo;s &lt;strong class="kw"&gt;architecture shape&lt;/strong&gt;; quantization never touches the architecture.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Model Quantization&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Zero-Shot Learning vs Few-Shot Learning: No Examples vs a Handful of Examples</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</link><pubDate>Mon, 03 Aug 2026 03:45:54 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Zero-shot and few-shot learning describe how much task-specific example data a model is given before it has to perform a task. Zero-shot relies solely on a &lt;strong class="kw"&gt;task description&lt;/strong&gt;, while few-shot conditions its predictions on a small set of &lt;strong class="kw"&gt;labeled examples&lt;/strong&gt;, usually trading a little setup cost for higher accuracy.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Zero-Shot&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Few-Shot&lt;/text&gt;&lt;line x1="320" y1="45" x2="320" y2="335" stroke-dasharray="4,4" style="stroke:var(--border)" stroke-width="1.5"/&gt;&lt;rect x="60" y="55" width="200" height="95" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="85" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="160" y="103" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;only&lt;/text&gt;&lt;text x="160" y="128" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;(0 examples)&lt;/text&gt;&lt;rect x="380" y="55" width="200" height="95" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="78" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="480" y="96" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;+ K examples&lt;/text&gt;&lt;rect x="400" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="412" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;1&lt;/text&gt;&lt;rect x="430" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="442" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;2&lt;/text&gt;&lt;rect x="460" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="472" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;3&lt;/text&gt;&lt;line x1="160" y1="150" x2="160" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,185 155,177 165,177" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="150" x2="480" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,185 475,177 485,177" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="160" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="160" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;rect x="380" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="480" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="480" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="160" y1="235" x2="160" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,270 155,262 165,262" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="235" x2="480" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,270 475,262 485,262" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="270" width="200" height="55" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;rect x="380" y="270" width="200" height="55" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;text x="160" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;no task-specific data&lt;/text&gt;&lt;text x="480" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;learns from few examples&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Zero-Shot Learning&lt;/th&gt;
&lt;th&gt;Few-Shot Learning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core definition&lt;/td&gt;
&lt;td&gt;Model performs a task it was never explicitly shown examples for, guided only by natural-language instructions or class descriptions&lt;/td&gt;
&lt;td&gt;Model performs a task after being shown a small number (typically 1-100) of labeled examples at inference or fine-tuning time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examples provided at inference&lt;/td&gt;
&lt;td&gt;None — only a task description or prompt&lt;/td&gt;
&lt;td&gt;A handful of input-output pairs included in the prompt or used for fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Underlying mechanism&lt;/td&gt;
&lt;td&gt;Relies entirely on knowledge encoded during pretraining plus semantic alignment between labels and text&lt;/td&gt;
&lt;td&gt;Uses in-context learning or lightweight fine-tuning to infer the task pattern directly from the provided examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Labeling/data cost&lt;/td&gt;
&lt;td&gt;Effectively zero — no labeled data needed for the target task&lt;/td&gt;
&lt;td&gt;Low but nonzero — requires curating a small, representative set of examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt/context length&lt;/td&gt;
&lt;td&gt;Short — just the instruction or class names&lt;/td&gt;
&lt;td&gt;Longer — instruction plus example pairs, consuming more context tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical accuracy&lt;/td&gt;
&lt;td&gt;Lower and more variable, especially on niche or ambiguous tasks&lt;/td&gt;
&lt;td&gt;Generally higher and more stable since examples disambiguate intent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sensitivity to example choice&lt;/td&gt;
&lt;td&gt;Not applicable — there are no examples to choose&lt;/td&gt;
&lt;td&gt;High — accuracy can swing significantly with example selection, order, and count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common techniques&lt;/td&gt;
&lt;td&gt;Prompt engineering, CLIP-style embedding matching, instruction-tuned LLMs&lt;/td&gt;
&lt;td&gt;Few-shot prompting, meta-learning (e.g. MAML), lightweight fine-tuning or LoRA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Zero-shot uses no &lt;strong class="kw"&gt;task examples&lt;/strong&gt; at all, relying purely on pretrained knowledge and instructions&lt;/li&gt;
&lt;li&gt;Few-shot conditions the model on a small &lt;strong class="kw"&gt;support set&lt;/strong&gt; of labeled examples at inference time&lt;/li&gt;
&lt;li&gt;Few-shot generally achieves higher &lt;strong class="kw"&gt;accuracy&lt;/strong&gt; because examples disambiguate an otherwise vague instruction&lt;/li&gt;
&lt;li&gt;Zero-shot has zero &lt;strong class="kw"&gt;labeling cost&lt;/strong&gt;, while few-shot requires curating representative examples&lt;/li&gt;
&lt;li&gt;Few-shot performance is sensitive to &lt;strong class="kw"&gt;example selection&lt;/strong&gt;, a variable that zero-shot simply doesn&amp;rsquo;t have&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Zero-Shot Learning&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Fine-Tuning vs RAG: Updating Model Weights vs Retrieving External Knowledge</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-fine-tuning-vs-rag-updating-model-weights-vs-retrieving-exte/</link><pubDate>Mon, 03 Aug 2026 03:44:47 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-fine-tuning-vs-rag-updating-model-weights-vs-retrieving-exte/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Fine-tuning and RAG (Retrieval-Augmented Generation) are two ways to make a large language model produce better, more relevant answers, but they intervene at different points in the pipeline. Fine-tuning permanently adjusts the model&amp;rsquo;s &lt;strong class="kw"&gt;weights&lt;/strong&gt; through additional training, while RAG leaves the model untouched and instead injects context by &lt;strong class="kw"&gt;retrieving&lt;/strong&gt; documents at query time. The choice matters because it determines how you update knowledge, control latency and cost, and trace where an answer came from.&lt;/p&gt;</description></item><item><title>Generative Model vs Discriminative Model: Modeling the Data vs Modeling the Boundary</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-generative-model-vs-discriminative-model-modeling-the-data-v/</link><pubDate>Mon, 03 Aug 2026 03:42:57 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-generative-model-vs-discriminative-model-modeling-the-data-v/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Generative and discriminative models represent two different answers to &amp;ldquo;what should a model actually learn from labeled data?&amp;rdquo; A &lt;strong class="kw"&gt;generative model&lt;/strong&gt; learns the full joint distribution of inputs and labels — effectively how each class produces its data — while a &lt;strong class="kw"&gt;discriminative model&lt;/strong&gt; learns only the boundary needed to tell classes apart, without modeling how the data itself was produced. That difference drives everything from data efficiency to whether the model can create new examples.&lt;/p&gt;</description></item><item><title>BERT vs GPT: Bidirectional Understanding vs Autoregressive Generation</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-bert-vs-gpt-bidirectional-understanding-vs-autoregressive-ge/</link><pubDate>Mon, 03 Aug 2026 03:42:01 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-bert-vs-gpt-bidirectional-understanding-vs-autoregressive-ge/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;BERT and GPT are both transformer-based language models, but they&amp;rsquo;re built from opposite halves of the transformer and trained for opposite jobs. BERT uses an &lt;strong class="kw"&gt;encoder&lt;/strong&gt; trained to fill in masked words using context from both directions, making it suited to understanding text, while GPT uses a &lt;strong class="kw"&gt;decoder&lt;/strong&gt; trained to predict the next word from only what came before, making it suited to generating text.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;
&lt;defs&gt;
&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;
&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-a)"/&gt;
&lt;/marker&gt;
&lt;marker id="arrowB" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;
&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-b)"/&gt;
&lt;/marker&gt;
&lt;/defs&gt;
&lt;line x1="320" y1="20" x2="320" y2="300" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4 4"/&gt;
&lt;text x="160" y="36" text-anchor="middle" style="fill:var(--primary)" font-size="22" font-weight="bold"&gt;BERT&lt;/text&gt;
&lt;text x="480" y="36" text-anchor="middle" style="fill:var(--primary)" font-size="22" font-weight="bold"&gt;GPT&lt;/text&gt;
&lt;text x="160" y="56" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Bidirectional Encoder&lt;/text&gt;
&lt;text x="480" y="56" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Autoregressive Decoder&lt;/text&gt;
&lt;path d="M50,153 Q100,85 150,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M100,153 Q125,120 150,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M150,153 Q175,120 200,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M150,153 Q200,85 250,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;rect x="30" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="50" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;the&lt;/text&gt;
&lt;rect x="80" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="100" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;cat&lt;/text&gt;
&lt;rect x="130" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="2" stroke-dasharray="3 2"/&gt;
&lt;text x="150" y="173" text-anchor="middle" style="fill:var(--primary)" font-size="9" font-weight="bold"&gt;[MASK]&lt;/text&gt;
&lt;rect x="180" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="200" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;on&lt;/text&gt;
&lt;rect x="230" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="250" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;mat&lt;/text&gt;
&lt;text x="160" y="222" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;sees full sentence context&lt;/text&gt;
&lt;text x="160" y="237" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;(left + right) to fill the mask&lt;/text&gt;
&lt;text x="160" y="270" text-anchor="middle" style="fill:var(--content)" font-size="12" font-weight="bold"&gt;-&amp;gt; classification, embeddings, NER&lt;/text&gt;
&lt;path d="M370,153 Q460,90 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M415,153 Q482,110 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M460,153 Q505,125 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M505,153 Q527,135 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M570,170 L574,170" style="stroke:var(--compare-b);fill:none" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;
&lt;rect x="350" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="370" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;the&lt;/text&gt;
&lt;rect x="395" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="415" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;cat&lt;/text&gt;
&lt;rect x="440" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="460" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;sat&lt;/text&gt;
&lt;rect x="485" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="505" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;on&lt;/text&gt;
&lt;rect x="530" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="2"/&gt;
&lt;text x="550" y="174" text-anchor="middle" style="fill:var(--primary)" font-size="11" font-weight="bold"&gt;mat&lt;/text&gt;
&lt;rect x="575" y="155" width="34" height="30" rx="4" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3 2"/&gt;
&lt;text x="592" y="174" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;?&lt;/text&gt;
&lt;text x="480" y="222" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;each token sees only itself&lt;/text&gt;
&lt;text x="480" y="237" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;+ prior tokens (causal mask)&lt;/text&gt;
&lt;text x="480" y="270" text-anchor="middle" style="fill:var(--content)" font-size="12" font-weight="bold"&gt;-&amp;gt; generation, chat, completion&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;BERT&lt;/th&gt;
&lt;th&gt;GPT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;Encoder-only transformer stack&lt;/td&gt;
&lt;td&gt;Decoder-only transformer stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pretraining objective&lt;/td&gt;
&lt;td&gt;Masked language modeling: predict randomly hidden tokens, plus next-sentence prediction&lt;/td&gt;
&lt;td&gt;Causal language modeling: predict the next token given all prior tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention pattern&lt;/td&gt;
&lt;td&gt;Bidirectional self-attention; every token attends to the full sequence&lt;/td&gt;
&lt;td&gt;Causal (masked) self-attention; each token attends only to itself and earlier tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output generation&lt;/td&gt;
&lt;td&gt;One contextual embedding per input token, produced in a single forward pass&lt;/td&gt;
&lt;td&gt;Text generated autoregressively, one token at a time, each output fed back as input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical adaptation&lt;/td&gt;
&lt;td&gt;Fine-tuned with a task-specific head on top of the pretrained encoder&lt;/td&gt;
&lt;td&gt;Adapted via prompting, instruction tuning, or fine-tuning to continue text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary use cases&lt;/td&gt;
&lt;td&gt;Classification, named entity recognition, semantic search, sentence embeddings&lt;/td&gt;
&lt;td&gt;Open-ended generation, chat, code completion, summarization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference cost per query&lt;/td&gt;
&lt;td&gt;Fixed: one pass regardless of desired output&lt;/td&gt;
&lt;td&gt;Scales with number of generated tokens, each requiring a forward pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;BERT&amp;rsquo;s &lt;strong class="kw"&gt;encoder&lt;/strong&gt; attends to both left and right context; GPT&amp;rsquo;s &lt;strong class="kw"&gt;decoder&lt;/strong&gt; attends only to prior tokens.&lt;/li&gt;
&lt;li&gt;BERT trains on &lt;strong class="kw"&gt;masked language modeling&lt;/strong&gt;; GPT trains on &lt;strong class="kw"&gt;next-token prediction&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;BERT produces embeddings in a &lt;strong class="kw"&gt;single pass&lt;/strong&gt;; GPT produces text through &lt;strong class="kw"&gt;autoregressive decoding&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;BERT is optimized for &lt;strong class="kw"&gt;understanding tasks&lt;/strong&gt;; GPT is optimized for &lt;strong class="kw"&gt;generation tasks&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;BERT&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Overfitting vs Underfitting: Memorizing Noise vs Missing the Signal</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-overfitting-vs-underfitting-memorizing-noise-vs-missing-the/</link><pubDate>Mon, 03 Aug 2026 03:37:23 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-overfitting-vs-underfitting-memorizing-noise-vs-missing-the/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Overfitting and underfitting describe the two ways a model can fail to generalize: one learns the training data too well, the other not well enough. Understanding which failure mode you&amp;rsquo;re in determines whether you should simplify or add regularization, or instead increase capacity and train longer. &lt;strong class="kw"&gt;Overfitting&lt;/strong&gt; traps a model in the noise of its training set, while &lt;strong class="kw"&gt;underfitting&lt;/strong&gt; leaves it unable to capture the underlying pattern at all.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="20" font-weight="700" style="fill:var(--primary)"&gt;Overfitting&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="20" font-weight="700" style="fill:var(--primary)"&gt;Underfitting&lt;/text&gt;&lt;line x1="40" y1="50" x2="40" y2="300" stroke-width="1.5" style="stroke:var(--border)"/&gt;&lt;line x1="40" y1="300" x2="280" y2="300" stroke-width="1.5" style="stroke:var(--border)"/&gt;&lt;line x1="360" y1="50" x2="360" y2="300" stroke-width="1.5" style="stroke:var(--border)"/&gt;&lt;line x1="360" y1="300" x2="600" y2="300" stroke-width="1.5" style="stroke:var(--border)"/&gt;&lt;path d="M60,260 L95,180 L130,220 L165,140 L200,190 L235,110 L262,150" fill="none" stroke-width="2.5" style="stroke:var(--compare-a)"/&gt;&lt;circle cx="60" cy="260" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="95" cy="180" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="130" cy="220" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="165" cy="140" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="200" cy="190" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="235" cy="110" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="262" cy="150" r="5" stroke-width="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)"/&gt;&lt;circle cx="360" cy="260" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="395" cy="180" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="430" cy="220" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="465" cy="140" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="500" cy="190" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="535" cy="110" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;circle cx="562" cy="150" r="5" stroke-width="2" style="fill:var(--compare-b-soft);stroke:var(--compare-b)"/&gt;&lt;line x1="345" y1="225" x2="575" y2="150" stroke-width="2.5" style="stroke:var(--compare-b)"/&gt;&lt;text x="160" y="325" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Fits every point exactly&lt;/text&gt;&lt;text x="160" y="345" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;captures noise, not signal&lt;/text&gt;&lt;text x="480" y="325" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Misses the curved trend&lt;/text&gt;&lt;text x="480" y="345" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;too simple for the pattern&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Overfitting&lt;/th&gt;
&lt;th&gt;Underfitting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Underlying cause&lt;/td&gt;
&lt;td&gt;Model too complex relative to the data, so it learns noise and idiosyncrasies&lt;/td&gt;
&lt;td&gt;Model too simple to represent the true relationship in the data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training error&lt;/td&gt;
&lt;td&gt;Very low, often near zero&lt;/td&gt;
&lt;td&gt;High, the model struggles even on data it was trained on&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validation/test error&lt;/td&gt;
&lt;td&gt;High, much worse than training error&lt;/td&gt;
&lt;td&gt;High, similar in magnitude to training error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bias-variance profile&lt;/td&gt;
&lt;td&gt;Low bias, high variance&lt;/td&gt;
&lt;td&gt;High bias, low variance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generalization to new data&lt;/td&gt;
&lt;td&gt;Poor, predictions swing wildly on unseen inputs&lt;/td&gt;
&lt;td&gt;Poor, predictions are consistently and systematically off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning curve signature&lt;/td&gt;
&lt;td&gt;Training and validation loss diverge as training continues&lt;/td&gt;
&lt;td&gt;Training and validation loss both plateau high and close together&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical remedies&lt;/td&gt;
&lt;td&gt;Regularization, more training data, dropout, early stopping, simpler model&lt;/td&gt;
&lt;td&gt;Increase model capacity, add features, train longer, reduce regularization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Overfitting memorizes &lt;strong class="kw"&gt;noise&lt;/strong&gt; in the training set, while underfitting never learns the underlying &lt;strong class="kw"&gt;pattern&lt;/strong&gt; at all.&lt;/li&gt;
&lt;li&gt;Overfitting shows near-zero training error but a wide train/validation gap, a sign of &lt;strong class="kw"&gt;high variance&lt;/strong&gt;; underfitting shows poor performance on both, a sign of &lt;strong class="kw"&gt;high bias&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Overfitting is treated with &lt;strong class="kw"&gt;regularization&lt;/strong&gt; or more data; underfitting is treated by increasing &lt;strong class="kw"&gt;model capacity&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Overfitting gets worse the longer an overly flexible model &lt;strong class="kw"&gt;keeps training&lt;/strong&gt;; underfitting persists regardless of duration, since it&amp;rsquo;s a &lt;strong class="kw"&gt;structural limit&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Overfitting&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Precision vs Recall: Predicted-Positive Accuracy vs Actual-Positive Coverage</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-precision-vs-recall-predicted-positive-accuracy-vs-actual-po/</link><pubDate>Mon, 03 Aug 2026 03:36:08 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-precision-vs-recall-predicted-positive-accuracy-vs-actual-po/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Precision and recall are two classification metrics computed from the same confusion matrix but answering different questions about a model&amp;rsquo;s positive predictions. &lt;strong class="kw"&gt;Precision&lt;/strong&gt; asks how many predicted positives were correct, while &lt;strong class="kw"&gt;recall&lt;/strong&gt; asks how many actual positives were found. Optimizing one in isolation almost always trades off against the other.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="180" y="40" text-anchor="middle" font-size="16" style="fill:var(--compare-a)"&gt;Predicted Positive&lt;/text&gt;&lt;text x="460" y="40" text-anchor="middle" font-size="16" style="fill:var(--compare-b)"&gt;Actual Positive&lt;/text&gt;&lt;circle cx="270" cy="165" r="95" fill-opacity="0.55" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="2"/&gt;&lt;circle cx="390" cy="165" r="95" fill-opacity="0.55" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;text x="215" y="170" text-anchor="middle" font-size="20" style="fill:var(--compare-a)"&gt;FP&lt;/text&gt;&lt;text x="330" y="170" text-anchor="middle" font-size="22" style="fill:var(--primary)"&gt;TP&lt;/text&gt;&lt;text x="445" y="170" text-anchor="middle" font-size="20" style="fill:var(--compare-b)"&gt;FN&lt;/text&gt;&lt;text x="215" y="195" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;wrong alarms&lt;/text&gt;&lt;text x="445" y="195" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;missed cases&lt;/text&gt;&lt;line x1="270" y1="270" x2="270" y2="300" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="270" y="320" text-anchor="middle" font-size="15" style="fill:var(--compare-a)"&gt;Precision = TP / (TP + FP)&lt;/text&gt;&lt;line x1="390" y1="270" x2="390" y2="300" style="stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="345" text-anchor="middle" font-size="15" style="fill:var(--compare-b)"&gt;Recall = TP / (TP + FN)&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;th&gt;Recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Question answered&lt;/td&gt;
&lt;td&gt;Of items predicted positive, how many actually are positive?&lt;/td&gt;
&lt;td&gt;Of items that are actually positive, how many did the model find?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Formula&lt;/td&gt;
&lt;td&gt;TP / (TP + FP)&lt;/td&gt;
&lt;td&gt;TP / (TP + FN)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Denominator basis&lt;/td&gt;
&lt;td&gt;Total predicted positive (TP + FP)&lt;/td&gt;
&lt;td&gt;Total actual positive (TP + FN)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error type penalized&lt;/td&gt;
&lt;td&gt;False positives (false alarms)&lt;/td&gt;
&lt;td&gt;False negatives (missed detections)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Increases when&lt;/td&gt;
&lt;td&gt;Model makes fewer incorrect positive calls&lt;/td&gt;
&lt;td&gt;Model catches more of the true positive cases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Threshold trade-off&lt;/td&gt;
&lt;td&gt;Raising the decision threshold typically raises precision&lt;/td&gt;
&lt;td&gt;Lowering the decision threshold typically raises recall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure mode at extreme&lt;/td&gt;
&lt;td&gt;High precision, low recall: model is overly conservative and misses real cases&lt;/td&gt;
&lt;td&gt;High recall, low precision: model is overly liberal and floods results with false alarms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Precision&amp;rsquo;s denominator is predicted positives; recall&amp;rsquo;s denominator is actual positives, so they measure against different totals&lt;/li&gt;
&lt;li&gt;Precision is hurt by &lt;strong class="kw"&gt;false positives&lt;/strong&gt;; recall is hurt by &lt;strong class="kw"&gt;false negatives&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Adjusting the classification &lt;strong class="kw"&gt;threshold&lt;/strong&gt; pushes precision and recall in opposite directions&lt;/li&gt;
&lt;li&gt;Neither metric alone summarizes model quality, which is why the &lt;strong class="kw"&gt;F1 score&lt;/strong&gt; combines them&lt;/li&gt;
&lt;li&gt;A model with 100% recall can trivially predict everyone positive, and a model with 100% precision can trivially predict almost no one positive&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Precision&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Bagging vs Boosting: Parallel Resampling vs Sequential Error Correction</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-bagging-vs-boosting-parallel-resampling-vs-sequential-error/</link><pubDate>Mon, 03 Aug 2026 03:35:31 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-bagging-vs-boosting-parallel-resampling-vs-sequential-error/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Bagging and Boosting are both ensemble techniques that combine many weak learners into one stronger model, but they build that ensemble in fundamentally different ways. Bagging trains learners independently in parallel on &lt;strong class="kw"&gt;bootstrap samples&lt;/strong&gt; and averages their outputs to cut variance, while Boosting trains learners one after another, each one correcting the last model&amp;rsquo;s mistakes through &lt;strong class="kw"&gt;sequential reweighting&lt;/strong&gt; to cut bias. The choice affects training time, overfitting risk, and how robust the model is to noisy data.&lt;/p&gt;</description></item><item><title>Classification vs Regression: Predicting Categories vs Predicting Numbers</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-classification-vs-regression-predicting-categories-vs-predic/</link><pubDate>Mon, 03 Aug 2026 03:28:23 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-classification-vs-regression-predicting-categories-vs-predic/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Classification and regression are the two core types of supervised learning, distinguished by what kind of output they predict. Classification assigns inputs to a &lt;strong class="kw"&gt;discrete class&lt;/strong&gt;, while regression estimates a &lt;strong class="kw"&gt;continuous value&lt;/strong&gt;. Picking the wrong one for your target variable leads to mismatched loss functions, evaluation metrics, and model outputs.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;line x1="320" y1="20" x2="320" y2="340" stroke-dasharray="4 4" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;text x="160" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="20" font-weight="bold"&gt;Classification&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="20" font-weight="bold"&gt;Regression&lt;/text&gt;&lt;line x1="60" y1="280" x2="280" y2="280" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="60" y1="80" x2="60" y2="280" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="150" y1="85" x2="270" y2="270" stroke-dasharray="5 3" style="stroke:var(--border)" stroke-width="1.5"/&gt;&lt;circle cx="90" cy="150" r="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="110" cy="130" r="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="95" cy="175" r="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="125" cy="160" r="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="80" cy="200" r="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="214" y="214" width="12" height="12" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="234" y="194" width="12" height="12" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="249" y="229" width="12" height="12" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="224" y="245" width="12" height="12" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="199" y="204" width="12" height="12" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="70" cy="305" r="5" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="82" y="309" style="fill:var(--content)" font-size="11"&gt;Class A&lt;/text&gt;&lt;rect x="150" y="300" width="10" height="10" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="165" y="309" style="fill:var(--content)" font-size="11"&gt;Class B&lt;/text&gt;&lt;text x="160" y="335" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;output: discrete category&lt;/text&gt;&lt;line x1="380" y1="280" x2="600" y2="280" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="380" y1="80" x2="380" y2="280" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;circle cx="400" cy="245" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="420" cy="225" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="440" cy="230" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="205" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="480" cy="195" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="500" cy="175" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="520" cy="165" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="540" cy="145" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="560" cy="135" r="5" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;line x1="392" y1="255" x2="588" y2="120" style="stroke:var(--compare-b)" stroke-width="2"/&gt;&lt;text x="480" y="335" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;output: continuous number&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Classification&lt;/th&gt;
&lt;th&gt;Regression&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Target variable type&lt;/td&gt;
&lt;td&gt;Discrete, categorical labels from a finite set of classes&lt;/td&gt;
&lt;td&gt;Continuous, ordered numeric values&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning objective&lt;/td&gt;
&lt;td&gt;Learn a decision boundary that separates classes&lt;/td&gt;
&lt;td&gt;Learn a function mapping inputs to a continuous output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical loss function&lt;/td&gt;
&lt;td&gt;Cross-entropy, log loss, or hinge loss&lt;/td&gt;
&lt;td&gt;Mean squared error or mean absolute error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model output format&lt;/td&gt;
&lt;td&gt;Class label or probability distribution over classes&lt;/td&gt;
&lt;td&gt;Single scalar value (or vector of scalars)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common algorithms&lt;/td&gt;
&lt;td&gt;Logistic regression, SVM, decision trees, kNN, softmax networks&lt;/td&gt;
&lt;td&gt;Linear regression, ridge/lasso, decision trees, kNN, regression networks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation metrics&lt;/td&gt;
&lt;td&gt;Accuracy, precision/recall, F1, ROC-AUC, confusion matrix&lt;/td&gt;
&lt;td&gt;RMSE, MAE, R-squared, MAPE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error interpretation&lt;/td&gt;
&lt;td&gt;Prediction is simply right, wrong, or confused with another class&lt;/td&gt;
&lt;td&gt;Prediction error has magnitude and direction, showing how far off it was&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Classification predicts a &lt;strong class="kw"&gt;discrete label&lt;/strong&gt; from a fixed set of classes, while regression predicts a &lt;strong class="kw"&gt;continuous value&lt;/strong&gt; on a numeric scale.&lt;/li&gt;
&lt;li&gt;Classification models typically optimize &lt;strong class="kw"&gt;cross-entropy loss&lt;/strong&gt; to separate classes, while regression models optimize &lt;strong class="kw"&gt;squared error&lt;/strong&gt; to minimize distance from the true value.&lt;/li&gt;
&lt;li&gt;Classification is evaluated with metrics like &lt;strong class="kw"&gt;accuracy/F1&lt;/strong&gt;, while regression is evaluated with metrics like &lt;strong class="kw"&gt;RMSE/R-squared&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;A classification error is simply right, wrong, or a class confusion, while a regression error carries a &lt;strong class="kw"&gt;magnitude&lt;/strong&gt; showing how far off the prediction was.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Classification&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Supervised vs Unsupervised Learning: Labeled Guidance vs Pattern Discovery</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-supervised-vs-unsupervised-learning-labeled-guidance-vs-patt/</link><pubDate>Mon, 03 Aug 2026 03:27:11 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-supervised-vs-unsupervised-learning-labeled-guidance-vs-patt/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Supervised learning trains a model on &lt;strong class="kw"&gt;labeled data&lt;/strong&gt;, teaching it to map inputs to known outputs it can later predict. Unsupervised learning works on raw, unlabeled data and instead performs &lt;strong class="kw"&gt;pattern discovery&lt;/strong&gt;, uncovering structure like clusters or reduced representations with no target to match against. The distinction matters because it determines what data you need, how you measure success, and which problems each approach can actually solve.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="5" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" viewBox="0 0 10 10" refX="5" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="15" x2="320" y2="345" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" font-size="17" font-weight="700" style="fill:var(--primary)"&gt;Supervised Learning&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" font-size="17" font-weight="700" style="fill:var(--primary)"&gt;Unsupervised Learning&lt;/text&gt;&lt;g&gt;&lt;rect x="45" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="73" y1="57" x2="88" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="98" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="98" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;rect x="115" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="143" y1="57" x2="158" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="168" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="168" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;rect x="185" y="48" width="28" height="18" rx="2" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="213" y1="57" x2="228" y2="57" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="238" cy="57" r="8" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="238" y="60" text-anchor="middle" font-size="9" style="fill:var(--content)"&gt;y&lt;/text&gt;&lt;/g&gt;&lt;text x="160" y="82" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;features + known labels&lt;/text&gt;&lt;line x1="160" y1="92" x2="160" y2="133" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;rect x="90" y="138" width="140" height="42" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="164" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="160" y1="180" x2="160" y2="205" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="222" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Predicted label&lt;/text&gt;&lt;path d="M148,246 L157,255 L175,230" style="stroke:var(--compare-a);fill:none" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round"/&gt;&lt;text x="160" y="275" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;checked against true y&lt;/text&gt;&lt;g&gt;&lt;circle cx="400" cy="55" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="430" cy="75" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="50" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="500" cy="70" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="530" cy="52" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="558" cy="78" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;/g&gt;&lt;text x="480" y="98" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;features only, no labels&lt;/text&gt;&lt;line x1="480" y1="108" x2="480" y2="133" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;rect x="410" y="138" width="140" height="42" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="164" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="480" y1="180" x2="480" y2="205" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;circle cx="425" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="417" cy="224" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="433" cy="236" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="480" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="472" cy="222" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="488" cy="238" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="535" cy="230" r="26" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3,3"/&gt;&lt;circle cx="527" cy="224" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="543" cy="236" r="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="285" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Discovered clusters&lt;/text&gt;&lt;text x="480" y="303" text-anchor="middle" font-size="10.5" style="fill:var(--secondary)"&gt;no ground truth to check&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Supervised Learning&lt;/th&gt;
&lt;th&gt;Unsupervised Learning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input data&lt;/td&gt;
&lt;td&gt;Labeled examples: features paired with a known target value&lt;/td&gt;
&lt;td&gt;Unlabeled examples: features only, no target provided&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning objective&lt;/td&gt;
&lt;td&gt;Minimize the error between predicted and true labels&lt;/td&gt;
&lt;td&gt;Discover inherent structure, grouping, or compressed representation in the data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training signal&lt;/td&gt;
&lt;td&gt;Explicit feedback from a loss function computed against ground truth&lt;/td&gt;
&lt;td&gt;No explicit feedback; relies on similarity, density, or variance within the data itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model output&lt;/td&gt;
&lt;td&gt;A predicted class label or continuous value&lt;/td&gt;
&lt;td&gt;Cluster assignments, reduced dimensions, or anomaly scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation&lt;/td&gt;
&lt;td&gt;Direct measurement on a held-out labeled test set (accuracy, F1, RMSE)&lt;/td&gt;
&lt;td&gt;Indirect measurement (silhouette score, reconstruction error) or human interpretation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common tasks&lt;/td&gt;
&lt;td&gt;Classification and regression&lt;/td&gt;
&lt;td&gt;Clustering, dimensionality reduction, and anomaly detection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data/labeling cost&lt;/td&gt;
&lt;td&gt;Requires a labeled dataset, often costly and time-consuming to build&lt;/td&gt;
&lt;td&gt;Uses raw data as-is, cheaper and faster to collect at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical algorithms&lt;/td&gt;
&lt;td&gt;Logistic regression, random forests, gradient boosting, supervised neural nets&lt;/td&gt;
&lt;td&gt;k-means, PCA, DBSCAN, autoencoders&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Supervised learning requires &lt;strong class="kw"&gt;labeled data&lt;/strong&gt;; unsupervised learning works directly on &lt;strong class="kw"&gt;raw data&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Supervised models are scored against &lt;strong class="kw"&gt;ground truth&lt;/strong&gt;; unsupervised models are judged by &lt;strong class="kw"&gt;internal structure&lt;/strong&gt; metrics instead.&lt;/li&gt;
&lt;li&gt;Supervised learning targets &lt;strong class="kw"&gt;prediction&lt;/strong&gt; of a known outcome; unsupervised learning targets &lt;strong class="kw"&gt;discovery&lt;/strong&gt; of unknown structure.&lt;/li&gt;
&lt;li&gt;Labeling is usually the &lt;strong class="kw"&gt;bottleneck cost&lt;/strong&gt; for supervised systems, while unsupervised systems scale with raw &lt;strong class="kw"&gt;data volume&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Supervised errors are measurable per-example; unsupervised quality is often assessed via &lt;strong class="kw"&gt;proxy metrics&lt;/strong&gt; or manual review.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Supervised Learning&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Machine Learning vs Deep Learning: Manual Features vs Learned Representations</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-machine-learning-vs-deep-learning-manual-features-vs-learned/</link><pubDate>Mon, 03 Aug 2026 03:25:19 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-machine-learning-vs-deep-learning-manual-features-vs-learned/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Deep Learning is technically a subset of Machine Learning, but in practice the two names are used to distinguish classical algorithms from neural-network-based approaches. Traditional &lt;strong class="kw"&gt;Machine Learning&lt;/strong&gt; relies on humans to hand-engineer features before a model like a decision tree or SVM can learn from them, while &lt;strong class="kw"&gt;Deep Learning&lt;/strong&gt; uses multi-layer neural networks that learn their own feature representations directly from raw data. The distinction matters because it drives very different requirements for data volume, compute, and interpretability.&lt;/p&gt;</description></item></channel></rss>