<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Neural-Networks on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/neural-networks/</link><description>Recent content in Neural-Networks on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:40:07 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/neural-networks/index.xml" rel="self" type="application/rss+xml"/><item><title>Transformer vs RNN: Parallel Attention vs Sequential Recurrence</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</link><pubDate>Mon, 03 Aug 2026 03:40:07 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;RNNs process sequences one token at a time, carrying context forward through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt; that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;. The difference reshapes everything from training speed to how well long-range context survives.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="20" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="5,5"/&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;RNN&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;Transformer&lt;/text&gt;&lt;circle cx="70" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="140" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="210" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="280" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="140" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="210" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="280" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="45" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="115" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="185" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="255" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h1&lt;/text&gt;&lt;text x="140" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h2&lt;/text&gt;&lt;text x="210" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h3&lt;/text&gt;&lt;text x="280" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h4&lt;/text&gt;&lt;line x1="70" y1="286" x2="70" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="140" y1="286" x2="140" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="210" y1="286" x2="210" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="280" y1="286" x2="280" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="95" y1="180" x2="113" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="165" y1="180" x2="183" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="235" y1="180" x2="253" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;state passed step by step&lt;/text&gt;&lt;text x="160" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes one token at a time&lt;/text&gt;&lt;circle cx="390" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="530" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="600" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="460" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="530" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="600" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="365" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="435" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="505" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="575" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z1&lt;/text&gt;&lt;text x="460" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z2&lt;/text&gt;&lt;text x="530" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z3&lt;/text&gt;&lt;text x="600" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z4&lt;/text&gt;&lt;line x1="390" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="390" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="460" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="530" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;text x="495" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;every token attends to every token&lt;/text&gt;&lt;text x="495" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes all tokens at once&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;RNN&lt;/th&gt;
&lt;th&gt;Transformer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input processing order&lt;/td&gt;
&lt;td&gt;Tokens consumed one at a time, in sequence&lt;/td&gt;
&lt;td&gt;All tokens consumed simultaneously&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context propagation&lt;/td&gt;
&lt;td&gt;Hidden state carried forward step to step&lt;/td&gt;
&lt;td&gt;Self-attention lets each position read all others directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-range dependencies&lt;/td&gt;
&lt;td&gt;Signal weakens over distance (vanishing/exploding gradients)&lt;/td&gt;
&lt;td&gt;Direct connection between any two positions regardless of distance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Positional information&lt;/td&gt;
&lt;td&gt;Implicit, from the order tokens are fed in&lt;/td&gt;
&lt;td&gt;Explicit, via added positional encodings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training parallelization&lt;/td&gt;
&lt;td&gt;Limited — must unroll and step through time&lt;/td&gt;
&lt;td&gt;Fully parallel across the sequence dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computational cost&lt;/td&gt;
&lt;td&gt;O(n) sequential steps, O(1) state per step&lt;/td&gt;
&lt;td&gt;O(n^2) attention cost over sequence length&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference/generation&lt;/td&gt;
&lt;td&gt;Constant memory, naturally one step at a time&lt;/td&gt;
&lt;td&gt;Requires KV caching to avoid recomputing past attention&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use cases&lt;/td&gt;
&lt;td&gt;LSTM/GRU for streaming or small-scale sequence tasks&lt;/td&gt;
&lt;td&gt;BERT/GPT-style models for large-scale language and vision tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Transformer processes all tokens in parallel via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;; RNN processes tokens sequentially through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;RNN suffers from &lt;strong class="kw"&gt;vanishing gradients&lt;/strong&gt; over long sequences; Transformer links any two positions directly&lt;/li&gt;
&lt;li&gt;Transformer needs explicit &lt;strong class="kw"&gt;positional encodings&lt;/strong&gt; since attention has no inherent order; RNN gets order for free&lt;/li&gt;
&lt;li&gt;Transformer training scales with &lt;strong class="kw"&gt;quadratic complexity&lt;/strong&gt; in sequence length; RNN training is linear but hard to parallelize&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;RNN&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Batch Normalization vs Dropout: Stabilizing Activations vs Preventing Co-Adaptation</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-batch-normalization-vs-dropout-stabilizing-activations-vs-pr/</link><pubDate>Mon, 03 Aug 2026 03:38:48 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-batch-normalization-vs-dropout-stabilizing-activations-vs-pr/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Batch normalization and dropout are both inserted between layers of a neural network, but they solve different problems during training. Batch normalization &lt;strong class="kw"&gt;rescales activations&lt;/strong&gt; using batch statistics to stabilize and speed up training, while dropout &lt;strong class="kw"&gt;randomly zeroes neurons&lt;/strong&gt; to stop the network from over-relying on any single feature.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;
&lt;line x1="320" y1="50" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;
&lt;text x="170" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Batch Normalization&lt;/text&gt;
&lt;text x="470" y="30" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Dropout&lt;/text&gt;
&lt;text x="170" y="60" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;raw activations (batch)&lt;/text&gt;
&lt;rect x="45" y="110" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="95" y="80" width="30" height="70" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="145" y="125" width="30" height="25" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="195" y="95" width="30" height="55" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="245" y="115" width="30" height="35" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;line x1="150" y1="150" x2="325" y2="150" style="stroke:var(--border)" stroke-width="1"/&gt;
&lt;line x1="170" y1="160" x2="170" y2="190" style="stroke:var(--content)" stroke-width="1.5"/&gt;
&lt;polygon points="170,196 165,186 175,186" style="fill:var(--content)"/&gt;
&lt;text x="200" y="180" text-anchor="middle" style="fill:var(--secondary)" font-size="10"&gt;μ=0, σ=1&lt;/text&gt;
&lt;rect x="45" y="220" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="95" y="220" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="145" y="220" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="195" y="220" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;rect x="245" y="220" width="30" height="40" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="170" y="280" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;normalized, uniform scale&lt;/text&gt;
&lt;text x="170" y="330" text-anchor="middle" style="fill:var(--secondary)" font-size="10"&gt;recomputed from batch stats every pass&lt;/text&gt;
&lt;text x="480" y="80" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;pass 1 - random mask&lt;/text&gt;
&lt;circle cx="380" cy="115" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="430" cy="115" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="480" cy="115" r="18" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,3"/&gt;
&lt;circle cx="530" cy="115" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="580" cy="115" r="18" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,3"/&gt;
&lt;text x="480" y="195" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;pass 2 - different mask&lt;/text&gt;
&lt;circle cx="380" cy="230" r="18" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,3"/&gt;
&lt;circle cx="430" cy="230" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="480" cy="230" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="530" cy="230" r="18" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,3"/&gt;
&lt;circle cx="580" cy="230" r="18" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="480" y="330" text-anchor="middle" style="fill:var(--secondary)" font-size="10"&gt;random subset silenced each forward pass&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Batch Normalization&lt;/th&gt;
&lt;th&gt;Dropout&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Insertion point&lt;/td&gt;
&lt;td&gt;after a linear/conv layer, before the activation function&lt;/td&gt;
&lt;td&gt;after the activation function, on the layer&amp;rsquo;s output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Core mechanism&lt;/td&gt;
&lt;td&gt;normalizes activations to zero mean/unit variance using batch statistics, then applies a learnable scale and shift&lt;/td&gt;
&lt;td&gt;randomly zeroes a fraction p of activations on each forward pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary goal&lt;/td&gt;
&lt;td&gt;stabilize and accelerate training by reducing internal covariate shift&lt;/td&gt;
&lt;td&gt;reduce overfitting by preventing neurons from co-adapting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learnable/hyperparameters&lt;/td&gt;
&lt;td&gt;learnable scale (gamma) and shift (beta) per channel; tracks running mean/variance&lt;/td&gt;
&lt;td&gt;no learnable parameters; single hyperparameter, the dropout rate p&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train vs inference behavior&lt;/td&gt;
&lt;td&gt;uses batch statistics in training, switches to stored running averages at inference&lt;/td&gt;
&lt;td&gt;active during training, disabled entirely (identity function) at inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sensitivity to batch size&lt;/td&gt;
&lt;td&gt;degrades with very small or inconsistent batches since statistics become noisy&lt;/td&gt;
&lt;td&gt;unaffected by batch size, operates independently per example&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combined usage&lt;/td&gt;
&lt;td&gt;typically placed before dropout; can conflict with dropout&amp;rsquo;s variance shift&lt;/td&gt;
&lt;td&gt;often reduced or omitted alongside batch norm in modern CNNs due to that interaction&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Batch normalization rescales using &lt;strong class="kw"&gt;batch statistics&lt;/strong&gt;; dropout relies on &lt;strong class="kw"&gt;random masking&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Batch normalization introduces learnable &lt;strong class="kw"&gt;scale and shift&lt;/strong&gt; parameters; dropout adds none.&lt;/li&gt;
&lt;li&gt;At inference, batch normalization switches to &lt;strong class="kw"&gt;running averages&lt;/strong&gt; while dropout is simply &lt;strong class="kw"&gt;turned off&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Batch normalization mainly targets &lt;strong class="kw"&gt;training stability&lt;/strong&gt;; dropout mainly targets &lt;strong class="kw"&gt;overfitting&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Combining them naively can cause a &lt;strong class="kw"&gt;variance shift&lt;/strong&gt; that hurts performance.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Batch Normalization&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>LSTM vs GRU: Three Gates vs Two Gates in Recurrent Memory</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-lstm-vs-gru-three-gates-vs-two-gates-in-recurrent-memory/</link><pubDate>Mon, 03 Aug 2026 03:33:36 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-lstm-vs-gru-three-gates-vs-two-gates-in-recurrent-memory/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;LSTM and GRU are both gated recurrent architectures built to capture long-range dependencies in sequences while avoiding the vanishing-gradient problem of vanilla RNNs. LSTM keeps a dedicated &lt;strong class="kw"&gt;cell state&lt;/strong&gt; alongside its hidden state, regulated by three gates, while GRU folds everything into a single &lt;strong class="kw"&gt;hidden state&lt;/strong&gt; updated by just two gates. That structural difference drives everything else: parameter count, training speed, and how precisely you can control what the network remembers.&lt;/p&gt;</description></item><item><title>CNN vs RNN: Spatial Convolution vs Sequential Recurrence</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-cnn-vs-rnn-spatial-convolution-vs-sequential-recurrence/</link><pubDate>Mon, 03 Aug 2026 03:31:06 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-cnn-vs-rnn-spatial-convolution-vs-sequential-recurrence/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) are neural architectures built for different data shapes: CNNs slide &lt;strong class="kw"&gt;convolutional filters&lt;/strong&gt; across a spatial grid to detect local patterns, while RNNs pass a &lt;strong class="kw"&gt;recurrent hidden state&lt;/strong&gt; across time steps to model sequential dependencies. Choosing between them (or their modern successors) depends on whether your data&amp;rsquo;s structure is spatial, temporal, or both.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" markerWidth="8" markerHeight="8" refX="6" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" markerWidth="8" markerHeight="8" refX="6" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="65" x2="320" y2="310" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4,4"/&gt;&lt;text x="150" y="32" text-anchor="middle" style="fill:var(--primary)" font-size="20" font-weight="bold"&gt;CNN&lt;/text&gt;&lt;text x="150" y="52" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Convolution over a spatial grid&lt;/text&gt;&lt;rect x="52" y="90" width="104" height="104" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="78" y1="90" x2="78" y2="194" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="104" y1="90" x2="104" y2="194" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="130" y1="90" x2="130" y2="194" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="52" y1="116" x2="156" y2="116" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="52" y1="142" x2="156" y2="142" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="52" y1="168" x2="156" y2="168" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;rect x="52" y="90" width="52" height="52" style="fill:none;stroke:var(--compare-a)" stroke-width="3"/&gt;&lt;line x1="108" y1="116" x2="123" y2="116" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="160" y1="142" x2="191" y2="142" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="176" y="132" text-anchor="middle" style="fill:var(--secondary)" font-size="10"&gt;convolve&lt;/text&gt;&lt;rect x="196" y="112" width="54" height="54" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;line x1="214" y1="112" x2="214" y2="166" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="232" y1="112" x2="232" y2="166" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="196" y1="130" x2="250" y2="130" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;line x1="196" y1="148" x2="250" y2="148" style="stroke:var(--border)" stroke-width="1"/&gt;&lt;text x="150" y="215" text-anchor="middle" style="fill:var(--secondary)" font-size="10.5"&gt;Filter weights shared across all positions&lt;/text&gt;&lt;text x="150" y="229" text-anchor="middle" style="fill:var(--secondary)" font-size="10.5"&gt;Captures local spatial patterns&lt;/text&gt;&lt;text x="150" y="300" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Best for grid-structured data (images)&lt;/text&gt;&lt;text x="460" y="32" text-anchor="middle" style="fill:var(--primary)" font-size="20" font-weight="bold"&gt;RNN&lt;/text&gt;&lt;text x="460" y="52" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Recurrence over a sequence&lt;/text&gt;&lt;rect x="360" y="130" width="40" height="40" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="440" y="130" width="40" height="40" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="520" y="130" width="40" height="40" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="380" y="155" text-anchor="middle" style="fill:var(--content)" font-size="14"&gt;h1&lt;/text&gt;&lt;text x="460" y="155" text-anchor="middle" style="fill:var(--content)" font-size="14"&gt;h2&lt;/text&gt;&lt;text x="540" y="155" text-anchor="middle" style="fill:var(--content)" font-size="14"&gt;h3&lt;/text&gt;&lt;line x1="400" y1="150" x2="439" y2="150" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="480" y1="150" x2="519" y2="150" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="380" y1="214" x2="380" y2="171" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="460" y1="214" x2="460" y2="171" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="540" y1="214" x2="540" y2="171" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;text x="380" y="228" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;x1&lt;/text&gt;&lt;text x="460" y="228" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;x2&lt;/text&gt;&lt;text x="540" y="228" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;x3&lt;/text&gt;&lt;line x1="380" y1="129" x2="380" y2="87" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="460" y1="129" x2="460" y2="87" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;line x1="540" y1="129" x2="540" y2="87" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;text x="380" y="78" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;y1&lt;/text&gt;&lt;text x="460" y="78" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;y2&lt;/text&gt;&lt;text x="540" y="78" text-anchor="middle" style="fill:var(--content)" font-size="12"&gt;y3&lt;/text&gt;&lt;text x="460" y="246" text-anchor="middle" style="fill:var(--secondary)" font-size="10.5"&gt;Hidden state carries context&lt;/text&gt;&lt;text x="460" y="260" text-anchor="middle" style="fill:var(--secondary)" font-size="10.5"&gt;forward through the sequence&lt;/text&gt;&lt;text x="460" y="300" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;Best for sequential/time-ordered data&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;CNN&lt;/th&gt;
&lt;th&gt;RNN&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input data shape&lt;/td&gt;
&lt;td&gt;Fixed-size spatial grid (2D/3D tensors like images)&lt;/td&gt;
&lt;td&gt;Variable-length ordered sequence (text, time series, audio)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Core operation&lt;/td&gt;
&lt;td&gt;Convolution: a filter slides over local receptive fields&lt;/td&gt;
&lt;td&gt;Recurrence: hidden state updated step-by-step from previous state plus current input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight sharing&lt;/td&gt;
&lt;td&gt;Same filter weights reused across all spatial positions&lt;/td&gt;
&lt;td&gt;Same weight matrices reused across all time steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context captured&lt;/td&gt;
&lt;td&gt;Local spatial neighborhoods, expanded via depth/pooling&lt;/td&gt;
&lt;td&gt;Temporal history accumulated in the hidden state over prior steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Order sensitivity&lt;/td&gt;
&lt;td&gt;Largely order-invariant beyond local structure; pooling discards exact position&lt;/td&gt;
&lt;td&gt;Strictly order-dependent; reordering the sequence changes the output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training parallelization&lt;/td&gt;
&lt;td&gt;Highly parallelizable across positions, channels, and layers&lt;/td&gt;
&lt;td&gt;Inherently sequential; each step waits on the previous hidden state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common failure mode&lt;/td&gt;
&lt;td&gt;Limited receptive field unless network is deep or uses dilation&lt;/td&gt;
&lt;td&gt;Vanishing/exploding gradients over long sequences&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical applications&lt;/td&gt;
&lt;td&gt;Image classification, object detection, segmentation&lt;/td&gt;
&lt;td&gt;Language modeling, time-series forecasting, speech recognition&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;CNNs assume spatially local structure and share &lt;strong class="kw"&gt;filter weights&lt;/strong&gt; across the whole input; RNNs share weights across time steps instead.&lt;/li&gt;
&lt;li&gt;CNN layers process all positions in parallel, while RNNs are &lt;strong class="kw"&gt;sequential&lt;/strong&gt; by construction since each step needs the prior hidden state.&lt;/li&gt;
&lt;li&gt;RNNs suffer from &lt;strong class="kw"&gt;vanishing gradients&lt;/strong&gt; over long sequences; CNNs sidestep this but need deeper stacks to grow their receptive field.&lt;/li&gt;
&lt;li&gt;Shuffling pixels barely changes what a CNN detects, but reordering a sequence fed to an RNN changes the output entirely, since RNNs are &lt;strong class="kw"&gt;order-sensitive&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;CNNs expect fixed-size grid inputs, whereas RNNs natively handle &lt;strong class="kw"&gt;variable-length&lt;/strong&gt; sequences.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;CNN&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Machine Learning vs Deep Learning: Manual Features vs Learned Representations</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-machine-learning-vs-deep-learning-manual-features-vs-learned/</link><pubDate>Mon, 03 Aug 2026 03:25:19 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-machine-learning-vs-deep-learning-manual-features-vs-learned/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Deep Learning is technically a subset of Machine Learning, but in practice the two names are used to distinguish classical algorithms from neural-network-based approaches. Traditional &lt;strong class="kw"&gt;Machine Learning&lt;/strong&gt; relies on humans to hand-engineer features before a model like a decision tree or SVM can learn from them, while &lt;strong class="kw"&gt;Deep Learning&lt;/strong&gt; uses multi-layer neural networks that learn their own feature representations directly from raw data. The distinction matters because it drives very different requirements for data volume, compute, and interpretability.&lt;/p&gt;</description></item></channel></rss>