<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Nlp on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/nlp/</link><description>Recent content in Nlp on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:45:54 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/nlp/index.xml" rel="self" type="application/rss+xml"/><item><title>Zero-Shot Learning vs Few-Shot Learning: No Examples vs a Handful of Examples</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</link><pubDate>Mon, 03 Aug 2026 03:45:54 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-zero-shot-learning-vs-few-shot-learning-no-examples-vs-a-han/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Zero-shot and few-shot learning describe how much task-specific example data a model is given before it has to perform a task. Zero-shot relies solely on a &lt;strong class="kw"&gt;task description&lt;/strong&gt;, while few-shot conditions its predictions on a small set of &lt;strong class="kw"&gt;labeled examples&lt;/strong&gt;, usually trading a little setup cost for higher accuracy.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Zero-Shot&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="16" font-weight="600" style="fill:var(--primary)"&gt;Few-Shot&lt;/text&gt;&lt;line x1="320" y1="45" x2="320" y2="335" stroke-dasharray="4,4" style="stroke:var(--border)" stroke-width="1.5"/&gt;&lt;rect x="60" y="55" width="200" height="95" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="85" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="160" y="103" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;only&lt;/text&gt;&lt;text x="160" y="128" text-anchor="middle" font-size="11" style="fill:var(--secondary)"&gt;(0 examples)&lt;/text&gt;&lt;rect x="380" y="55" width="200" height="95" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="78" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Task instruction&lt;/text&gt;&lt;text x="480" y="96" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;+ K examples&lt;/text&gt;&lt;rect x="400" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="412" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;1&lt;/text&gt;&lt;rect x="430" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="442" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;2&lt;/text&gt;&lt;rect x="460" y="108" width="25" height="22" rx="3" style="fill:none;stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="472" y="123" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;3&lt;/text&gt;&lt;line x1="160" y1="150" x2="160" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,185 155,177 165,177" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="150" x2="480" y2="183" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,185 475,177 485,177" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="160" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="160" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;rect x="380" y="185" width="200" height="50" rx="6" style="fill:none;stroke:var(--border)" stroke-width="1.5"/&gt;&lt;text x="480" y="206" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Pretrained&lt;/text&gt;&lt;text x="480" y="222" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Model&lt;/text&gt;&lt;line x1="160" y1="235" x2="160" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="160,270 155,262 165,262" style="fill:var(--secondary)"/&gt;&lt;line x1="480" y1="235" x2="480" y2="268" style="stroke:var(--secondary)" stroke-width="1.5"/&gt;&lt;polygon points="480,270 475,262 485,262" style="fill:var(--secondary)"/&gt;&lt;rect x="60" y="270" width="200" height="55" rx="6" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;rect x="380" y="270" width="200" height="55" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="302" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Prediction&lt;/text&gt;&lt;text x="160" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;no task-specific data&lt;/text&gt;&lt;text x="480" y="350" text-anchor="middle" font-size="10" style="fill:var(--secondary)"&gt;learns from few examples&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Zero-Shot Learning&lt;/th&gt;
&lt;th&gt;Few-Shot Learning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core definition&lt;/td&gt;
&lt;td&gt;Model performs a task it was never explicitly shown examples for, guided only by natural-language instructions or class descriptions&lt;/td&gt;
&lt;td&gt;Model performs a task after being shown a small number (typically 1-100) of labeled examples at inference or fine-tuning time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examples provided at inference&lt;/td&gt;
&lt;td&gt;None — only a task description or prompt&lt;/td&gt;
&lt;td&gt;A handful of input-output pairs included in the prompt or used for fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Underlying mechanism&lt;/td&gt;
&lt;td&gt;Relies entirely on knowledge encoded during pretraining plus semantic alignment between labels and text&lt;/td&gt;
&lt;td&gt;Uses in-context learning or lightweight fine-tuning to infer the task pattern directly from the provided examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Labeling/data cost&lt;/td&gt;
&lt;td&gt;Effectively zero — no labeled data needed for the target task&lt;/td&gt;
&lt;td&gt;Low but nonzero — requires curating a small, representative set of examples&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt/context length&lt;/td&gt;
&lt;td&gt;Short — just the instruction or class names&lt;/td&gt;
&lt;td&gt;Longer — instruction plus example pairs, consuming more context tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical accuracy&lt;/td&gt;
&lt;td&gt;Lower and more variable, especially on niche or ambiguous tasks&lt;/td&gt;
&lt;td&gt;Generally higher and more stable since examples disambiguate intent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sensitivity to example choice&lt;/td&gt;
&lt;td&gt;Not applicable — there are no examples to choose&lt;/td&gt;
&lt;td&gt;High — accuracy can swing significantly with example selection, order, and count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common techniques&lt;/td&gt;
&lt;td&gt;Prompt engineering, CLIP-style embedding matching, instruction-tuned LLMs&lt;/td&gt;
&lt;td&gt;Few-shot prompting, meta-learning (e.g. MAML), lightweight fine-tuning or LoRA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Zero-shot uses no &lt;strong class="kw"&gt;task examples&lt;/strong&gt; at all, relying purely on pretrained knowledge and instructions&lt;/li&gt;
&lt;li&gt;Few-shot conditions the model on a small &lt;strong class="kw"&gt;support set&lt;/strong&gt; of labeled examples at inference time&lt;/li&gt;
&lt;li&gt;Few-shot generally achieves higher &lt;strong class="kw"&gt;accuracy&lt;/strong&gt; because examples disambiguate an otherwise vague instruction&lt;/li&gt;
&lt;li&gt;Zero-shot has zero &lt;strong class="kw"&gt;labeling cost&lt;/strong&gt;, while few-shot requires curating representative examples&lt;/li&gt;
&lt;li&gt;Few-shot performance is sensitive to &lt;strong class="kw"&gt;example selection&lt;/strong&gt;, a variable that zero-shot simply doesn&amp;rsquo;t have&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Zero-Shot Learning&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>BERT vs GPT: Bidirectional Understanding vs Autoregressive Generation</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-bert-vs-gpt-bidirectional-understanding-vs-autoregressive-ge/</link><pubDate>Mon, 03 Aug 2026 03:42:01 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-bert-vs-gpt-bidirectional-understanding-vs-autoregressive-ge/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;BERT and GPT are both transformer-based language models, but they&amp;rsquo;re built from opposite halves of the transformer and trained for opposite jobs. BERT uses an &lt;strong class="kw"&gt;encoder&lt;/strong&gt; trained to fill in masked words using context from both directions, making it suited to understanding text, while GPT uses a &lt;strong class="kw"&gt;decoder&lt;/strong&gt; trained to predict the next word from only what came before, making it suited to generating text.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;
&lt;defs&gt;
&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;
&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-a)"/&gt;
&lt;/marker&gt;
&lt;marker id="arrowB" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;
&lt;path d="M0,0L10,5L0,10z" style="fill:var(--compare-b)"/&gt;
&lt;/marker&gt;
&lt;/defs&gt;
&lt;line x1="320" y1="20" x2="320" y2="300" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4 4"/&gt;
&lt;text x="160" y="36" text-anchor="middle" style="fill:var(--primary)" font-size="22" font-weight="bold"&gt;BERT&lt;/text&gt;
&lt;text x="480" y="36" text-anchor="middle" style="fill:var(--primary)" font-size="22" font-weight="bold"&gt;GPT&lt;/text&gt;
&lt;text x="160" y="56" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Bidirectional Encoder&lt;/text&gt;
&lt;text x="480" y="56" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;Autoregressive Decoder&lt;/text&gt;
&lt;path d="M50,153 Q100,85 150,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M100,153 Q125,120 150,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M150,153 Q175,120 200,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;path d="M150,153 Q200,85 250,153" style="stroke:var(--compare-a);fill:none" stroke-width="1.3" marker-start="url(#arrowA)" marker-end="url(#arrowA)"/&gt;
&lt;rect x="30" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="50" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;the&lt;/text&gt;
&lt;rect x="80" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="100" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;cat&lt;/text&gt;
&lt;rect x="130" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="2" stroke-dasharray="3 2"/&gt;
&lt;text x="150" y="173" text-anchor="middle" style="fill:var(--primary)" font-size="9" font-weight="bold"&gt;[MASK]&lt;/text&gt;
&lt;rect x="180" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="200" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;on&lt;/text&gt;
&lt;rect x="230" y="155" width="40" height="30" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="250" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;mat&lt;/text&gt;
&lt;text x="160" y="222" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;sees full sentence context&lt;/text&gt;
&lt;text x="160" y="237" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;(left + right) to fill the mask&lt;/text&gt;
&lt;text x="160" y="270" text-anchor="middle" style="fill:var(--content)" font-size="12" font-weight="bold"&gt;-&amp;gt; classification, embeddings, NER&lt;/text&gt;
&lt;path d="M370,153 Q460,90 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M415,153 Q482,110 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M460,153 Q505,125 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M505,153 Q527,135 550,153" style="stroke:var(--compare-b);fill:none" stroke-width="1.3" marker-end="url(#arrowB)"/&gt;
&lt;path d="M570,170 L574,170" style="stroke:var(--compare-b);fill:none" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;
&lt;rect x="350" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="370" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;the&lt;/text&gt;
&lt;rect x="395" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="415" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;cat&lt;/text&gt;
&lt;rect x="440" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="460" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;sat&lt;/text&gt;
&lt;rect x="485" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="505" y="174" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;on&lt;/text&gt;
&lt;rect x="530" y="155" width="40" height="30" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="2"/&gt;
&lt;text x="550" y="174" text-anchor="middle" style="fill:var(--primary)" font-size="11" font-weight="bold"&gt;mat&lt;/text&gt;
&lt;rect x="575" y="155" width="34" height="30" rx="4" style="fill:none;stroke:var(--border)" stroke-width="1.5" stroke-dasharray="3 2"/&gt;
&lt;text x="592" y="174" text-anchor="middle" style="fill:var(--secondary)" font-size="12"&gt;?&lt;/text&gt;
&lt;text x="480" y="222" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;each token sees only itself&lt;/text&gt;
&lt;text x="480" y="237" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;+ prior tokens (causal mask)&lt;/text&gt;
&lt;text x="480" y="270" text-anchor="middle" style="fill:var(--content)" font-size="12" font-weight="bold"&gt;-&amp;gt; generation, chat, completion&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;BERT&lt;/th&gt;
&lt;th&gt;GPT&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;Encoder-only transformer stack&lt;/td&gt;
&lt;td&gt;Decoder-only transformer stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pretraining objective&lt;/td&gt;
&lt;td&gt;Masked language modeling: predict randomly hidden tokens, plus next-sentence prediction&lt;/td&gt;
&lt;td&gt;Causal language modeling: predict the next token given all prior tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention pattern&lt;/td&gt;
&lt;td&gt;Bidirectional self-attention; every token attends to the full sequence&lt;/td&gt;
&lt;td&gt;Causal (masked) self-attention; each token attends only to itself and earlier tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output generation&lt;/td&gt;
&lt;td&gt;One contextual embedding per input token, produced in a single forward pass&lt;/td&gt;
&lt;td&gt;Text generated autoregressively, one token at a time, each output fed back as input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical adaptation&lt;/td&gt;
&lt;td&gt;Fine-tuned with a task-specific head on top of the pretrained encoder&lt;/td&gt;
&lt;td&gt;Adapted via prompting, instruction tuning, or fine-tuning to continue text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary use cases&lt;/td&gt;
&lt;td&gt;Classification, named entity recognition, semantic search, sentence embeddings&lt;/td&gt;
&lt;td&gt;Open-ended generation, chat, code completion, summarization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference cost per query&lt;/td&gt;
&lt;td&gt;Fixed: one pass regardless of desired output&lt;/td&gt;
&lt;td&gt;Scales with number of generated tokens, each requiring a forward pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;BERT&amp;rsquo;s &lt;strong class="kw"&gt;encoder&lt;/strong&gt; attends to both left and right context; GPT&amp;rsquo;s &lt;strong class="kw"&gt;decoder&lt;/strong&gt; attends only to prior tokens.&lt;/li&gt;
&lt;li&gt;BERT trains on &lt;strong class="kw"&gt;masked language modeling&lt;/strong&gt;; GPT trains on &lt;strong class="kw"&gt;next-token prediction&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;BERT produces embeddings in a &lt;strong class="kw"&gt;single pass&lt;/strong&gt;; GPT produces text through &lt;strong class="kw"&gt;autoregressive decoding&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;BERT is optimized for &lt;strong class="kw"&gt;understanding tasks&lt;/strong&gt;; GPT is optimized for &lt;strong class="kw"&gt;generation tasks&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;BERT&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Transformer vs RNN: Parallel Attention vs Sequential Recurrence</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</link><pubDate>Mon, 03 Aug 2026 03:40:07 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;RNNs process sequences one token at a time, carrying context forward through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt; that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;. The difference reshapes everything from training speed to how well long-range context survives.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="20" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="5,5"/&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;RNN&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;Transformer&lt;/text&gt;&lt;circle cx="70" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="140" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="210" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="280" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="140" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="210" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="280" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="45" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="115" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="185" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="255" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h1&lt;/text&gt;&lt;text x="140" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h2&lt;/text&gt;&lt;text x="210" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h3&lt;/text&gt;&lt;text x="280" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h4&lt;/text&gt;&lt;line x1="70" y1="286" x2="70" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="140" y1="286" x2="140" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="210" y1="286" x2="210" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="280" y1="286" x2="280" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="95" y1="180" x2="113" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="165" y1="180" x2="183" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="235" y1="180" x2="253" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;state passed step by step&lt;/text&gt;&lt;text x="160" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes one token at a time&lt;/text&gt;&lt;circle cx="390" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="530" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="600" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="460" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="530" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="600" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="365" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="435" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="505" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="575" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z1&lt;/text&gt;&lt;text x="460" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z2&lt;/text&gt;&lt;text x="530" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z3&lt;/text&gt;&lt;text x="600" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z4&lt;/text&gt;&lt;line x1="390" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="390" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="460" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="530" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;text x="495" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;every token attends to every token&lt;/text&gt;&lt;text x="495" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes all tokens at once&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;RNN&lt;/th&gt;
&lt;th&gt;Transformer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input processing order&lt;/td&gt;
&lt;td&gt;Tokens consumed one at a time, in sequence&lt;/td&gt;
&lt;td&gt;All tokens consumed simultaneously&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context propagation&lt;/td&gt;
&lt;td&gt;Hidden state carried forward step to step&lt;/td&gt;
&lt;td&gt;Self-attention lets each position read all others directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-range dependencies&lt;/td&gt;
&lt;td&gt;Signal weakens over distance (vanishing/exploding gradients)&lt;/td&gt;
&lt;td&gt;Direct connection between any two positions regardless of distance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Positional information&lt;/td&gt;
&lt;td&gt;Implicit, from the order tokens are fed in&lt;/td&gt;
&lt;td&gt;Explicit, via added positional encodings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training parallelization&lt;/td&gt;
&lt;td&gt;Limited — must unroll and step through time&lt;/td&gt;
&lt;td&gt;Fully parallel across the sequence dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computational cost&lt;/td&gt;
&lt;td&gt;O(n) sequential steps, O(1) state per step&lt;/td&gt;
&lt;td&gt;O(n^2) attention cost over sequence length&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference/generation&lt;/td&gt;
&lt;td&gt;Constant memory, naturally one step at a time&lt;/td&gt;
&lt;td&gt;Requires KV caching to avoid recomputing past attention&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use cases&lt;/td&gt;
&lt;td&gt;LSTM/GRU for streaming or small-scale sequence tasks&lt;/td&gt;
&lt;td&gt;BERT/GPT-style models for large-scale language and vision tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Transformer processes all tokens in parallel via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;; RNN processes tokens sequentially through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;RNN suffers from &lt;strong class="kw"&gt;vanishing gradients&lt;/strong&gt; over long sequences; Transformer links any two positions directly&lt;/li&gt;
&lt;li&gt;Transformer needs explicit &lt;strong class="kw"&gt;positional encodings&lt;/strong&gt; since attention has no inherent order; RNN gets order for free&lt;/li&gt;
&lt;li&gt;Transformer training scales with &lt;strong class="kw"&gt;quadratic complexity&lt;/strong&gt; in sequence length; RNN training is linear but hard to parallelize&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;RNN&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>