<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Sequence-Modeling on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/sequence-modeling/</link><description>Recent content in Sequence-Modeling on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 03:40:07 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/sequence-modeling/index.xml" rel="self" type="application/rss+xml"/><item><title>Transformer vs RNN: Parallel Attention vs Sequential Recurrence</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</link><pubDate>Mon, 03 Aug 2026 03:40:07 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-transformer-vs-rnn-parallel-attention-vs-sequential-recurren/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;RNNs process sequences one token at a time, carrying context forward through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt; that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;. The difference reshapes everything from training speed to how well long-range context survives.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="20" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="5,5"/&gt;&lt;text x="160" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;RNN&lt;/text&gt;&lt;text x="480" y="30" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;Transformer&lt;/text&gt;&lt;circle cx="70" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="140" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="210" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;circle cx="280" cy="300" r="14" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="140" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="210" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="280" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="45" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="115" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="185" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="255" y="163" width="50" height="34" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="70" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h1&lt;/text&gt;&lt;text x="140" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h2&lt;/text&gt;&lt;text x="210" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h3&lt;/text&gt;&lt;text x="280" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;h4&lt;/text&gt;&lt;line x1="70" y1="286" x2="70" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="140" y1="286" x2="140" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="210" y1="286" x2="210" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="280" y1="286" x2="280" y2="198" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="95" y1="180" x2="113" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="165" y1="180" x2="183" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;line x1="235" y1="180" x2="253" y2="180" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;state passed step by step&lt;/text&gt;&lt;text x="160" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes one token at a time&lt;/text&gt;&lt;circle cx="390" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="460" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="530" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;circle cx="600" cy="300" r="14" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x1&lt;/text&gt;&lt;text x="460" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x2&lt;/text&gt;&lt;text x="530" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x3&lt;/text&gt;&lt;text x="600" y="304" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;x4&lt;/text&gt;&lt;rect x="365" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="435" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="505" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;rect x="575" y="163" width="50" height="34" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="390" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z1&lt;/text&gt;&lt;text x="460" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z2&lt;/text&gt;&lt;text x="530" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z3&lt;/text&gt;&lt;text x="600" y="184" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;z4&lt;/text&gt;&lt;line x1="390" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="390" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="390" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="460" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="460" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="530" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;line x1="530" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="390" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="460" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="530" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.35"/&gt;&lt;line x1="600" y1="286" x2="600" y2="198" style="stroke:var(--compare-b)" stroke-width="1" opacity="0.7"/&gt;&lt;text x="495" y="120" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;every token attends to every token&lt;/text&gt;&lt;text x="495" y="340" text-anchor="middle" font-size="12" style="fill:var(--secondary)"&gt;processes all tokens at once&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;RNN&lt;/th&gt;
&lt;th&gt;Transformer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input processing order&lt;/td&gt;
&lt;td&gt;Tokens consumed one at a time, in sequence&lt;/td&gt;
&lt;td&gt;All tokens consumed simultaneously&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context propagation&lt;/td&gt;
&lt;td&gt;Hidden state carried forward step to step&lt;/td&gt;
&lt;td&gt;Self-attention lets each position read all others directly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-range dependencies&lt;/td&gt;
&lt;td&gt;Signal weakens over distance (vanishing/exploding gradients)&lt;/td&gt;
&lt;td&gt;Direct connection between any two positions regardless of distance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Positional information&lt;/td&gt;
&lt;td&gt;Implicit, from the order tokens are fed in&lt;/td&gt;
&lt;td&gt;Explicit, via added positional encodings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training parallelization&lt;/td&gt;
&lt;td&gt;Limited — must unroll and step through time&lt;/td&gt;
&lt;td&gt;Fully parallel across the sequence dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computational cost&lt;/td&gt;
&lt;td&gt;O(n) sequential steps, O(1) state per step&lt;/td&gt;
&lt;td&gt;O(n^2) attention cost over sequence length&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference/generation&lt;/td&gt;
&lt;td&gt;Constant memory, naturally one step at a time&lt;/td&gt;
&lt;td&gt;Requires KV caching to avoid recomputing past attention&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use cases&lt;/td&gt;
&lt;td&gt;LSTM/GRU for streaming or small-scale sequence tasks&lt;/td&gt;
&lt;td&gt;BERT/GPT-style models for large-scale language and vision tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Transformer processes all tokens in parallel via &lt;strong class="kw"&gt;self-attention&lt;/strong&gt;; RNN processes tokens sequentially through a &lt;strong class="kw"&gt;hidden state&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;RNN suffers from &lt;strong class="kw"&gt;vanishing gradients&lt;/strong&gt; over long sequences; Transformer links any two positions directly&lt;/li&gt;
&lt;li&gt;Transformer needs explicit &lt;strong class="kw"&gt;positional encodings&lt;/strong&gt; since attention has no inherent order; RNN gets order for free&lt;/li&gt;
&lt;li&gt;Transformer training scales with &lt;strong class="kw"&gt;quadratic complexity&lt;/strong&gt; in sequence length; RNN training is linear but hard to parallelize&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;RNN&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>LSTM vs GRU: Three Gates vs Two Gates in Recurrent Memory</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-lstm-vs-gru-three-gates-vs-two-gates-in-recurrent-memory/</link><pubDate>Mon, 03 Aug 2026 03:33:36 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-lstm-vs-gru-three-gates-vs-two-gates-in-recurrent-memory/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;LSTM and GRU are both gated recurrent architectures built to capture long-range dependencies in sequences while avoiding the vanishing-gradient problem of vanilla RNNs. LSTM keeps a dedicated &lt;strong class="kw"&gt;cell state&lt;/strong&gt; alongside its hidden state, regulated by three gates, while GRU folds everything into a single &lt;strong class="kw"&gt;hidden state&lt;/strong&gt; updated by just two gates. That structural difference drives everything else: parameter count, training speed, and how precisely you can control what the network remembers.&lt;/p&gt;</description></item></channel></rss>