Transformer vs RNN: Parallel Attention vs Sequential Recurrence

Overview RNNs process sequences one token at a time, carrying context forward through a hidden state that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via self-attention. The difference reshapes everything from training speed to how well long-range context survives. Comparison Diagram RNNTransformerx1x2x3x4h1h2h3h4state passed step by stepprocesses one token at a timex1x2x3x4z1z2z3z4every token attends to every tokenprocesses all tokens at once Comparison Table Aspect RNN Transformer Input processing order Tokens consumed one at a time, in sequence All tokens consumed simultaneously Context propagation Hidden state carried forward step to step Self-attention lets each position read all others directly Long-range dependencies Signal weakens over distance (vanishing/exploding gradients) Direct connection between any two positions regardless of distance Positional information Implicit, from the order tokens are fed in Explicit, via added positional encodings Training parallelization Limited — must unroll and step through time Fully parallel across the sequence dimension Computational cost O(n) sequential steps, O(1) state per step O(n^2) attention cost over sequence length Inference/generation Constant memory, naturally one step at a time Requires KV caching to avoid recomputing past attention Typical use cases LSTM/GRU for streaming or small-scale sequence tasks BERT/GPT-style models for large-scale language and vision tasks Key Differences Transformer processes all tokens in parallel via self-attention; RNN processes tokens sequentially through a hidden state RNN suffers from vanishing gradients over long sequences; Transformer links any two positions directly Transformer needs explicit positional encodings since attention has no inherent order; RNN gets order for free Transformer training scales with quadratic complexity in sequence length; RNN training is linear but hard to parallelize When to Use Each RNN ...

August 3, 2026 · 2 min · 374 words · jeonck

LSTM vs GRU: Three Gates vs Two Gates in Recurrent Memory

Overview LSTM and GRU are both gated recurrent architectures built to capture long-range dependencies in sequences while avoiding the vanishing-gradient problem of vanilla RNNs. LSTM keeps a dedicated cell state alongside its hidden state, regulated by three gates, while GRU folds everything into a single hidden state updated by just two gates. That structural difference drives everything else: parameter count, training speed, and how precisely you can control what the network remembers. ...

August 3, 2026 · 3 min · 441 words · jeonck