Transformer vs RNN: Parallel Attention vs Sequential Recurrence
Overview RNNs process sequences one token at a time, carrying context forward through a hidden state that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via self-attention. The difference reshapes everything from training speed to how well long-range context survives. Comparison Diagram RNNTransformerx1x2x3x4h1h2h3h4state passed step by stepprocesses one token at a timex1x2x3x4z1z2z3z4every token attends to every tokenprocesses all tokens at once Comparison Table Aspect RNN Transformer Input processing order Tokens consumed one at a time, in sequence All tokens consumed simultaneously Context propagation Hidden state carried forward step to step Self-attention lets each position read all others directly Long-range dependencies Signal weakens over distance (vanishing/exploding gradients) Direct connection between any two positions regardless of distance Positional information Implicit, from the order tokens are fed in Explicit, via added positional encodings Training parallelization Limited — must unroll and step through time Fully parallel across the sequence dimension Computational cost O(n) sequential steps, O(1) state per step O(n^2) attention cost over sequence length Inference/generation Constant memory, naturally one step at a time Requires KV caching to avoid recomputing past attention Typical use cases LSTM/GRU for streaming or small-scale sequence tasks BERT/GPT-style models for large-scale language and vision tasks Key Differences Transformer processes all tokens in parallel via self-attention; RNN processes tokens sequentially through a hidden state RNN suffers from vanishing gradients over long sequences; Transformer links any two positions directly Transformer needs explicit positional encodings since attention has no inherent order; RNN gets order for free Transformer training scales with quadratic complexity in sequence length; RNN training is linear but hard to parallelize When to Use Each RNN ...