Zero-Shot Learning vs Few-Shot Learning: No Examples vs a Handful of Examples

Overview Zero-shot and few-shot learning describe how much task-specific example data a model is given before it has to perform a task. Zero-shot relies solely on a task description, while few-shot conditions its predictions on a small set of labeled examples, usually trading a little setup cost for higher accuracy. Comparison Diagram Zero-ShotFew-ShotTask instructiononly(0 examples)Task instruction+ K examples123PretrainedModelPretrainedModelPredictionPredictionno task-specific datalearns from few examples Comparison Table Aspect Zero-Shot Learning Few-Shot Learning Core definition Model performs a task it was never explicitly shown examples for, guided only by natural-language instructions or class descriptions Model performs a task after being shown a small number (typically 1-100) of labeled examples at inference or fine-tuning time Examples provided at inference None — only a task description or prompt A handful of input-output pairs included in the prompt or used for fine-tuning Underlying mechanism Relies entirely on knowledge encoded during pretraining plus semantic alignment between labels and text Uses in-context learning or lightweight fine-tuning to infer the task pattern directly from the provided examples Labeling/data cost Effectively zero — no labeled data needed for the target task Low but nonzero — requires curating a small, representative set of examples Prompt/context length Short — just the instruction or class names Longer — instruction plus example pairs, consuming more context tokens Typical accuracy Lower and more variable, especially on niche or ambiguous tasks Generally higher and more stable since examples disambiguate intent Sensitivity to example choice Not applicable — there are no examples to choose High — accuracy can swing significantly with example selection, order, and count Common techniques Prompt engineering, CLIP-style embedding matching, instruction-tuned LLMs Few-shot prompting, meta-learning (e.g. MAML), lightweight fine-tuning or LoRA Key Differences Zero-shot uses no task examples at all, relying purely on pretrained knowledge and instructions Few-shot conditions the model on a small support set of labeled examples at inference time Few-shot generally achieves higher accuracy because examples disambiguate an otherwise vague instruction Zero-shot has zero labeling cost, while few-shot requires curating representative examples Few-shot performance is sensitive to example selection, a variable that zero-shot simply doesn’t have When to Use Each Zero-Shot Learning ...

August 3, 2026 · 3 min · 463 words · jeonck

BERT vs GPT: Bidirectional Understanding vs Autoregressive Generation

Overview BERT and GPT are both transformer-based language models, but they’re built from opposite halves of the transformer and trained for opposite jobs. BERT uses an encoder trained to fill in masked words using context from both directions, making it suited to understanding text, while GPT uses a decoder trained to predict the next word from only what came before, making it suited to generating text. Comparison Diagram BERT GPT Bidirectional Encoder Autoregressive Decoder the cat [MASK] on mat sees full sentence context (left + right) to fill the mask -> classification, embeddings, NER the cat sat on mat ? each token sees only itself + prior tokens (causal mask) -> generation, chat, completion Comparison Table Aspect BERT GPT Architecture Encoder-only transformer stack Decoder-only transformer stack Pretraining objective Masked language modeling: predict randomly hidden tokens, plus next-sentence prediction Causal language modeling: predict the next token given all prior tokens Attention pattern Bidirectional self-attention; every token attends to the full sequence Causal (masked) self-attention; each token attends only to itself and earlier tokens Output generation One contextual embedding per input token, produced in a single forward pass Text generated autoregressively, one token at a time, each output fed back as input Typical adaptation Fine-tuned with a task-specific head on top of the pretrained encoder Adapted via prompting, instruction tuning, or fine-tuning to continue text Primary use cases Classification, named entity recognition, semantic search, sentence embeddings Open-ended generation, chat, code completion, summarization Inference cost per query Fixed: one pass regardless of desired output Scales with number of generated tokens, each requiring a forward pass Key Differences BERT’s encoder attends to both left and right context; GPT’s decoder attends only to prior tokens. BERT trains on masked language modeling; GPT trains on next-token prediction. BERT produces embeddings in a single pass; GPT produces text through autoregressive decoding. BERT is optimized for understanding tasks; GPT is optimized for generation tasks. When to Use Each BERT ...

August 3, 2026 · 3 min · 431 words · jeonck

Transformer vs RNN: Parallel Attention vs Sequential Recurrence

Overview RNNs process sequences one token at a time, carrying context forward through a hidden state that updates at each step. Transformers instead process every token simultaneously, letting each position directly attend to every other via self-attention. The difference reshapes everything from training speed to how well long-range context survives. Comparison Diagram RNNTransformerx1x2x3x4h1h2h3h4state passed step by stepprocesses one token at a timex1x2x3x4z1z2z3z4every token attends to every tokenprocesses all tokens at once Comparison Table Aspect RNN Transformer Input processing order Tokens consumed one at a time, in sequence All tokens consumed simultaneously Context propagation Hidden state carried forward step to step Self-attention lets each position read all others directly Long-range dependencies Signal weakens over distance (vanishing/exploding gradients) Direct connection between any two positions regardless of distance Positional information Implicit, from the order tokens are fed in Explicit, via added positional encodings Training parallelization Limited — must unroll and step through time Fully parallel across the sequence dimension Computational cost O(n) sequential steps, O(1) state per step O(n^2) attention cost over sequence length Inference/generation Constant memory, naturally one step at a time Requires KV caching to avoid recomputing past attention Typical use cases LSTM/GRU for streaming or small-scale sequence tasks BERT/GPT-style models for large-scale language and vision tasks Key Differences Transformer processes all tokens in parallel via self-attention; RNN processes tokens sequentially through a hidden state RNN suffers from vanishing gradients over long sequences; Transformer links any two positions directly Transformer needs explicit positional encodings since attention has no inherent order; RNN gets order for free Transformer training scales with quadratic complexity in sequence length; RNN training is linear but hard to parallelize When to Use Each RNN ...

August 3, 2026 · 2 min · 374 words · jeonck