Overview Zero-shot and few-shot learning describe how much task-specific example data a model is given before it has to perform a task. Zero-shot relies solely on a task description, while few-shot conditions its predictions on a small set of labeled examples, usually trading a little setup cost for higher accuracy.
Comparison Diagram Zero-ShotFew-ShotTask instructiononly(0 examples)Task instruction+ K examples123PretrainedModelPretrainedModelPredictionPredictionno task-specific datalearns from few examples Comparison Table Aspect Zero-Shot Learning Few-Shot Learning Core definition Model performs a task it was never explicitly shown examples for, guided only by natural-language instructions or class descriptions Model performs a task after being shown a small number (typically 1-100) of labeled examples at inference or fine-tuning time Examples provided at inference None — only a task description or prompt A handful of input-output pairs included in the prompt or used for fine-tuning Underlying mechanism Relies entirely on knowledge encoded during pretraining plus semantic alignment between labels and text Uses in-context learning or lightweight fine-tuning to infer the task pattern directly from the provided examples Labeling/data cost Effectively zero — no labeled data needed for the target task Low but nonzero — requires curating a small, representative set of examples Prompt/context length Short — just the instruction or class names Longer — instruction plus example pairs, consuming more context tokens Typical accuracy Lower and more variable, especially on niche or ambiguous tasks Generally higher and more stable since examples disambiguate intent Sensitivity to example choice Not applicable — there are no examples to choose High — accuracy can swing significantly with example selection, order, and count Common techniques Prompt engineering, CLIP-style embedding matching, instruction-tuned LLMs Few-shot prompting, meta-learning (e.g. MAML), lightweight fine-tuning or LoRA Key Differences Zero-shot uses no task examples at all, relying purely on pretrained knowledge and instructions Few-shot conditions the model on a small support set of labeled examples at inference time Few-shot generally achieves higher accuracy because examples disambiguate an otherwise vague instruction Zero-shot has zero labeling cost, while few-shot requires curating representative examples Few-shot performance is sensitive to example selection, a variable that zero-shot simply doesn’t have When to Use Each Zero-Shot Learning
...