Overview

Overfitting and underfitting describe the two ways a model can fail to generalize: one learns the training data too well, the other not well enough. Understanding which failure mode you’re in determines whether you should simplify or add regularization, or instead increase capacity and train longer. Overfitting traps a model in the noise of its training set, while underfitting leaves it unable to capture the underlying pattern at all.

Comparison Diagram

OverfittingUnderfittingFits every point exactlycaptures noise, not signalMisses the curved trendtoo simple for the pattern

Comparison Table

AspectOverfittingUnderfitting
Underlying causeModel too complex relative to the data, so it learns noise and idiosyncrasiesModel too simple to represent the true relationship in the data
Training errorVery low, often near zeroHigh, the model struggles even on data it was trained on
Validation/test errorHigh, much worse than training errorHigh, similar in magnitude to training error
Bias-variance profileLow bias, high varianceHigh bias, low variance
Generalization to new dataPoor, predictions swing wildly on unseen inputsPoor, predictions are consistently and systematically off
Learning curve signatureTraining and validation loss diverge as training continuesTraining and validation loss both plateau high and close together
Typical remediesRegularization, more training data, dropout, early stopping, simpler modelIncrease model capacity, add features, train longer, reduce regularization

Key Differences

  • Overfitting memorizes noise in the training set, while underfitting never learns the underlying pattern at all.
  • Overfitting shows near-zero training error but a wide train/validation gap, a sign of high variance; underfitting shows poor performance on both, a sign of high bias.
  • Overfitting is treated with regularization or more data; underfitting is treated by increasing model capacity.
  • Overfitting gets worse the longer an overly flexible model keeps training; underfitting persists regardless of duration, since it’s a structural limit.

When to Use Each

Overfitting

  • Complex Model, Small Dataset: A high-capacity model like a deep neural net or high-degree polynomial trained on a small dataset will readily memorize training examples instead of generalizing.
  • Perfect Training Accuracy: If training accuracy is near 100% but validation accuracy is much lower, the model has likely learned noise specific to the training set.
  • Too Many Training Epochs: Validation loss rising while training loss keeps falling is a classic sign the model is overfitting to the training data.

Underfitting

  • Oversimplified Model: Using a linear model to fit a clearly nonlinear relationship leaves the model unable to capture the true pattern.
  • Both Errors Stay High: If training and validation accuracy are both low and close together, the model lacks the capacity to learn the data.
  • Excessive Regularization: Applying too much L1/L2 penalty, dropout, or early stopping can prevent even a capable model from fitting the data well.