Overview
Overfitting and underfitting describe the two ways a model can fail to generalize: one learns the training data too well, the other not well enough. Understanding which failure mode you’re in determines whether you should simplify or add regularization, or instead increase capacity and train longer. Overfitting traps a model in the noise of its training set, while underfitting leaves it unable to capture the underlying pattern at all.
Comparison Diagram
Comparison Table
| Aspect | Overfitting | Underfitting |
|---|---|---|
| Underlying cause | Model too complex relative to the data, so it learns noise and idiosyncrasies | Model too simple to represent the true relationship in the data |
| Training error | Very low, often near zero | High, the model struggles even on data it was trained on |
| Validation/test error | High, much worse than training error | High, similar in magnitude to training error |
| Bias-variance profile | Low bias, high variance | High bias, low variance |
| Generalization to new data | Poor, predictions swing wildly on unseen inputs | Poor, predictions are consistently and systematically off |
| Learning curve signature | Training and validation loss diverge as training continues | Training and validation loss both plateau high and close together |
| Typical remedies | Regularization, more training data, dropout, early stopping, simpler model | Increase model capacity, add features, train longer, reduce regularization |
Key Differences
- Overfitting memorizes noise in the training set, while underfitting never learns the underlying pattern at all.
- Overfitting shows near-zero training error but a wide train/validation gap, a sign of high variance; underfitting shows poor performance on both, a sign of high bias.
- Overfitting is treated with regularization or more data; underfitting is treated by increasing model capacity.
- Overfitting gets worse the longer an overly flexible model keeps training; underfitting persists regardless of duration, since it’s a structural limit.
When to Use Each
Overfitting
- Complex Model, Small Dataset: A high-capacity model like a deep neural net or high-degree polynomial trained on a small dataset will readily memorize training examples instead of generalizing.
- Perfect Training Accuracy: If training accuracy is near 100% but validation accuracy is much lower, the model has likely learned noise specific to the training set.
- Too Many Training Epochs: Validation loss rising while training loss keeps falling is a classic sign the model is overfitting to the training data.
Underfitting
- Oversimplified Model: Using a linear model to fit a clearly nonlinear relationship leaves the model unable to capture the true pattern.
- Both Errors Stay High: If training and validation accuracy are both low and close together, the model lacks the capacity to learn the data.
- Excessive Regularization: Applying too much L1/L2 penalty, dropout, or early stopping can prevent even a capable model from fitting the data well.