Overview
Bagging and Boosting are both ensemble techniques that combine many weak learners into one stronger model, but they build that ensemble in fundamentally different ways. Bagging trains learners independently in parallel on bootstrap samples and averages their outputs to cut variance, while Boosting trains learners one after another, each one correcting the last model’s mistakes through sequential reweighting to cut bias. The choice affects training time, overfitting risk, and how robust the model is to noisy data.
Comparison Diagram
Comparison Table
| Aspect | Bagging | Boosting |
|---|---|---|
| Training order | Base learners trained independently in parallel | Base learners trained sequentially, each depending on the previous one |
| Data sampling per learner | Each learner gets a random bootstrap sample (with replacement) of the training set | Each learner sees the full dataset with sample weights adjusted to emphasize prior errors |
| What each new learner targets | Same objective as every other learner; no awareness of others’ errors | Explicitly targets the residuals or misclassifications left by the previous learner |
| How predictions are combined | Simple averaging (regression) or majority vote (classification), equal weight per learner | Weighted sum where more accurate learners get more influence over the final output |
| Primary statistical effect | Reduces variance by smoothing out individual model noise | Reduces bias by iteratively correcting systematic errors |
| Overfitting risk | Robust to overfitting; adding more learners rarely hurts | Can overfit if run for too many rounds or the learning rate is set too high |
| Sensitivity to noisy data | Relatively robust since outliers only skew a subset of bootstrap samples | More sensitive since misclassified outliers get progressively upweighted |
| Canonical algorithms | Random Forest, Bagged Decision Trees | AdaBoost, Gradient Boosting, XGBoost, LightGBM |
Key Differences
- Bagging trains learners in parallel; Boosting trains them sequentially
- Bagging equally weights models to reduce variance; Boosting weights models to reduce bias
- Bagging resamples data via bootstrapping; Boosting reuses the full dataset with reweighted samples
- Boosting is more prone to overfitting on noisy data than Bagging
When to Use Each
Bagging
- High-Variance Base Models: Unstable models like deep, unpruned decision trees benefit most from bagging’s averaging, which stabilizes their predictions.
- Need Fast Parallel Training: Because learners are independent, bagging fully parallelizes across cores or machines, cutting training time.
- Noisy or Outlier-Heavy Data: Bootstrap resampling limits how much influence any single outlier or mislabeled point has on the final ensemble.
Boosting
- High-Bias Base Models: Simple weak learners like shallow decision stumps benefit from boosting’s iterative error correction.
- Squeezing Out Max Accuracy: On clean, well-prepared datasets, boosting often reaches lower error than bagging by directly attacking residual mistakes.
- Structured Tabular Competitions: Gradient boosting frameworks like XGBoost and LightGBM are the de facto standard on tabular prediction leaderboards.