Overview

Bagging and Boosting are both ensemble techniques that combine many weak learners into one stronger model, but they build that ensemble in fundamentally different ways. Bagging trains learners independently in parallel on bootstrap samples and averages their outputs to cut variance, while Boosting trains learners one after another, each one correcting the last model’s mistakes through sequential reweighting to cut bias. The choice affects training time, overfitting risk, and how robust the model is to noisy data.

Comparison Diagram

BaggingBoostingTraining DataSample ASample BSample CModel AModel BModel CAggregate(avg / vote)Final PredictionParallel & independent — reduces varianceTraining DataModel 1reweight errorsModel 2reweight errorsModel 3Weighted SumFinal PredictionSequential & dependent — reduces bias

Comparison Table

AspectBaggingBoosting
Training orderBase learners trained independently in parallelBase learners trained sequentially, each depending on the previous one
Data sampling per learnerEach learner gets a random bootstrap sample (with replacement) of the training setEach learner sees the full dataset with sample weights adjusted to emphasize prior errors
What each new learner targetsSame objective as every other learner; no awareness of others’ errorsExplicitly targets the residuals or misclassifications left by the previous learner
How predictions are combinedSimple averaging (regression) or majority vote (classification), equal weight per learnerWeighted sum where more accurate learners get more influence over the final output
Primary statistical effectReduces variance by smoothing out individual model noiseReduces bias by iteratively correcting systematic errors
Overfitting riskRobust to overfitting; adding more learners rarely hurtsCan overfit if run for too many rounds or the learning rate is set too high
Sensitivity to noisy dataRelatively robust since outliers only skew a subset of bootstrap samplesMore sensitive since misclassified outliers get progressively upweighted
Canonical algorithmsRandom Forest, Bagged Decision TreesAdaBoost, Gradient Boosting, XGBoost, LightGBM

Key Differences

  • Bagging trains learners in parallel; Boosting trains them sequentially
  • Bagging equally weights models to reduce variance; Boosting weights models to reduce bias
  • Bagging resamples data via bootstrapping; Boosting reuses the full dataset with reweighted samples
  • Boosting is more prone to overfitting on noisy data than Bagging

When to Use Each

Bagging

  • High-Variance Base Models: Unstable models like deep, unpruned decision trees benefit most from bagging’s averaging, which stabilizes their predictions.
  • Need Fast Parallel Training: Because learners are independent, bagging fully parallelizes across cores or machines, cutting training time.
  • Noisy or Outlier-Heavy Data: Bootstrap resampling limits how much influence any single outlier or mislabeled point has on the final ensemble.

Boosting

  • High-Bias Base Models: Simple weak learners like shallow decision stumps benefit from boosting’s iterative error correction.
  • Squeezing Out Max Accuracy: On clean, well-prepared datasets, boosting often reaches lower error than bagging by directly attacking residual mistakes.
  • Structured Tabular Competitions: Gradient boosting frameworks like XGBoost and LightGBM are the de facto standard on tabular prediction leaderboards.