Ch. 28 · Machine Learning

Random Forest vs Gradient Boosting: Bagging vs Boosting

Why a random forest averages away variance while gradient boosting chips away bias, with runnable code, early stopping and when to choose each.

~8 min readintermediateupdated Oct 6, 2026

“What is the difference between a random forest and gradient boosting?” is the most common tree question in machine learning interviews, usually paired with “bagging versus boosting” and “which would you use here?”. Both build ensembles of decision trees, so a vague answer sounds plausible. Interviewers are listening for the mechanism: which error each method attacks, why the trees are deep in one and shallow in the other, what happens when you add more trees, and which hyperparameters you would tune first. Getting those right also explains most of the production surprises these models cause.

Before you start

You should know how a single decision tree splits data on feature thresholds and why a fully grown tree memorises its training set. The ideas of bias and variance from the bias-variance tradeoff are used throughout. Code runs on Python 3.14 with scikit-learn 1.9 and NumPy 2.5; outputs come from real runs with fixed seeds. XGBoost 3.x and LightGBM 4.x are discussed but not run.

The short answer

A random forest is bagging plus feature randomness: it trains many deep trees independently, each on a bootstrap sample of rows and choosing each split from a random subset of features, then averages them. Deep trees have low bias and high variance, and averaging decorrelated trees removes much of the variance. Gradient boosting trains shallow trees sequentially, each one fitted to the negative gradient of the loss (for squared error, the residuals) of the ensemble so far, and adds it scaled by a learning rate. It starts biased and removes bias step by step. Forests are robust baselines that are hard to overfit by adding trees; boosting usually wins on tabular accuracy but needs tuning and early stopping.

How it works

Averaging helps only when the models make different mistakes. If each tree’s prediction has variance σ² and any two trees have error correlation ρ, the average of B trees has variance ρσ² + (1 - ρ)σ² / B. More trees shrink the second term towards zero, but the first term stays. That is why a forest does two things: bootstrap samples give each tree a different view of the rows, and random feature subsets at each split (max_features) stop every tree from splitting on the same dominant feature first, lowering ρ.

Boosting builds an additive model instead. It starts with a constant (the mean of y for squared error), then repeatedly fits a small tree to what is still wrong and adds a fraction of it:

import numpy as np
from sklearn.tree import DecisionTreeRegressor

def boost(X_train, y_train, X_test, rounds=500, lr=0.1, depth=3):
    pred_tr = np.full(len(y_train), y_train.mean())
    pred_te = np.full(len(X_test), y_train.mean())
    for _ in range(rounds):
        tree = DecisionTreeRegressor(max_depth=depth).fit(X_train, y_train - pred_tr)  # residuals
        pred_tr += lr * tree.predict(X_train)
        pred_te += lr * tree.predict(X_test)
    return pred_te
python

For squared error the residual is exactly the negative gradient of the loss with respect to the prediction, which is where the name comes from. For log loss, the tree is fitted to y - p instead, and other losses plug in their own gradients. The learning rate (shrinkage) deliberately takes small steps so later trees can correct earlier ones, and the shallow depth keeps each tree a weak learner.

Step-by-step walkthrough

We use make_friedman1(n_samples=4000, n_features=10, noise=1.0, random_state=0), a non-linear regression problem with five informative features, five useless ones and noise variance 1.0, so a perfect model would score a test MSE of about 1.0. A 70/30 split is used throughout.

Step 1: Measure a single deep tree

from sklearn.datasets import make_friedman1
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error as mse

X, y = make_friedman1(n_samples=4000, n_features=10, noise=1.0, random_state=0)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
tree = DecisionTreeRegressor(random_state=0).fit(X_train, y_train)
print(mse(y_train, tree.predict(X_train)), mse(y_test, tree.predict(X_test)))
# 0.0 7.251
python

Zero training error and a test MSE of 7.25 is textbook variance: the tree fitted every noisy point.

Step 2: Average many trees into a forest

from sklearn.ensemble import RandomForestRegressor

for n in (1, 10, 50, 300):
    rf = RandomForestRegressor(n_estimators=n, oob_score=n >= 50, random_state=0, n_jobs=-1)
    rf.fit(X_train, y_train)
    print(n, round(mse(y_test, rf.predict(X_test)), 3))
# 1 8.147
# 10 3.486
# 50 3.16   (oob R2 0.870, test R2 0.872)
# 300 3.064 (oob R2 0.881, test R2 0.876)
python

Ten trees already halve the error, and after that the curve flattens instead of turning up: adding trees to a forest does not overfit, it just stops helping. The out-of-bag score, computed on the roughly 37% of rows each tree never saw, tracks the test score closely, so a forest gives you a free validation estimate.

Looking inside the 300-tree forest explains the floor. An individual tree has test MSE 8.05, and the average pairwise correlation of tree errors is 0.39. With max_features=0.33 the correlation falls to 0.30, but each tree, now often denied the informative features, gets worse (11.3), and the forest ends slightly worse at 3.26. Note that RandomForestRegressor defaults to max_features=1.0, which is plain bagging, while the classifier defaults to "sqrt".

Step 3: Boost shallow trees instead

Running the boost function from above with depth-3 trees:

# round 1    train 22.03  test 21.997
# round 10   train 10.051 test 10.865
# round 100  train 1.114  test 1.884
# round 500  train 0.461  test 1.446
python

The first rounds are badly underfit, because one depth-3 tree scaled by 0.1 is a weak model. Each round removes some bias, and after 500 rounds the test MSE is 1.45, less than half the forest’s error and close to the noise floor of 1.0. This dataset rewards boosting because its target is a smooth sum of effects, which an additive model of small trees captures well.

Step 4: Let early stopping choose the number of rounds

from sklearn.ensemble import HistGradientBoostingRegressor

hgb = HistGradientBoostingRegressor(learning_rate=0.05, max_iter=2000, max_leaf_nodes=15,
                                    early_stopping=True, validation_fraction=0.15,
                                    n_iter_no_change=50, random_state=0).fit(X_train, y_train)
print(hgb.n_iter_, round(mse(y_test, hgb.predict(X_test)), 3))
# 444 1.499
python

HistGradientBoostingRegressor is scikit-learn’s LightGBM-style implementation: it bins features into histograms, handles missing values natively and is far faster than the classic GradientBoostingRegressor on large data. Early stopping holds out 15% of the training rows, stops when the validation loss has not improved for 50 rounds, and chose 444 here. Its default early_stopping="auto" turns this on only above 10,000 samples, so set it explicitly on smaller data.

Worked scenario

A team compares models on 600 training rows with noisy targets (noise variance 4.0) and concludes boosting is worse than a forest. Their boosting configuration was GradientBoostingRegressor(learning_rate=0.5, max_depth=6, n_estimators=1000):

# round   train_mse  test_mse
# 5       1.217      8.458
# 10      0.429      8.059
# 20      0.104      8.156
# 100     0.0        8.228
# 1000    0.0        8.229
# best round 9 with test MSE 7.983; random forest test MSE 7.145
python

With a large learning rate and deep trees, each round fits the remaining noise aggressively. Training error hits zero by round 100, and the best test score came at round 9, after which the model only memorised. This is the key asymmetry with forests: boosting rounds are not averaged, they are added, so too many of them overfit.

The fix uses the standard boosting regularizers together: a small learning rate, shallow trees, row subsampling, and early stopping on held-out data:

es = GradientBoostingRegressor(learning_rate=0.05, max_depth=3, n_estimators=5000, subsample=0.8,
                               validation_fraction=0.2, n_iter_no_change=50,
                               random_state=0).fit(X_train, y_train)
# early stopped at 203 rounds, train MSE 2.44, test MSE 5.799
python

Test MSE falls from 8.2 to 5.8, now clearly better than the forest’s 7.1. The training error of 2.44 is no longer zero, which is the point: the model stopped before fitting the noise.

Common mistake

  • “Boosting reduces variance like bagging.” Boosting primarily reduces bias. Its regularization tools (shrinkage, subsampling, early stopping) exist because it can overfit.
  • “More trees in a forest cause overfitting.” They do not; the error plateaus. Forest overfitting is controlled by tree depth and min_samples_leaf, not tree count.
  • Tuning the number of boosting rounds on the test set. Early stopping needs its own validation split, or the reported score is optimistic.
  • Expecting trees to extrapolate. Train on x from 0 to 9 with y = 2x and a forest predicts 17 for x = 20 while GradientBoostingRegressor predicts about 18, because every leaf predicts an average of seen targets. A linear model predicts 40.
  • Reading impurity-based feature importance as truth. It favours high-cardinality and continuous features; use permutation importance on validation data instead.
  • Scaling features for trees. Splits are unaffected by monotonic transforms, so scaling does nothing.

Verify the behavior

You can prove that gradient boosting is just a sum of residual-fitting trees by rebuilding a fitted scikit-learn model’s predictions by hand:

from sklearn.ensemble import GradientBoostingRegressor

gb = GradientBoostingRegressor(n_estimators=200, learning_rate=0.1, max_depth=3,
                               random_state=0).fit(X_train, y_train)

# The ensemble is the training mean plus learning_rate times the sum of tree outputs
manual = y_train.mean() + 0.1 * sum(t[0].predict(X_test) for t in gb.estimators_)
assert np.allclose(manual, gb.predict(X_test))

# Tree 5 is exactly the tree you get by fitting the residuals left after 4 trees
resid = y_train - list(gb.staged_predict(X_train))[3]
mine = DecisionTreeRegressor(max_depth=3).fit(X_train, resid)
assert np.allclose(mine.predict(X_test), gb.estimators_[4, 0].predict(X_test))
print("the model is an additive sum of residual-fitting trees")
python

Rebuilding all 200 rounds from scratch can drift slightly from scikit-learn, because near-tied splits deep in training are broken in a different order; replaying one stage at a time avoids that.

Follow-up questions

  • What do XGBoost and LightGBM add? Second-order gradients, L1 and L2 penalties on leaf weights, learned default directions for missing values and histogram-based splitting; LightGBM also grows trees leaf-wise, which is fast but overfits small data unless num_leaves is capped.
  • Which boosting hyperparameters do you tune first? Learning rate with early stopping to set the rounds, then tree size (max_depth or num_leaves, minimum samples per leaf), then row and column subsampling, then penalties.
  • Can you parallelise boosting? Not across rounds, which are sequential, but split finding within a tree is parallelised, which is where XGBoost and LightGBM get their speed.
  • When would you pick a forest anyway? A quick, robust baseline with little tuning, noisy labels, a need for out-of-bag estimates, or embarrassingly parallel training.

Interview exercise

Each tree in a forest has prediction variance 8 and the trees’ errors have pairwise correlation 0.4. What is the variance of the average of 10 trees, of 300 trees and of infinitely many? A colleague proposes going from 300 to 3,000 trees to improve accuracy. What would you suggest instead?

Answer and reasoning

Using ρσ² + (1 - ρ)σ² / B with σ² = 8 and ρ = 0.4: the correlated part is 3.2 and the independent part is 4.8 / B. For 10 trees that gives 3.2 + 0.48 = 3.68, for 300 trees 3.2 + 0.016 = 3.216, and for infinitely many it approaches 3.2. Going to 3,000 trees would remove about 0.014 of variance while costing ten times the memory and prediction latency. The remaining 3.2 is set by correlation, so the useful levers are lowering ρ (try a smaller max_features, while checking that individual trees do not get much worse, as Step 2 showed) or switching to gradient boosting if the problem is really bias. An interviewer wants to hear that the floor is correlation, not tree count.

Continue learning

More in Machine Learning

esc