Ch. 28 · Machine Learning

Bias-Variance Tradeoff Explained: ML Interview Guide

What bias and variance really measure, how to read validation and learning curves, and which fix to reach for when a model underfits or overfits.

~9 min readintermediateupdated Oct 6, 2026

“Explain the bias-variance tradeoff” is one of the most common machine learning interview questions, and also one of the most poorly answered. Most candidates recite “simple models have high bias, complex models have high variance” and stop. Interviewers ask it because the follow-up is practical: here are a training score and a validation score, what is wrong with the model and what will you try next? A candidate who can turn the theory into a diagnosis, and the diagnosis into the right fix, is a candidate who will not spend a month collecting data that cannot help.

Before you start

You should know what a training set and a validation set are, and what a model’s “complexity” might mean: tree depth, polynomial degree, number of neighbours in k-NN. The code uses Python with NumPy 2.5 and scikit-learn 1.9; outputs in comments come from real runs with fixed seeds, so your numbers should match on the same versions and may differ slightly on others. No calculus is needed; the decomposition is shown by simulation rather than by derivation.

The short answer

The expected error of a model on new data can be split into three parts: bias squared, the error from the model’s assumptions being wrong; variance, the error from the model changing a lot depending on which training sample it saw; and irreducible noise. Making a model more flexible usually lowers bias and raises variance, so the best model sits in between. In practice you diagnose it from scores: high training and validation error that are close together means high bias, and a low training error with a much higher validation error means high variance.

How it works

Picture training the same kind of model many times, each time on a fresh sample drawn from the same process. At any input x you get a spread of predictions. Bias is how far the average of those predictions is from the truth. Variance is how widely the predictions scatter around their own average. For squared error, the expected error at x is exactly bias squared plus variance plus the noise variance, because the cross terms cancel when you average.

You cannot see this on a real dataset, because you only have one training sample. With synthetic data you can draw as many as you like, which makes the decomposition visible. Here the truth is a sine wave with noise of standard deviation 0.3 (noise variance 0.09), and polynomials of increasing degree are fitted to 300 different samples of 30 points:

import numpy as np

rng = np.random.default_rng(0)
x_test = np.linspace(0.1, 0.9, 50)   # stay inside the training range
true_f = np.sin(2 * np.pi * x_test)

def bias_variance(degree, n_datasets=300, n_points=30, noise=0.3):
    preds = np.empty((n_datasets, x_test.size))
    for i in range(n_datasets):
        x = rng.uniform(0, 1, n_points)
        y = np.sin(2 * np.pi * x) + rng.normal(0, noise, n_points)
        coefs = np.polyfit(x, y, degree)
        preds[i] = np.polyval(coefs, x_test)
    mean_pred = preds.mean(axis=0)
    bias_sq = np.mean((mean_pred - true_f) ** 2)
    variance = np.mean(preds.var(axis=0))
    return bias_sq, variance

for degree in (1, 3, 6, 9):
    b, v = bias_variance(degree)
    print(f"degree {degree}: bias^2={b:.3f} variance={v:.3f} total={b + v + 0.09:.3f}")
# degree 1: bias^2=0.151 variance=0.019 total=0.260
# degree 3: bias^2=0.003 variance=0.010 total=0.104
# degree 6: bias^2=0.000 variance=0.030 total=0.120
# degree 9: bias^2=0.001 variance=0.400 total=0.491
python

A straight line cannot bend into a sine wave, so its average prediction is wrong everywhere: high bias. It barely changes between samples, so its variance is small. A cubic can follow the curve, so bias almost vanishes. By degree 9 the average prediction is still right (bias is near zero) but each individual fit swings wildly to chase its own noise, and variance dominates. The total error is a U shape with its floor at a moderate complexity, and it can never go below the 0.09 of noise.

Two points are worth saying in an interview. Once the model can represent the truth, extra flexibility buys only variance. And variance depends on the amount of data as well as the model: the same degree 9 polynomial is much calmer with 400 points, which is why “more data” fixes variance and does nothing for bias.

Step-by-step walkthrough

On real data you have one sample, so you estimate the same picture with cross-validation. The tools are a validation curve (score against complexity) and a learning curve (score against training size).

Step 1: Build a dataset with known noise

make_classification gives a controllable problem. flip_y=0.1 randomly flips 10% of labels, so the best possible accuracy is around 0.9 rather than 1.0, which plays the role of irreducible noise:

from sklearn.datasets import make_classification

X, y = make_classification(n_samples=2000, n_features=20, n_informative=5,
                           flip_y=0.1, random_state=0)
python

Knowing the ceiling matters: 0.84 is close to excellent here.

Step 2: Sweep complexity with a validation curve

validation_curve runs cross-validation for each value of one hyperparameter and returns training and validation scores per fold:

from sklearn.model_selection import validation_curve
from sklearn.tree import DecisionTreeClassifier

depths = [1, 3, 5, 8, 12, None]
train, val = validation_curve(
    DecisionTreeClassifier(random_state=0), X, y,
    param_name="max_depth", param_range=depths, cv=5, scoring="accuracy")
for d, tr, va in zip(depths, train.mean(axis=1), val.mean(axis=1)):
    print(f"max_depth={str(d):>4}  train={tr:.3f}  val={va:.3f}  gap={tr - va:.3f}")
# max_depth=   1  train=0.747  val=0.740  gap=0.007
# max_depth=   3  train=0.834  val=0.817  gap=0.018
# max_depth=   5  train=0.885  val=0.838  gap=0.048
# max_depth=   8  train=0.939  val=0.833  gap=0.106
# max_depth=  12  train=0.983  val=0.804  gap=0.179
# max_depth=None  train=1.000  val=0.798  gap=0.202
python

Read it from the top. At depth 1, training and validation agree but both are low: the model is too simple to use the five informative features, which is high bias. As depth grows, training accuracy climbs toward 1.0 while validation peaks at depth 5 and then declines. The widening gap is variance: deep trees memorise the flipped labels. Depth 5 is the sweet spot, and notice that it still has a gap; zero gap is not the goal, the best validation score is.

Step 3: Ask whether more data would help

A learning curve holds the model fixed and varies the training size:

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import learning_curve

for name, model in [("deep tree", DecisionTreeClassifier(random_state=0)),
                    ("logistic", LogisticRegression(max_iter=1000))]:
    sizes, train, val = learning_curve(model, X, y, cv=5,
                                       train_sizes=[0.1, 0.3, 0.6, 1.0])
    print(name)
    for n, tr, va in zip(sizes, train.mean(axis=1), val.mean(axis=1)):
        print(f"  n={n:>4}  train={tr:.3f}  val={va:.3f}")
# deep tree
#   n= 160  train=1.000  val=0.731
#   n= 480  train=1.000  val=0.758
#   n= 960  train=1.000  val=0.783
#   n=1600  train=1.000  val=0.798
# logistic
#   n= 160  train=0.826  val=0.769
#   n= 480  train=0.830  val=0.795
#   n= 960  train=0.814  val=0.802
#   n=1600  train=0.816  val=0.806
python

The unconstrained tree has a large gap that is still closing at 1,600 rows, so more data would keep helping: a variance problem. Logistic regression’s curves have nearly met at about 0.81. Doubling the data would add little, because the remaining error is bias from its linear boundary. To improve it you need a more expressive model or better features, not more rows.

Step 4: Match the fix to the diagnosis

For high bias: add informative features or interactions, use a more flexible model (deeper trees, boosting, a kernel), reduce regularization, or train longer if training stopped early. For high variance: collect more data, simplify (limit depth, fewer features), regularize more, stop early, or average many models with bagging. Random forests exist precisely because averaging many deep, decorrelated trees keeps their low bias and removes much of their variance.

Worked scenario

A team builds a churn classifier with k-nearest neighbours and reports that it is “perfect in training”. They used k=1:

from sklearn.model_selection import cross_validate, GridSearchCV
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

broken = make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=1))
res = cross_validate(broken, X_train, y_train, cv=5, return_train_score=True)
print(f"k=1  train={res['train_score'].mean():.3f}  val={res['test_score'].mean():.3f}")
# k=1  train=1.000  val=0.751
python

With k = 1, every training point is its own nearest neighbour, so training accuracy is 1.0 by construction and says nothing. Validation accuracy of 0.751 and a gap of 0.25 is a textbook variance problem. The fix is to treat k as the bias-variance dial and choose it by cross-validation:

search = GridSearchCV(make_pipeline(StandardScaler(), KNeighborsClassifier()),
                      {"kneighborsclassifier__n_neighbors": [1, 5, 15, 31, 61, 121, 301]},
                      cv=5, return_train_score=True)
search.fit(X_train, y_train)
# k=1    train=1.000 val=0.751
# k=5    train=0.853 val=0.803
# k=15   train=0.834 val=0.810
# k=31   train=0.823 val=0.810
# k=61   train=0.813 val=0.807
# k=121  train=0.801 val=0.794
# k=301  train=0.778 val=0.769
print(search.best_params_, round(search.score(X_test, y_test), 3))
# {'kneighborsclassifier__n_neighbors': 15} 0.841
python

Small k overfits, large k averages over so many neighbours that the boundary is too smooth and both scores fall together, which is bias. The plateau between 15 and 61 is where the tradeoff balances. (The data here is make_classification(n_samples=3000, n_features=8, n_informative=4, flip_y=0.15, class_sep=0.8, random_state=1) with a stratified 75/25 split.)

Common mistake

  • “Complex models have high variance, so always use simple ones.” Complexity is only too high relative to the data you have. A deep network on ten million rows can have lower variance than a modest tree on two hundred.
  • Diagnosing from the training score alone. A training accuracy of 1.0 is a symptom only when compared with validation.
  • Treating a zero gap as the goal. The goal is the best validation score; a small gap at a higher validation score beats no gap at a lower one.
  • Collecting more data for a bias problem. If the learning curves have already converged, more rows will not move them.

Verify the behavior

Turn the diagnosis into an assertion you can keep in a notebook or test suite. This checks that the deep tree has a large gap and that limiting depth improves validation accuracy on the same folds:

from sklearn.model_selection import cross_validate

def gap_and_val(model):
    r = cross_validate(model, X, y, cv=5, return_train_score=True)
    return r["train_score"].mean() - r["test_score"].mean(), r["test_score"].mean()

deep_gap, deep_val = gap_and_val(DecisionTreeClassifier(random_state=0))
cap_gap, cap_val = gap_and_val(DecisionTreeClassifier(max_depth=5, random_state=0))
assert deep_gap > 0.15 and cap_gap < 0.06
assert cap_val > deep_val
print(round(deep_val, 3), round(cap_val, 3))   # 0.798 0.838
python

Follow-up questions

  • Why does bagging reduce variance but not bias? Averaging many models trained on bootstrap samples cancels their independent errors, but if every model is systematically wrong in the same direction, the average is wrong in that direction too.
  • How does regularization fit in? A penalty restricts the effective complexity of the model: it adds a little bias in exchange for a larger drop in variance.
  • What is double descent? In very over-parameterised models such as large neural networks, test error can fall again after the interpolation point. It does not contradict the decomposition, but it shows that parameter count is a poor measure of effective complexity.
  • Does boosting reduce bias or variance? Mainly bias: each shallow tree corrects what the ensemble still gets wrong, which is why it needs a learning rate and early stopping to keep variance in check.

Interview exercise

You train a gradient boosting model for credit default. Cross-validated AUC is 0.71 on training folds and 0.70 on validation folds. Adding six more months of data last quarter moved validation AUC from 0.698 to 0.701. Your manager proposes buying another year of historical data. What do you tell them?

Answer and reasoning

The training and validation scores are almost identical, and they are both modest, so this is a high-bias situation: the model is not overfitting, it is failing to capture the signal. The learning curve is already flat, as the last data addition showed, so another year of the same features is unlikely to help much. I would spend the effort on bias instead: engineer better features (payment history ratios, utilisation trends, interactions), check whether the model is over-regularized (very shallow trees, a tiny learning rate stopped too early, strong L2 on leaves), and confirm the labels are not so noisy that 0.70 is near the ceiling. Buying data becomes worth it only if it brings new information, such as bureau data, rather than more rows of the same columns.

Continue learning

More in Machine Learning

esc