Ch. 28 · Machine Learning

Overfitting and L1 vs L2 Regularization Explained with Code

How to spot overfitting, why L1 zeroes weights while L2 only shrinks them, and how scaling, early stopping and dropout fit into regularization.

~9 min readintermediateupdated Oct 6, 2026

“What is overfitting and how do you prevent it?” is usually the first question in a machine learning screen, and “what is the difference between L1 and L2 regularization?” is usually the second. Interviewers pair them because the second answer shows whether the first was understood. A candidate who says “regularization prevents overfitting” without explaining what the penalty does to the weights, why lasso produces zeros, or why the same penalty behaves differently on unscaled data, has memorised a slogan. This guide builds the understanding from a model that overfits badly, then fixes it in several ways and measures each fix.

Before you start

You should be comfortable with linear regression as “find weights that minimise squared error” and know what a training and test split is. The code runs on Python with NumPy 2.5 and scikit-learn 1.9. Printed outputs come from real runs with fixed seeds. If bias and variance are new to you, read the bias-variance tradeoff first, because regularization is a way to buy lower variance with a little extra bias.

The short answer

Overfitting means the model has learned noise in the training data, so it scores much better on training data than on new data. Regularization limits how much the model can fit by adding a penalty on weight size to the loss. L2 (ridge) adds the sum of squared weights and shrinks all of them smoothly. L1 (lasso) adds the sum of absolute weights and drives many of them to exactly zero, which also selects features. Both need standardized features, and the penalty strength is chosen by cross-validation. Early stopping, dropout and limits on tree depth are other forms of the same idea.

How it works

A regularized linear model minimises a training loss plus a penalty:

  • Ridge: mean squared error + alpha * sum(w ** 2)
  • Lasso: mean squared error + alpha * sum(abs(w))
  • Elastic net: a weighted mix of the two

The penalty makes large weights expensive. An overfitting model typically has huge, opposing weights that cancel on the training points and explode between them, so pricing weight size squeezes out exactly that behaviour. To see it, fit a degree 15 polynomial to 20 noisy points from a sine curve, with and without a penalty:

import numpy as np
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.metrics import root_mean_squared_error
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures, StandardScaler

rng = np.random.default_rng(42)
def sample(n):
    x = rng.uniform(-1, 1, (n, 1))
    return x, np.sin(3 * x).ravel() + rng.normal(0, 0.2, n)

X_train, y_train = sample(20)
X_test, y_test = sample(500)

models = {
    "no penalty": LinearRegression(),
    "ridge a=0.1": Ridge(alpha=0.1),
    "lasso a=0.01": Lasso(alpha=0.01, max_iter=100_000),
}
for name, reg in models.items():
    m = make_pipeline(PolynomialFeatures(15, include_bias=False), StandardScaler(), reg)
    m.fit(X_train, y_train)
    coef = m[-1].coef_
    print(f"{name:13} train={root_mean_squared_error(y_train, m.predict(X_train)):.3f} "
          f"test={root_mean_squared_error(y_test, m.predict(X_test)):.3f} "
          f"max|w|={np.abs(coef).max():9.1f} zeros={np.sum(coef == 0)}")
# no penalty    train=0.069 test=11.712 max|w|=    731.8 zeros=0
# ridge a=0.1   train=0.130 test=0.260 max|w|=      1.1 zeros=0
# lasso a=0.01  train=0.156 test=0.263 max|w|=      1.0 zeros=10
python

The noise has a standard deviation of 0.2, so no model can do much better than an RMSE of 0.2. The unpenalised fit beats that on training data (0.069), which is the giveaway: it is fitting noise. Its weights reach 731 and its test error is 11.7, because between and beyond the training points the polynomial swings far off the curve. Both penalties pay a little on training error and land near the noise floor on test data. Lasso also switched off 10 of the 15 polynomial terms.

Why does L1 produce exact zeros while L2 does not? Look at the gradient of each penalty with respect to one weight. For L2 it is 2 * alpha * w, which fades as the weight approaches zero, so the push toward zero becomes too weak to get it all the way there. For L1 it is alpha * sign(w), a constant force however small the weight is. If the loss gains less than that from a feature, its weight is pushed to zero and held there. Geometrically, the L1 constraint region is a diamond with corners on the axes, and the loss contours usually touch it at a corner, where some weights are zero.

Step-by-step walkthrough

Step 1: Confirm it is overfitting, not leakage or bad luck

Compare training and validation scores across folds, not one split. A persistent gap, with training error below the noise level you would expect, is overfitting. A validation score that is suspiciously high is a different problem (usually leakage), and a gap that appears only on one fold is usually a small or unlucky split. Cross-validation with return_train_score=True gives both numbers per fold.

Step 2: Watch L1 select features as alpha grows

On a problem where only 5 of 30 features matter, increase the penalty and count the surviving weights:

from sklearn.datasets import make_regression

X, y = make_regression(n_samples=200, n_features=30, n_informative=5, noise=10, random_state=0)
for alpha in (0.01, 1, 5, 20):
    l = Lasso(alpha=alpha, max_iter=50_000).fit(X, y)
    r = Ridge(alpha=alpha * 100).fit(X, y)
    print(f"alpha={alpha:<5} lasso non-zero={np.sum(l.coef_ != 0):2d}   "
          f"ridge(alpha={alpha*100:g}) non-zero={np.sum(r.coef_ != 0):2d}")
# alpha=0.01  lasso non-zero=30   ridge(alpha=1) non-zero=30
# alpha=1     lasso non-zero= 9   ridge(alpha=100) non-zero=30
# alpha=5     lasso non-zero= 5   ridge(alpha=500) non-zero=30
# alpha=20    lasso non-zero= 3   ridge(alpha=2000) non-zero=30
python

Lasso walks down from 30 features to exactly the 5 informative ones, then starts dropping real ones when the penalty is too strong. Ridge keeps all 30 at every strength, just smaller. Choose L1 when you want a sparse, explainable model or suspect most features are irrelevant; choose L2 when many features each carry a little signal or features are correlated, because lasso tends to keep one of a correlated group arbitrarily while ridge shares the weight.

Step 3: Choose the strength by cross-validation

The penalty strength is a hyperparameter. RidgeCV, LassoCV and LogisticRegressionCV search a grid of values efficiently, and GridSearchCV works for anything else. Search on a log scale (0.001, 0.01, 0.1, 1, 10), because the effect of alpha is multiplicative. Beware the naming: alpha in Ridge and Lasso is the penalty strength, but C in LogisticRegression and SVC is its inverse, so a smaller C means more regularization.

Step 4: Use the regularizers that are not penalties

Early stopping limits how long an iterative model trains. Gradient boosting with a validation split stops adding trees when the validation loss has not improved for a number of rounds:

from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier

X, y = make_classification(n_samples=5000, n_features=20, n_informative=6, flip_y=0.1, random_state=0)
gb = HistGradientBoostingClassifier(max_iter=2000, learning_rate=0.1, early_stopping=True,
                                    validation_fraction=0.1, n_iter_no_change=20, random_state=0).fit(X, y)
print(gb.n_iter_)   # 58
python

It used 58 of the 2,000 allowed trees. Dropout does something similar for neural networks: during training each activation is zeroed with probability p, so the network cannot rely on any single unit and effectively trains an ensemble of thinned networks. Frameworks use inverted dropout, scaling the survivors by 1 / (1 - p) during training so nothing needs rescaling at inference, when dropout is switched off:

rng = np.random.default_rng(0)
activations = np.ones(100_000)
p_drop = 0.5
mask = rng.random(activations.shape) >= p_drop
train_out = activations * mask / (1 - p_drop)
print(round(mask.mean(), 3), round(train_out.mean(), 3))   # 0.499 0.998
python

For trees, the equivalents are max_depth, min_samples_leaf and cost-complexity pruning; for any model, more data and data augmentation reduce overfitting without any penalty at all.

Worked scenario

A team predicts monthly customer spend with lasso. The true drivers are a premium-plan flag (worth 40 a month) and income in dollars (0.002 per dollar). They fit lasso directly on the raw columns, plus eight noise columns:

from sklearn.model_selection import cross_val_score

rng = np.random.default_rng(7)
n = 1000
is_premium = rng.integers(0, 2, n)
income = rng.normal(60_000, 15_000, n)
X = np.column_stack([is_premium, income, rng.normal(size=(n, 8))])
y = 40 * is_premium + 0.002 * income + rng.normal(0, 10, n)

unscaled = Lasso(alpha=5).fit(X, y)
print(f"unscaled: premium={unscaled.coef_[0]:.2f} income={unscaled.coef_[1]:.5f}")
# unscaled: premium=19.88 income=0.00194
python

The premium effect was cut in half while income was barely touched. The penalty treats every weight alike, but the weights live on different scales: income needs a tiny weight because its values are huge, so penalising it costs almost nothing, while the 0/1 flag needs a big weight and gets hammered. The model’s conclusions about which feature matters are an artefact of units. The fix is to standardize inside the pipeline:

scaled = make_pipeline(StandardScaler(), Lasso(alpha=5)).fit(X, y)
w = scaled[-1].coef_ / scaled[0].scale_   # back to original units
print(f"scaled:   premium={w[0]:.2f} income={w[1]:.5f}")
# scaled:   premium=29.48 income=0.00162
print(cross_val_score(Lasso(alpha=5), X, y, cv=5).mean().round(3))   # 0.84
print(cross_val_score(scaled, X, y, cv=5).mean().round(3))           # 0.878
python

Now both effects are shrunk by a similar fraction, and cross-validated R-squared improves from 0.84 to 0.878. The scaler sits inside the pipeline so it is refitted on each training fold.

Common mistake

  • “Regularization always improves the model.” It trades bias for variance. On an underfitting model it makes things worse.
  • Regularizing unscaled features, as in the scenario. This also applies to logistic regression, SVMs and neural networks.
  • Penalising the intercept or one-hot reference columns by accident in hand-written code; scikit-learn does not penalise the intercept.
  • Reading lasso’s selected features as the true causes. With correlated features lasso picks one somewhat arbitrarily, and the choice can flip between bootstrap samples.
  • Mixing up alpha and C. Raising C in LogisticRegression weakens regularization.
  • Forgetting that dropout is off at inference. Predictions from a network left in training mode are noisy and slightly wrong.

Verify the behavior

This check confirms that cross-validated lasso keeps every truly informative feature on a problem where you know the answer:

from sklearn.linear_model import LassoCV

X, y, true_coef = make_regression(n_samples=200, n_features=30, n_informative=5,
                                  noise=10, coef=True, random_state=0)
model = make_pipeline(StandardScaler(), LassoCV(cv=5, random_state=0)).fit(X, y)
kept = set(np.flatnonzero(model[-1].coef_))
truth = set(np.flatnonzero(true_coef))
assert truth <= kept, "lasso dropped a real feature"
print(sorted(map(int, truth)), len(kept), round(model[-1].alpha_, 3))
# [16, 19, 21, 24, 29] 10 0.881
python

All 5 real features survive, but so do 5 noise features. That is typical: the alpha that predicts best is smaller than the one that selects exactly, so treat lasso selection as a screening step, not a final answer.

Follow-up questions

  • What is the Bayesian view? L2 is a Gaussian prior on the weights and L1 a Laplace prior; the regularized solution is the maximum a posteriori estimate.
  • When would you use elastic net? When you want sparsity but have groups of correlated features: the L2 part keeps correlated features together.
  • Does regularization help a random forest? Not as a weight penalty; forests are controlled by tree depth, leaf size and max_features, and averaging is itself variance reduction.
  • How do you pick the dropout rate? Treat it as a hyperparameter; 0.1 to 0.5 is common, higher for large dense layers and lower for convolutional ones.

Interview exercise

You train logistic regression on 50,000 TF-IDF text features for 2,000 labelled support tickets. Training accuracy is 99.8% and validation accuracy is 71%. Name three changes you would try, in order, and say what you expect each to do.

Answer and reasoning

With 25 times more features than examples, the model can separate the training set almost perfectly, so this is variance. First, check the strength: scikit-learn already applies L2 with C=1.0, so I would search C on a log scale with LogisticRegressionCV and expect a smaller C to narrow the gap. Second, switch to L1 or elastic net (l1_ratio above 0 with the saga solver), because most of the 50,000 terms are irrelevant and sparsity both regularizes and gives an interpretable vocabulary. Third, reduce the feature space itself: a minimum document frequency, removing near-duplicate tickets that leak between splits, or using pretrained sentence embeddings instead of raw terms. Labelling more tickets would help too, since the gap says the model is data-starved.

Continue learning

More in Machine Learning

esc