Ch. 29 · Deep Learning

Dropout and Overfitting in Deep Neural Networks Explained

How to spot overfitting in loss curves, why inverted dropout scales by 1/(1-p), and how weight decay, augmentation and early stopping help.

~9 min readintermediateupdated Oct 6, 2026

“Your model gets 99 percent training accuracy and 70 percent validation accuracy. What do you do?” is how most interviewers open the regularization conversation, and dropout is usually where it goes next: how does dropout work, why are activations scaled, and what changes at inference? It tests whether you can read a training curve, name the cause and pick a remedy for a reason. Strong candidates also know where the classic advice breaks down, such as why huge networks often generalize and why dropout and batch norm can clash.

Before you start

You should know what training and validation sets are, what a loss curve is, and roughly how gradient descent updates weights. The examples are pure Python 3.12 or newer with only the standard library and a seeded random.Random, so every printed number is reproducible. PyTorch snippets are illustrative (PyTorch 2.x) and were not executed for this article.

The short answer

Overfitting means the model has learned patterns specific to the training set, so training loss keeps falling while validation loss bottoms out and rises. Dropout fights it by zeroing each activation with probability p during training, which stops units from relying on particular partners and acts like training an ensemble of thinned networks. Modern frameworks use inverted dropout: surviving activations are multiplied by 1/(1-p) during training, so the expected activation is unchanged and evaluation is just the identity. Dropout sits alongside other remedies: more or augmented data, weight decay, early stopping and, sometimes, a smaller model.

How it works

The whole layer is a mask and a rescale:

import random

def dropout(xs, p, training, rng):
    """Inverted dropout: zero each unit with probability p, scale survivors by 1 / (1 - p)."""
    if not training or p == 0.0:
        return list(xs)                       # identity at eval time
    keep = 1.0 - p
    return [x / keep if rng.random() < keep else 0.0 for x in xs]

rng = random.Random(42)
h = [0.8, 1.5, 0.2, 2.0, 1.1]
print([round(x, 3) for x in dropout(h, 0.5, True, rng)])   # [0.0, 3.0, 0.4, 4.0, 0.0]
print(dropout(h, 0.5, False, rng))                          # [0.8, 1.5, 0.2, 2.0, 1.1]
python

A fresh mask is drawn for every forward pass, so each mini-batch trains a different random subnetwork that shares weights with all the others. No unit can count on a specific neighbour, which discourages fragile co-adaptations. At test time the full network, with expected activations matched, cheaply approximates averaging all those subnetworks. Note that p is the drop probability in PyTorch’s nn.Dropout(p=0.5) and Keras’s Dropout(rate), while the original paper used p for the keep probability.

Step-by-step walkthrough

Step 1: Diagnose from the curves

Before choosing a regularizer, confirm you are overfitting and not underfitting:

train_loss = [1.20, 0.85, 0.62, 0.48, 0.37, 0.29, 0.22, 0.17, 0.13, 0.10]
val_loss   = [1.25, 0.92, 0.74, 0.66, 0.63, 0.64, 0.68, 0.73, 0.80, 0.88]

best = min(range(len(val_loss)), key=val_loss.__getitem__)
print("best epoch:", best, "val:", val_loss[best])   # best epoch: 4 val: 0.63
for epoch in (2, best, 9):
    print(epoch, "gap:", round(val_loss[epoch] - train_loss[epoch], 2))
# 2 gap: 0.12
# 4 gap: 0.26
# 9 gap: 0.78
python

Up to epoch 4 both losses fall: the model is still learning general structure. After that, training loss keeps dropping while validation loss climbs, and the gap widens from 0.26 to 0.78. That shape is overfitting. If both curves had flattened at a high value, the model would be underfitting, and adding dropout would make things worse; you would want more capacity, more training or a better learning rate instead.

Step 2: Check that inverted dropout preserves the expectation

def mean_output(x, p, trials, rng, inverted=True):
    total = 0.0
    for _ in range(trials):
        if inverted:
            total += dropout([x], p, True, rng)[0]
        else:
            total += x if rng.random() >= p else 0.0      # "plain" dropout: no rescaling
    return total / trials

rng = random.Random(7)
x, p = 2.0, 0.3
print(round(mean_output(x, p, 100_000, rng), 3))                  # 2.004
print(round(mean_output(x, p, 100_000, rng, inverted=False), 3))  # 1.399
print(dropout([x], p, False, rng)[0])                             # 2.0
python

With rescaling, the average training-time output is 2.0 (within sampling noise), matching the eval-time output. Without it, training sees an average of 1.4, which is (1 - p) times smaller, so a network trained this way must have its activations multiplied by (1 - p) at inference instead. The original 2014 paper did exactly that; inverted dropout moves the correction into training so inference code needs no special case.

Step 3: Weight decay, L2 versus decoupled

Weight decay shrinks weights toward zero every step. With plain SGD, adding an L2 penalty wd * w to the gradient and shrinking the weight directly are the same update. With Adam they differ, because Adam divides the whole gradient, penalty included, by a running estimate of its magnitude. This simulation gives two weights zero-mean data gradients of very different sizes:

import math

def adam_run(grad_scale, wd, decoupled, steps=1000, lr=0.01, b1=0.9, b2=0.999, eps=1e-8):
    w, m, v = 1.0, 0.0, 0.0
    for t in range(1, steps + 1):
        g = grad_scale * (1 if t % 2 else -1)   # data gradient that averages to zero
        if decoupled:
            w -= lr * wd * w                     # AdamW: shrink the weight directly
        else:
            g += wd * w                          # Adam + L2: the penalty joins the gradient
        m = b1 * m + (1 - b1) * g
        v = b2 * v + (1 - b2) * g * g
        m_hat, v_hat = m / (1 - b1 ** t), v / (1 - b2 ** t)
        w -= lr * m_hat / (math.sqrt(v_hat) + eps)
    return w

for scale in (10.0, 0.01):
    l2 = adam_run(scale, 0.1, decoupled=False)
    adamw = adam_run(scale, 0.1, decoupled=True)
    print(scale, round(l2, 3), round(adamw, 3))
# 10.0 0.889 0.361
# 0.01 0.0 0.361
python

Under L2, the weight with large gradients barely decays (0.889) and the one with tiny gradients is driven to zero, because the penalty gets divided by each weight’s gradient scale. Decoupled decay shrinks both by the same factor. That is the argument of Loshchilov and Hutter’s AdamW paper, and why torch.optim.AdamW is the usual choice for transformers. A common refinement is to exclude biases and normalization parameters from decay (PyTorch 2.x, illustrative):

decay = [p for p in model.parameters() if p.ndim > 1]       # weight matrices and kernels
no_decay = [p for p in model.parameters() if p.ndim <= 1]   # biases, norm gammas and betas
optimizer = torch.optim.AdamW(
    [{"params": decay, "weight_decay": 0.05}, {"params": no_decay, "weight_decay": 0.0}],
    lr=3e-4,
)
python

Step 4: Early stopping with patience

Validation loss is noisy, so stopping at the first uptick is a mistake. Track the best value, allow a few non-improving epochs, and restore the best checkpoint:

def synthetic_val_loss(epoch, rng):
    # falls quickly, bottoms out near epoch 12, then rises as the model overfits; plus noise
    return 0.3 + 1.5 * math.exp(-epoch / 3) + 0.004 * max(0, epoch - 12) ** 1.5 + rng.gauss(0, 0.02)

def train_with_early_stopping(max_epochs=60, patience=5, min_delta=1e-3, seed=6):
    rng = random.Random(seed)
    best_loss, best_epoch, bad_epochs = float("inf"), -1, 0
    for epoch in range(max_epochs):
        loss = synthetic_val_loss(epoch, rng)      # stands in for train_one_epoch() + evaluate()
        if loss < best_loss - min_delta:
            best_loss, best_epoch, bad_epochs = loss, epoch, 0   # save a checkpoint here
        else:
            bad_epochs += 1
            if bad_epochs >= patience:
                return epoch, best_epoch, round(best_loss, 4)
    return max_epochs - 1, best_epoch, round(best_loss, 4)

print(train_with_early_stopping())             # (19, 14, 0.296)
print(train_with_early_stopping(patience=1))   # (9, 8, 0.3803)
python

With patience 5 the loop stops at epoch 19 and restores epoch 14. With patience 1, a single noisy bump at epoch 9 ends training while the loss is still falling, leaving a worse model. Because the best epoch is chosen on the validation set, its score is slightly optimistic; report final numbers on a separate test set.

Data augmentation works from the other side: it enlarges the effective training set with label-preserving transforms, applied to training data only (illustrative, torchvision 0.16 or newer):

from torchvision.transforms import v2
train_tf = v2.Compose([v2.ToImage(), v2.RandomResizedCrop(224), v2.RandomHorizontalFlip(),
                       v2.ToDtype(torch.float32, scale=True)])
python

Step 5: Dropout and batch norm together

Inverted dropout keeps the mean but not the variance. A batch norm layer placed after dropout learns its running variance from the noisier training activations, then sees calmer ones at eval:

def variance(xs):
    m = sum(xs) / len(xs)
    return sum((x - m) ** 2 for x in xs) / len(xs)

rng = random.Random(5)
acts = [rng.gauss(0.0, 1.0) for _ in range(50_000)]   # activations entering a batch norm layer
print(round(variance(dropout(acts, 0.5, True, rng)), 2))    # 1.99 what running_var learns
print(round(variance(dropout(acts, 0.5, False, rng)), 2))   # 1.0 what it sees at eval time
python

At eval the layer divides by a standard deviation about 1.4 times too large. Li et al. (2019) called this “variance shift”. The usual fix is architectural: in convolutional nets with batch norm, use little or no dropout in the conv blocks and put it after the last normalization layer, typically in the classifier head.

Worked scenario

A team replaces nn.Dropout(0.5) with a hand-written layer to log which units fire. Validation accuracy holds up, but the model’s probabilities become badly overconfident. The custom layer zeroes units in training and returns the input unchanged in eval, but never rescales:

def plain_dropout(xs, p, training, rng):          # the hand-rolled layer: no rescaling
    if not training:
        return list(xs)
    return [x if rng.random() >= p else 0.0 for x in xs]

def pre_activation(layer, training, rng, trials=2000):
    h, w = [1.0] * 100, [0.02] * 100              # 100 hidden units feeding one output neuron
    total = 0.0
    for _ in range(trials):
        total += sum(wi * hi for wi, hi in zip(w, layer(h, 0.5, training, rng)))
    return total / trials

rng = random.Random(11)
for layer in (plain_dropout, dropout):
    train_z = pre_activation(layer, True, rng)
    eval_z = pre_activation(layer, False, rng, trials=1)
    print(layer.__name__, round(train_z, 3), round(eval_z, 3))
# plain_dropout 1.0 2.0
# dropout 2.001 2.0
python

The next layer learned to expect an input around 1.0 and receives 2.0 at inference, so logits double and the softmax sharpens. The ranking of classes barely changes, which is why accuracy hid the bug while calibration metrics exposed it. The fix is to divide survivors by (1 - p), or simply keep nn.Dropout, which also follows model.train() and model.eval() automatically.

Common mistake

  • “Dropout is applied at inference too.” Only deliberately, as in Monte Carlo dropout for uncertainty estimates. Normal inference uses model.eval(), where dropout is the identity.
  • “Validation loss below training loss means a bug.” Dropout and augmentation are active only during training, and training loss is often averaged over an epoch while the model improves, so this is common early on.
  • “L2 regularization and weight decay are the same thing.” True for SGD, not for adaptive optimizers like Adam.
  • “Add dropout whenever the model overfits.” Check the data first: leakage, duplicates between splits or a tiny dataset call for different fixes.

Verify the behavior

Append these assertions to the snippets above and run the file with python3:

def test_dropout_is_identity_at_eval():
    assert dropout(h, 0.5, False, random.Random(0)) == h

def test_inverted_dropout_preserves_the_mean():
    assert abs(mean_output(2.0, 0.3, 100_000, random.Random(1)) - 2.0) < 0.02

def test_survivors_are_scaled():
    out = dropout([1.0] * 1000, 0.2, True, random.Random(2))
    assert set(out) <= {0.0, 1.25}

def test_early_stopping_keeps_the_best_epoch():
    stopped, best, _ = train_with_early_stopping()
    assert best < stopped and stopped - best == 5

for test in (test_dropout_is_identity_at_eval, test_inverted_dropout_preserves_the_mean,
             test_survivors_are_scaled, test_early_stopping_keeps_the_best_epoch):
    test()
print("ok")
python

In PyTorch, nn.Dropout(0.2) on torch.ones(1000) should return only 0.0 and 1.25 in training mode and the input unchanged after .eval().

Follow-up questions

Why can huge networks generalize at all? Zhang et al. (2017) showed standard networks can memorize randomly labelled data, so parameter count alone does not explain generalization. Belkin et al. (2019) and Nakkiran et al. (2019) described double descent: as model size grows, test error can fall, rise near the point where the model just fits the training data, then fall again. It is clearest with label noise and weak regularization, and it does not mean bigger is always better. Implicit regularization from SGD and architectural inductive biases are part of the explanation.

What dropout rate would you start with? Around 0.1 in transformer blocks, 0.2 to 0.5 in large fully connected layers, and often zero when pretraining very large language models on abundant data. Tune it on validation loss.

How is dropout different in RNNs? Dropping different units at every time step disrupts the recurrent state, so variational dropout reuses one mask across time steps.

Interview exercise

A ResNet-style classifier with batch norm reaches 99 percent training accuracy and 72 percent validation accuracy on 8,000 labelled images. A colleague adds nn.Dropout(0.5) after every convolution, uses Adam with weight_decay=1e-4, and trains for 200 epochs. Validation accuracy drops to 68 percent. Explain why and propose a better plan.

Answer and reasoning

First confirm the diagnosis: the curves should show training loss still falling while validation loss rises, which is overfitting on a small dataset. Dropout after every convolution is the wrong tool here. Each dropout layer feeds a batch norm layer, so Step 5’s variance shift hits every block at eval, and heavy dropout in conv layers also slows learning without matching gains. torch.optim.Adam with weight_decay applies an L2 penalty, which Step 3 showed is weakened for weights with large gradients. Two hundred epochs without early stopping runs far past the best epoch. A better plan: remove dropout from the conv blocks and keep at most one dropout layer in the head; switch to AdamW with a meaningful decay (for example 0.05) that skips biases and norm parameters; add augmentation such as random resized crops and flips; use early stopping with patience and restore the best checkpoint; and, if possible, start from a pretrained backbone, which usually helps most with 8,000 images.

Continue learning

More in Deep Learning

esc