Ch. 28 · Machine Learning

Gradient Descent Explained: Batch vs SGD vs Mini-Batch

How gradient descent updates weights, how batch, stochastic and mini-batch differ, and why learning rate and feature scaling decide convergence.

~9 min readintermediateupdated Oct 6, 2026

“Explain gradient descent” sounds like a warm-up, and interviewers use it exactly that way: the definition takes thirty seconds, and the real questions follow. Why does training loss jump to NaN after a few steps? Why is stochastic gradient descent noisy, and why does everyone use mini-batches? Why does the same learning rate work on one dataset and explode on another? Each has a mechanical answer, and giving it shows you have trained models rather than called .fit().

Before you start

You need the idea of a derivative as a slope: if increasing a weight increases the loss, the derivative is positive, so you should decrease the weight. You should also know linear regression and mean squared error (MSE). All code here is plain Python 3.14 with only the standard library, so you can run it anywhere; the printed values come from real runs with fixed seeds.

The short answer

Gradient descent minimises a loss by repeatedly stepping against its gradient: w = w - lr * grad. Batch gradient descent computes the gradient on the whole dataset, giving smooth but expensive steps. Stochastic gradient descent (SGD) uses one example per step, which is cheap but noisy. Mini-batch uses a few dozen to a few thousand rows, the practical default because it uses vectorised hardware well and averages away most of the noise. The learning rate controls step size: too small and training crawls, too large and the updates overshoot and diverge. Scaling features and using a decaying schedule or an adaptive optimizer such as Adam make the choice far less fragile.

How it works

Take the simplest possible loss, f(x) = x², whose minimum is at 0 and whose gradient is 2x. One update is x = x - lr * 2x = (1 - 2 * lr) * x, so the learning rate alone decides what happens:

def descend(lr, x=1.0, steps=5):
    out = []
    for _ in range(steps):
        x = x - lr * 2 * x          # gradient of x**2 is 2x
        out.append(round(x, 4))
    return out

for lr in (0.01, 0.1, 0.5, 0.9, 1.1):
    print(lr, descend(lr))
# 0.01 [0.98, 0.9604, 0.9412, 0.9224, 0.9039]
# 0.1 [0.8, 0.64, 0.512, 0.4096, 0.3277]
# 0.5 [0.0, 0.0, 0.0, 0.0, 0.0]
# 0.9 [-0.8, 0.64, -0.512, 0.4096, -0.3277]
# 1.1 [-1.2, 1.44, -1.728, 2.0736, -2.4883]
python

At 0.01 progress is real but slow. At 0.1 it shrinks by 20% per step. At 0.5 it lands on the minimum in one step, because that rate exactly matches the curvature. At 0.9 it overshoots to the other side every time but still converges, which shows up as an oscillating loss. At 1.1 every step overshoots by more than it corrects, and x grows by 20% per step: divergence.

The general rule comes from curvature. If the steepest direction of the loss has second derivative L, the iteration is stable only when lr < 2 / L. For x², L = 2, so the boundary is 1.0. For MSE on a dataset, the curvatures are determined by the feature values themselves, which is why the “right” learning rate is a property of the data, not of the algorithm.

Step-by-step walkthrough

We fit y = 3·x1 - 2·x2 + 5 + noise on 1,000 rows where both features are drawn from a standard normal distribution.

Step 1: Write the loss and its gradient

For MSE, the gradient with respect to each weight is the average of 2 * error * feature, and for the bias it is the average of 2 * error:

import random

random.seed(0)
X = [[random.gauss(0, 1), random.gauss(0, 1)] for _ in range(1000)]
y = [3 * a - 2 * b + 5 + random.gauss(0, 0.5) for a, b in X]

def predict(w, b, x):
    return w[0] * x[0] + w[1] * x[1] + b

def mse_and_grad(w, b, rows, targets):
    n = len(rows)
    loss, gw0, gw1, gb = 0.0, 0.0, 0.0, 0.0
    for x, t in zip(rows, targets):
        err = predict(w, b, x) - t
        loss += err * err
        gw0 += 2 * err * x[0]
        gw1 += 2 * err * x[1]
        gb += 2 * err
    return loss / n, [gw0 / n, gw1 / n], gb / n
python

The function takes any subset of rows. Batch passes all rows, SGD passes one, mini-batch passes a slice.

Step 2: Run batch gradient descent

def batch_gd(lr, epochs):
    w, b = [0.0, 0.0], 0.0
    for epoch in range(epochs):
        loss, gw, gb = mse_and_grad(w, b, X, y)
        w = [w[0] - lr * gw[0], w[1] - lr * gw[1]]
        b = b - lr * gb
    return w, b, mse_and_grad(w, b, X, y)[0]

for epochs in (1, 10, 50, 200):
    w, b, loss = batch_gd(0.1, epochs)
    print(epochs, [round(v, 3) for v in w], round(b, 3), round(loss, 4))
# 1 [0.644, -0.432] 1.008 24.7969
# 10 [2.717, -1.822] 4.467 0.6293
# 50 [2.974, -1.997] 4.992 0.2464
# 200 [2.974, -1.997] 4.992 0.2464
python

After 50 full passes the weights are close to the true 3, -2 and 5, and the loss sits at 0.2464, roughly the noise variance of 0.25. More epochs change nothing because the model has reached the least-squares solution.

Step 3: Compare SGD and mini-batch at the same budget

Give each variant five epochs, shuffling the rows every epoch:

def minibatch_gd(lr, epochs, batch_size, seed=0):
    rng = random.Random(seed)
    w, b = [0.0, 0.0], 0.0
    idx = list(range(len(X)))
    updates = 0
    for epoch in range(epochs):
        rng.shuffle(idx)
        for start in range(0, len(idx), batch_size):
            rows = [X[i] for i in idx[start:start + batch_size]]
            ts = [y[i] for i in idx[start:start + batch_size]]
            _, gw, gb = mse_and_grad(w, b, rows, ts)
            w = [w[0] - lr * gw[0], w[1] - lr * gw[1]]
            b = b - lr * gb
            updates += 1
    return w, b, mse_and_grad(w, b, X, y)[0], updates

# batch_size=1 (lr 0.05):  loss 0.3718 after 5,000 updates
# batch_size=32 (lr 0.1):  loss 0.2509 after 160 updates
# batch_size=1000 (lr 0.1): loss 4.0977 after 5 updates
python

Full batch made only five updates and is nowhere near done. Pure SGD made 5,000 updates and got close quickly, but its final loss is stuck well above the optimum because each one-row gradient points in a slightly wrong direction, so the weights keep jittering around the minimum. Mini-batches of 32 combine the two: enough updates to make progress and enough averaging to land close. On a GPU, a batch of 32 also costs about the same wall time as a batch of 1, which is the real reason mini-batch is the default.

Step 4: Fix the SGD noise floor with a schedule

The jitter is proportional to the learning rate, so shrinking the rate over time lets SGD settle. A simple inverse decay is lr_t = lr0 / (1 + decay * t), where t counts updates:

# same SGD loop, lr = 0.05 / (1 + decay * t)
# decay=0      -> loss 0.3718
# decay=0.001  -> loss 0.2551
# decay=0.01   -> loss 0.2465
python

With decay 0.01, five epochs of SGD reach 0.2465, essentially the batch optimum. Deep learning uses the same idea with warmup plus cosine or step decay, and adaptive optimizers such as Adam rescale each parameter’s step by a running estimate of its gradient magnitude.

Worked scenario

A data scientist predicts house prices (in thousands of dollars) from floor area in square feet (around 1,500) and number of bedrooms (1 to 5). The same training loop that worked above produces NaN with a learning rate of 0.001, and still with 1e-6. At 1e-7 it finally runs, but after 1,000 epochs the result is wrong:

# raw features, lr=1e-3 -> loss inf after 4 epochs
# raw features, lr=1e-6 -> loss inf after 21 epochs
# raw features, lr=1e-7 -> weights [0.259, 0.01], bias 0.0, loss 1360.4
python

The true model is 0.2·size + 15·bedrooms + 50. The curvature along the size weight is about 2 × mean(size²), roughly 4.8 million, so any learning rate above about 4e-7 diverges. But the bedrooms weight and the bias have curvatures thousands of times smaller, and at 1e-7 they barely move. The model compensates by inflating the size weight to 0.259, absorbing the intercept, and stalls at an MSE of 1,360 when the noise alone is 400. This is an ill-conditioned problem: one learning rate cannot suit directions whose curvature differs by orders of magnitude.

The fix is to standardise each feature to mean 0 and standard deviation 1 before training, then convert the weights back if you need them in original units:

# standardized features, lr=0.1, 100 epochs
# weights [78.473, 22.162], bias 401.44, loss 389.9
# back to original units: 0.201 per sq ft, 15.49 per bedroom, intercept 46.0
python

The loss reaches the noise level in 100 epochs at a learning rate a million times larger, and the recovered coefficients are close to the truth. In scikit-learn the equivalent is putting StandardScaler before SGDRegressor in a Pipeline, so the scaling statistics are learned on training data only.

Common mistake

  • “Gradient descent always finds the global minimum.” Only for convex losses such as MSE in linear regression or log loss in logistic regression. Neural network losses are non-convex; in high dimensions the bigger obstacles are saddle points and flat plateaus rather than bad local minima.
  • Reaching for a smaller learning rate when loss is NaN without checking the data. Divergence on a reasonable rate usually means unscaled features, a bug in the gradient, or extreme outliers.
  • Calling every mini-batch method “SGD” and then claiming it uses one row. In PyTorch, torch.optim.SGD is applied to whatever batch you feed it; the name refers to the stochastic gradient estimate, not batch size 1.
  • Judging convergence from a noisy SGD loss curve per step. Average over an epoch, or evaluate on a validation set at fixed intervals.
  • Forgetting to shuffle. If rows are sorted by target or time, consecutive mini-batches are biased and the weights swing with each block.

Verify the behavior

Two checks catch most gradient-descent bugs; save the Step 1 and Step 2 code as linreg.py to run them. A finite-difference test confirms the analytic gradient, and comparing the converged weights with the exact least-squares solution confirms the loop:

from linreg import X, y, mse_and_grad, batch_gd

w, b, h = [0.7, -0.3], 1.5, 1e-6
_, gw, gb = mse_and_grad(w, b, X, y)
num_w0 = (mse_and_grad([w[0] + h, w[1]], b, X, y)[0]
          - mse_and_grad([w[0] - h, w[1]], b, X, y)[0]) / (2 * h)
assert abs(num_w0 - gw[0]) < 1e-5
print(round(gw[0], 6), round(num_w0, 6))       # -4.93326 -4.93326

# Normal equations (X^T X) beta = X^T y, solved with 3x3 Gaussian elimination
# -> weights [2.974, -1.997], bias 4.992, within 1e-6 of batch_gd(0.1, 200)
python

The same check is still the quickest way to validate a custom loss.

Follow-up questions

  • What does momentum add? It keeps a running average of past gradients and steps along it, which damps oscillation across a narrow valley and speeds progress along it.
  • Why does Adam often need less tuning? It divides each parameter’s step by the square root of a running average of its squared gradients, so poorly scaled directions get comparable step sizes. See Adam vs SGD for when plain SGD still generalises better.
  • How do you pick a batch size? Start with the largest that fits in memory among common values (32 to 512), and scale the learning rate with it; very large batches often need warmup.
  • Why not just use the closed-form solution? It exists for linear regression but costs roughly cubic time in the number of features and needs all data in memory; most models, including logistic regression, have no closed form at all.

Interview exercise

You minimise f(w) = 5w² with plain gradient descent starting from w = 1. For which learning rates does it converge, which rate converges fastest, and what does the loss curve look like at lr = 0.15?

Answer and reasoning

The gradient is 10w, so one step gives w = w - lr · 10w = (1 - 10·lr) · w. The iterate shrinks only when |1 - 10·lr| < 1, which means 0 < lr < 0.2; that matches the rule lr < 2 / L with curvature L = 10. At exactly 0.2 it bounces between 1 and -1 forever, and above it diverges. The fastest rate is lr = 0.1, where the factor is zero and the minimum is reached in one step. At lr = 0.15 the factor is -0.5, so w goes 1, -0.5, 0.25, -0.125: the sign flips every step but the loss 5w² still falls by a factor of four each time.

Continue learning

More in Machine Learning

esc