The facts a deep learning round keeps returning to, from “why does this loss sit at 2.30” to “how big is the KV cache”. Framework names are PyTorch 2.x.
Neurons, layers & activations
- A perceptron is a linear classifier (
w·x + bplus a step); it cannot learn XOR. An MLP stacks layers with nonlinear activations. - Without activations, stacked linear layers collapse into one:
W2(W1 x) = (W2 W1) x. - Universal approximation: one wide hidden layer can approximate any continuous function on a bounded domain. Depth makes it cheap in practice.
| Activation | Formula | Range | Derivative | Use |
|---|---|---|---|---|
| Sigmoid | 1 / (1 + e^-x) |
(0, 1) | max 0.25, saturates | binary or multi-label output, gates |
| Tanh | tanh(x) |
(-1, 1) | max 1, saturates | RNN hidden state |
| ReLU | max(0, x) |
[0, inf) | 1 or 0 | default for CNNs and MLPs; units can “die” |
| Leaky ReLU | max(ax, x), small a |
(-inf, inf) | 1 or a |
avoids dead units |
| GELU | x * Phi(x) |
about (-0.17, inf) | smooth | BERT, GPT-style transformers |
| SwiGLU | Swish(xW) * (xV) |
gated | smooth | many recent LLM MLP blocks |
| Softmax | e^z_i / sum e^z |
sums to 1 | p - y with cross-entropy |
multi-class output |
Backpropagation
- Forward pass computes the loss and stores activations; backward pass applies the chain rule from the loss to every parameter.
- Linear layer
z = W h + b:dL/dW = dL/dzouterh,dL/db = dL/dz,dL/dh = W^T dL/dz. - A value used twice receives the sum of both gradients. A residual
y = x + F(x)passes the gradient straight tox. - Cost: a small multiple of one forward pass. Memory: activations, which is why training needs far more memory than inference.
- Check a hand-written backward with central differences: relative error around
1e-7or below in float64 is good (torch.autograd.gradcheck).
loss = criterion(model(x), y) # forward
optimizer.zero_grad() # .grad accumulates otherwise
loss.backward() # backprop
optimizer.step() # updateVanishing & exploding gradients, initialization
- The gradient at layer 1 is a product of one
W^T diag(f'(z))factor per layer: below 1 it vanishes, above 1 it explodes. - Never initialize all weights equal: units stay identical (symmetry). Zero biases are fine.
| Init | Weight variance | For |
|---|---|---|
| Xavier / Glorot | 2 / (fan_in + fan_out) |
tanh, sigmoid, linear |
| He / Kaiming | 2 / fan_in |
ReLU family |
PyTorch nn.Linear default |
uniform, bound 1/sqrt(fan_in) |
general use |
- Fixes: ReLU-family activations, proper init, normalization, residual connections, LSTM/GRU gates, and gradient clipping (exploding only):
clip_grad_norm_(model.parameters(), 1.0). - Diagnose with per-layer gradient norms, not the loss curve alone.
Loss functions
| Task | Output | Loss (PyTorch) |
|---|---|---|
| Regression | raw value | MSELoss, L1Loss, HuberLoss / SmoothL1Loss |
| Binary | 1 logit | BCEWithLogitsLoss |
| Multi-class | C logits | CrossEntropyLoss (applies log-softmax itself) |
| Multi-label | C logits | BCEWithLogitsLoss, one sigmoid per label |
- Cross-entropy =
-log p_target; uniform prediction over C classes givesln(C)(2.303 for 10). - Pass logits, not probabilities; fused losses use the log-sum-exp trick for stability.
- MSE on saturated sigmoid outputs learns slowly; cross-entropy keeps a large gradient when confidently wrong.
label_smoothing=0.1inCrossEntropyLosssoftens one-hot targets.
Optimizers
| Optimizer | Update | State per param | Notes |
|---|---|---|---|
| SGD | w -= lr * g |
0 | lr capped by steepest curvature (< 2 / L on a quadratic) |
| Momentum | v = mu*v + g; w -= lr*v |
1 | mu = 0.9; damps zig-zag, Nesterov option |
| Adam | m, v EMAs, bias-corrected; w -= lr * m_hat / (sqrt(v_hat) + eps) |
2 | defaults lr=1e-3, betas=(0.9, 0.999), eps=1e-8 |
| AdamW | Adam + w -= lr * wd * w (decoupled) |
2 | default weight_decay=0.01; transformer standard |
Adam(weight_decay=...)is L2, rescaled by1/sqrt(v_hat); not the same as AdamW.- Adam’s first step is about
lr * sign(g)for every parameter, regardless of gradient size. - Exclude biases and norm weights from weight decay via parameter groups.
- Typical learning rates: Adam from scratch
1e-4to1e-3; BERT-style fine-tuning1e-5to5e-5; LoRA often1e-4to2e-4.
Learning-rate schedules
- Warmup: linear ramp over the first few hundred to few thousand steps (or a few percent of training). Protects early steps; standard for transformers.
- Cosine decay to a floor; linear decay is common for fine-tuning; step decay for classic CNN recipes; one-cycle for fast schedules.
- Batch size up means learning rate up, roughly linearly with warmup, until it stops working.
- Resuming training needs weights and optimizer and scheduler state.
Normalization
| Batch norm | Layer norm | RMSNorm | |
|---|---|---|---|
| Statistics over | the batch (per feature or channel) | each example’s features | each example’s features, RMS only |
| Train vs eval | different: running mean and var in eval | identical | identical |
| Small batches | poor | fine | fine |
| Typical home | CNNs | transformers, RNNs | many recent LLMs (Llama) |
- Learned
gamma(scale) andbeta(shift) follow the normalization. - Pre-LN (norm inside the residual branch, before attention and MLP) trains more stably than post-LN.
- Forgetting
model.eval()makes batch norm outputs depend on batch-mates and leaves dropout on.
Regularization & overfitting
- Symptom: train loss falls while validation loss rises; the gap widens.
- Inverted dropout: zero with probability
p, scale survivors by1 / (1 - p)in training; identity at eval. Typicalp: 0.1 in transformers, 0.2 to 0.5 in dense layers. - Order of payoff: more data and augmentation, early stopping, weight decay, dropout and label smoothing, transfer learning, a smaller model.
- Big networks can memorize random labels and still generalize on real data; double descent means parameter count alone does not predict overfitting.
CNNs
- Local connectivity + weight sharing = few parameters and translation equivariance.
- Output size:
floor((n + 2p - d(k - 1) - 1) / s) + 1; with no dilation,floor((n + 2p - k) / s) + 1. - Parameters:
k * k * C_in * C_out + C_out, independent of image size. 3x3, 64 to 128 channels: 73,856. - Receptive field:
r += (k - 1) * product of previous strides. Three 3x3 stride-1 layers: 7x7. - 1x1 conv mixes channels (bottlenecks); max pooling keeps the strongest response; global average pooling removes the dependence on input resolution.
- ResNet:
y = x + F(x)fixes the degradation problem (deeper plain nets train worse).
nn.Conv2d(3, 64, kernel_size=3, stride=1, padding=1) # keeps H and W
nn.AdaptiveAvgPool2d(1) # global average poolingRNN, LSTM, GRU
- RNN:
h_t = tanh(W h_(t-1) + U x_t + b), trained with backpropagation through time (BPTT); gradients multiply through the recurrent Jacobian at every step. - LSTM: forget, input and output gates plus candidate; additive cell state
c_t = f*c_(t-1) + i*gcarries long-range gradient. - GRU: update and reset gates, no separate cell; about 3/4 of LSTM parameters.
- PyTorch
nn.LSTMparams per layer:4 * (h*x + h*h + 2h)(two bias vectors).nn.LSTM(10, 20): 2,560. - Use gradient clipping, truncated BPTT, and
pack_padded_sequenceor masks for padded batches.
Attention & transformers
Attention(Q, K, V) = softmax(Q K^T / sqrt(d_k)) V. Without the scaling, dot-product variance isd_kand the softmax saturates.- Causal mask: future scores set to
-infbefore softmax in decoders. - Multi-head:
hheads of sized_model / h, concat, output projection; same parameters as one full-width head. MQA/GQA share K and V heads. - Positional information: sinusoidal, learned absolute, relative biases (T5, ALiBi), RoPE (most current LLMs).
- Block (pre-LN):
x = x + Attn(LN(x)), thenx = x + MLP(LN(x)); MLP is usually 4x wide (or a gated variant). - Cost:
O(n^2 d)in sequence length; FlashAttention is exact and IO-aware (linear memory).F.scaled_dot_product_attention(q, k, v, is_causal=True). - KV cache bytes =
2 * layers * kv_heads * head_dim * seq_len * batch * bytes. 32 layers, 8 KV heads, head dim 128, fp16: 128 KiB per token, 4 GiB per 32k-token sequence.
Embeddings, transfer learning & LoRA
nn.Embedding(n, d)is a trainable lookup table (one-hot times matrix); only used rows get gradients. LMs often tie input and output embeddings.- Transfer: little similar data, freeze the backbone and train a head; more or different data, fine-tune with a small lr and warmup. Reuse the pretrained preprocessing.
- LoRA:
W x + (alpha / r) B A x,Wfrozen,Bstarts at zero. Rank 8 on 4096 by 4096: 65,536 trainable vs 16.8M (0.39%). Mergeable for zero-latency inference. - QLoRA: 4-bit frozen base plus LoRA adapters.
Mixed precision & GPU memory
| Format | Exponent / mantissa bits | Range | Loss scaling |
|---|---|---|---|
| fp32 | 8 / 23 | about 1e-38 to 3e38 |
no |
| fp16 | 5 / 10 | about 6e-8 to 65504 |
yes (torch.amp.GradScaler) |
| bf16 | 8 / 7 | same as fp32 | usually no |
torch.autocastruns matmuls in 16-bit; master weights, optimizer state and reductions stay in fp32.- Training memory = weights + gradients + optimizer state + activations. Adam with mixed precision: about 16 bytes per parameter before activations (7B params: about 112 GB).
- Inference: weights (7B in bf16: about 14 GB) plus the KV cache.
- Out of memory? Smaller micro-batch + gradient accumulation, mixed precision, fused attention, activation checkpointing, LoRA, FSDP / ZeRO sharding. Measure with
torch.cuda.max_memory_allocated().
Debugging checklist
- Initial loss near
ln(C)? Look at a preprocessed batch. Can the model overfit one batch? - Logits into
CrossEntropyLoss,zero_grad()every step,model.train()/model.eval()in the right places. - Sweep the learning rate by orders of magnitude; log gradient norms; add warmup and clipping before blaming the architecture.
- NaN: high lr, missing warmup, fp16 without loss scaling,
log(0)or division by zero in a custom loss.
Practise these in the Deep Learning chapter.