Deep Learning · cheat sheet

Deep Learning

Backprop, init, optimizers, normalization, CNNs, RNNs, attention, LoRA, mixed precision and GPU memory: the deep learning facts interviews probe.

The facts a deep learning round keeps returning to, from “why does this loss sit at 2.30” to “how big is the KV cache”. Framework names are PyTorch 2.x.

Neurons, layers & activations

  • A perceptron is a linear classifier (w·x + b plus a step); it cannot learn XOR. An MLP stacks layers with nonlinear activations.
  • Without activations, stacked linear layers collapse into one: W2(W1 x) = (W2 W1) x.
  • Universal approximation: one wide hidden layer can approximate any continuous function on a bounded domain. Depth makes it cheap in practice.
Activation Formula Range Derivative Use
Sigmoid 1 / (1 + e^-x) (0, 1) max 0.25, saturates binary or multi-label output, gates
Tanh tanh(x) (-1, 1) max 1, saturates RNN hidden state
ReLU max(0, x) [0, inf) 1 or 0 default for CNNs and MLPs; units can “die”
Leaky ReLU max(ax, x), small a (-inf, inf) 1 or a avoids dead units
GELU x * Phi(x) about (-0.17, inf) smooth BERT, GPT-style transformers
SwiGLU Swish(xW) * (xV) gated smooth many recent LLM MLP blocks
Softmax e^z_i / sum e^z sums to 1 p - y with cross-entropy multi-class output

Backpropagation

  • Forward pass computes the loss and stores activations; backward pass applies the chain rule from the loss to every parameter.
  • Linear layer z = W h + b: dL/dW = dL/dz outer h, dL/db = dL/dz, dL/dh = W^T dL/dz.
  • A value used twice receives the sum of both gradients. A residual y = x + F(x) passes the gradient straight to x.
  • Cost: a small multiple of one forward pass. Memory: activations, which is why training needs far more memory than inference.
  • Check a hand-written backward with central differences: relative error around 1e-7 or below in float64 is good (torch.autograd.gradcheck).
loss = criterion(model(x), y)   # forward
optimizer.zero_grad()           # .grad accumulates otherwise
loss.backward()                 # backprop
optimizer.step()                # update
python

Vanishing & exploding gradients, initialization

  • The gradient at layer 1 is a product of one W^T diag(f'(z)) factor per layer: below 1 it vanishes, above 1 it explodes.
  • Never initialize all weights equal: units stay identical (symmetry). Zero biases are fine.
Init Weight variance For
Xavier / Glorot 2 / (fan_in + fan_out) tanh, sigmoid, linear
He / Kaiming 2 / fan_in ReLU family
PyTorch nn.Linear default uniform, bound 1/sqrt(fan_in) general use
  • Fixes: ReLU-family activations, proper init, normalization, residual connections, LSTM/GRU gates, and gradient clipping (exploding only): clip_grad_norm_(model.parameters(), 1.0).
  • Diagnose with per-layer gradient norms, not the loss curve alone.

Loss functions

Task Output Loss (PyTorch)
Regression raw value MSELoss, L1Loss, HuberLoss / SmoothL1Loss
Binary 1 logit BCEWithLogitsLoss
Multi-class C logits CrossEntropyLoss (applies log-softmax itself)
Multi-label C logits BCEWithLogitsLoss, one sigmoid per label
  • Cross-entropy = -log p_target; uniform prediction over C classes gives ln(C) (2.303 for 10).
  • Pass logits, not probabilities; fused losses use the log-sum-exp trick for stability.
  • MSE on saturated sigmoid outputs learns slowly; cross-entropy keeps a large gradient when confidently wrong.
  • label_smoothing=0.1 in CrossEntropyLoss softens one-hot targets.

Optimizers

Optimizer Update State per param Notes
SGD w -= lr * g 0 lr capped by steepest curvature (< 2 / L on a quadratic)
Momentum v = mu*v + g; w -= lr*v 1 mu = 0.9; damps zig-zag, Nesterov option
Adam m, v EMAs, bias-corrected; w -= lr * m_hat / (sqrt(v_hat) + eps) 2 defaults lr=1e-3, betas=(0.9, 0.999), eps=1e-8
AdamW Adam + w -= lr * wd * w (decoupled) 2 default weight_decay=0.01; transformer standard
  • Adam(weight_decay=...) is L2, rescaled by 1/sqrt(v_hat); not the same as AdamW.
  • Adam’s first step is about lr * sign(g) for every parameter, regardless of gradient size.
  • Exclude biases and norm weights from weight decay via parameter groups.
  • Typical learning rates: Adam from scratch 1e-4 to 1e-3; BERT-style fine-tuning 1e-5 to 5e-5; LoRA often 1e-4 to 2e-4.

Learning-rate schedules

  • Warmup: linear ramp over the first few hundred to few thousand steps (or a few percent of training). Protects early steps; standard for transformers.
  • Cosine decay to a floor; linear decay is common for fine-tuning; step decay for classic CNN recipes; one-cycle for fast schedules.
  • Batch size up means learning rate up, roughly linearly with warmup, until it stops working.
  • Resuming training needs weights and optimizer and scheduler state.

Normalization

Batch norm Layer norm RMSNorm
Statistics over the batch (per feature or channel) each example’s features each example’s features, RMS only
Train vs eval different: running mean and var in eval identical identical
Small batches poor fine fine
Typical home CNNs transformers, RNNs many recent LLMs (Llama)
  • Learned gamma (scale) and beta (shift) follow the normalization.
  • Pre-LN (norm inside the residual branch, before attention and MLP) trains more stably than post-LN.
  • Forgetting model.eval() makes batch norm outputs depend on batch-mates and leaves dropout on.

Regularization & overfitting

  • Symptom: train loss falls while validation loss rises; the gap widens.
  • Inverted dropout: zero with probability p, scale survivors by 1 / (1 - p) in training; identity at eval. Typical p: 0.1 in transformers, 0.2 to 0.5 in dense layers.
  • Order of payoff: more data and augmentation, early stopping, weight decay, dropout and label smoothing, transfer learning, a smaller model.
  • Big networks can memorize random labels and still generalize on real data; double descent means parameter count alone does not predict overfitting.

CNNs

  • Local connectivity + weight sharing = few parameters and translation equivariance.
  • Output size: floor((n + 2p - d(k - 1) - 1) / s) + 1; with no dilation, floor((n + 2p - k) / s) + 1.
  • Parameters: k * k * C_in * C_out + C_out, independent of image size. 3x3, 64 to 128 channels: 73,856.
  • Receptive field: r += (k - 1) * product of previous strides. Three 3x3 stride-1 layers: 7x7.
  • 1x1 conv mixes channels (bottlenecks); max pooling keeps the strongest response; global average pooling removes the dependence on input resolution.
  • ResNet: y = x + F(x) fixes the degradation problem (deeper plain nets train worse).
nn.Conv2d(3, 64, kernel_size=3, stride=1, padding=1)   # keeps H and W
nn.AdaptiveAvgPool2d(1)                                # global average pooling
python

RNN, LSTM, GRU

  • RNN: h_t = tanh(W h_(t-1) + U x_t + b), trained with backpropagation through time (BPTT); gradients multiply through the recurrent Jacobian at every step.
  • LSTM: forget, input and output gates plus candidate; additive cell state c_t = f*c_(t-1) + i*g carries long-range gradient.
  • GRU: update and reset gates, no separate cell; about 3/4 of LSTM parameters.
  • PyTorch nn.LSTM params per layer: 4 * (h*x + h*h + 2h) (two bias vectors). nn.LSTM(10, 20): 2,560.
  • Use gradient clipping, truncated BPTT, and pack_padded_sequence or masks for padded batches.

Attention & transformers

  • Attention(Q, K, V) = softmax(Q K^T / sqrt(d_k)) V. Without the scaling, dot-product variance is d_k and the softmax saturates.
  • Causal mask: future scores set to -inf before softmax in decoders.
  • Multi-head: h heads of size d_model / h, concat, output projection; same parameters as one full-width head. MQA/GQA share K and V heads.
  • Positional information: sinusoidal, learned absolute, relative biases (T5, ALiBi), RoPE (most current LLMs).
  • Block (pre-LN): x = x + Attn(LN(x)), then x = x + MLP(LN(x)); MLP is usually 4x wide (or a gated variant).
  • Cost: O(n^2 d) in sequence length; FlashAttention is exact and IO-aware (linear memory). F.scaled_dot_product_attention(q, k, v, is_causal=True).
  • KV cache bytes = 2 * layers * kv_heads * head_dim * seq_len * batch * bytes. 32 layers, 8 KV heads, head dim 128, fp16: 128 KiB per token, 4 GiB per 32k-token sequence.

Embeddings, transfer learning & LoRA

  • nn.Embedding(n, d) is a trainable lookup table (one-hot times matrix); only used rows get gradients. LMs often tie input and output embeddings.
  • Transfer: little similar data, freeze the backbone and train a head; more or different data, fine-tune with a small lr and warmup. Reuse the pretrained preprocessing.
  • LoRA: W x + (alpha / r) B A x, W frozen, B starts at zero. Rank 8 on 4096 by 4096: 65,536 trainable vs 16.8M (0.39%). Mergeable for zero-latency inference.
  • QLoRA: 4-bit frozen base plus LoRA adapters.

Mixed precision & GPU memory

Format Exponent / mantissa bits Range Loss scaling
fp32 8 / 23 about 1e-38 to 3e38 no
fp16 5 / 10 about 6e-8 to 65504 yes (torch.amp.GradScaler)
bf16 8 / 7 same as fp32 usually no
  • torch.autocast runs matmuls in 16-bit; master weights, optimizer state and reductions stay in fp32.
  • Training memory = weights + gradients + optimizer state + activations. Adam with mixed precision: about 16 bytes per parameter before activations (7B params: about 112 GB).
  • Inference: weights (7B in bf16: about 14 GB) plus the KV cache.
  • Out of memory? Smaller micro-batch + gradient accumulation, mixed precision, fused attention, activation checkpointing, LoRA, FSDP / ZeRO sharding. Measure with torch.cuda.max_memory_allocated().

Debugging checklist

  • Initial loss near ln(C)? Look at a preprocessed batch. Can the model overfit one batch?
  • Logits into CrossEntropyLoss, zero_grad() every step, model.train() / model.eval() in the right places.
  • Sweep the learning rate by orders of magnitude; log gradient norms; add warmup and clipping before blaming the architecture.
  • NaN: high lr, missing warmup, fp16 without loss scaling, log(0) or division by zero in a custom loss.

Practise these in the Deep Learning chapter.

esc