SGD, Momentum, Adam and AdamW: Optimizer Interview Guide
How SGD, momentum, Adam and AdamW update weights, why AdamW decouples weight decay, and how warmup and cosine decay keep training stable.
Deep learning interviews: neural networks, backpropagation, activations and initialization, optimizers, normalization, CNNs, RNNs, attention and transformers, and training at scale.
Official reference: Dive into Deep Learning
How SGD, momentum, Adam and AdamW update weights, why AdamW decouples weight decay, and how warmup and cosine decay keep training stable.
How backpropagation applies the chain rule layer by layer, worked by hand on a tiny network with softmax cross-entropy and a gradient check.
How batch norm and layer norm compute their statistics, why model.eval() matters, why small batches hurt, and why transformers use LayerNorm.
How a conv layer slides a kernel, the output size and parameter formulas, pooling, and how receptive field grows, with runnable Python.
How to spot overfitting in loss curves, why inverted dropout scales by 1/(1-p), and how weight decay, augmentation and early stopping help.
How RNNs, LSTMs and GRUs process sequences, why gradients vanish through time, how gates fix it, and how to count each model's parameters.
Compute softmax(QK^T/sqrt(d_k))V by hand, see why the scaling matters, then add causal masks, multiple heads and a KV cache memory budget.
Why gradients vanish or explode in deep networks, measured layer by layer, and how He and Xavier init, residuals and clipping fix it.
A perceptron is a single linear unit followed by a step function: it computes w·x + b and outputs a class, so it can only separate data with a straight line (or hyperplane). It famously cannot learn XOR.
An MLP stacks layers of such units with differentiable activations and trains them with backpropagation. The activations are essential. Without them, two linear layers W2(W1 x) collapse into one linear map (W2 W1) x, so depth adds nothing. A nonlinearity between layers lets the network compose simple pieces into curved decision boundaries; with enough hidden units, a single hidden layer can approximate any continuous function on a bounded domain (the universal approximation theorem). In practice depth matters more than width, because deep networks represent hierarchical features with far fewer parameters.
Likely follow-up: Why do deep networks usually beat a single very wide hidden layer?
Sigmoid squashes to (0, 1) and tanh to (-1, 1). Both saturate: for large inputs their derivatives approach zero, and sigmoid's derivative never exceeds 0.25, so gradients shrink through deep stacks. Sigmoid is also not zero-centred. Today sigmoid belongs mainly at outputs (binary or multi-label probabilities) and inside gates, such as an LSTM's.
ReLU, max(0, x), has derivative 1 for positive inputs, is cheap, and made deep networks practical. Its weakness is "dying" units that get stuck with negative inputs and receive no gradient; leaky ReLU keeps a small slope there.
GELU, x * Phi(x), is a smooth ReLU-like curve that weights inputs by how likely they are to be positive. It is the default in BERT and GPT-style transformers, and gated variants such as SwiGLU are common in recent LLMs. For a new CNN or MLP, ReLU is the safe default; for transformers, follow the architecture's choice.
Likely follow-up: Where do you still use sigmoid in a modern network?
Backpropagation computes the gradient of a scalar loss with respect to every parameter by applying the chain rule over the computation graph in reverse.
The forward pass computes the loss and stores intermediate values such as activations. The backward pass starts with dL/dL = 1 and walks the graph backwards: each operation multiplies the incoming gradient by its local derivative and passes the result to its inputs. For a linear layer z = W h, that gives dL/dW as the outer product of dL/dz and h, and sends W^T dL/dz to the previous layer. When a value is used twice, the gradients from both uses are summed.
It is efficient because every intermediate gradient is computed once and reused, so the whole gradient costs only a small multiple of one forward pass. The price is memory: stored activations are why training needs far more memory than inference. Backprop only computes gradients; the optimizer decides how to use them.
Likely follow-up: Why is reverse mode preferred over forward-mode differentiation for training? · How would you check a hand-written backward pass?
The gradient reaching an early layer is a product of one factor per layer: the transposed weight matrix times the activation derivative. If those factors are consistently below 1 the product shrinks exponentially with depth and early layers stop learning; above 1 it grows exponentially, giving huge updates, loss spikes or NaN.
Causes are saturating activations (sigmoid, tanh), badly scaled initial weights and, in RNNs, multiplying by the same recurrent matrix at every time step.
Fixes:
To diagnose, log per-layer gradient norms rather than guessing from the loss curve.
Likely follow-up: Why does a residual connection keep gradients alive?
If every weight in a layer starts equal (zero or any constant), every unit computes the same output and receives the same gradient, so they stay identical forever. Random initialization breaks the symmetry.
The scale matters too. A pre-activation sums fan_in terms, so its variance is roughly fan_in * Var(w) times the input's mean square. Too small and signals shrink layer after layer; too large and they explode, both forwards and backwards.
Xavier (Glorot) sets Var(w) = 2 / (fan_in + fan_out), balancing forward and backward signal for tanh or linear layers. He (Kaiming) sets Var(w) = 2 / fan_in for ReLU: the factor 2 compensates for ReLU zeroing about half its inputs. Biases are usually initialized to zero, which is fine because the weights already break symmetry. PyTorch's default nn.Linear init is a Kaiming-uniform variant with bound 1/sqrt(fan_in).
Likely follow-up: Transformers often use std 0.02. Why does that work for them but not for a deep plain MLP?
Cross-entropy is the negative log-likelihood of the correct class, -log p_target. It penalizes confident wrong answers heavily, and paired with softmax its gradient with respect to the logits is simply p - y, so the gradient stays large while the model is badly wrong. MSE on sigmoid or softmax outputs multiplies by the activation's derivative, which is tiny when outputs saturate, so a confidently wrong model learns slowly. Cross-entropy is also the principled choice under a categorical likelihood.
Frameworks fuse softmax and the loss (nn.CrossEntropyLoss takes raw logits) for numerical stability: they compute log_softmax with the log-sum-exp trick, subtracting the maximum logit so exp never overflows. A common bug is applying softmax yourself and then passing probabilities to a loss that expects logits; it still trains, but poorly. For multi-label problems use sigmoid with binary cross-entropy (BCEWithLogitsLoss), one independent probability per label.
import math
def cross_entropy(logits, target):
m = max(logits) # log-sum-exp trick
log_z = m + math.log(sum(math.exp(z - m) for z in logits))
return log_z - logits[target]
print(round(cross_entropy([2.0, 1.0, 0.1], 0), 4)) # 0.417
print(round(cross_entropy([1000.0, 1001.0], 1), 4)) # 0.3133, no overflowLikely follow-up: What does label smoothing change in this loss?
Match the loss to the output and its likelihood:
Practical points: pass logits to the fused loss functions for stability, scale regression targets so the loss is not dominated by units, and handle class imbalance with class weights, focal loss or resampling. The training loss is not always the business metric; you might train with cross-entropy and choose a decision threshold for precision or recall afterwards.
Likely follow-up: When would focal loss help?
Instead of computing the gradient over the whole dataset (batch gradient descent) or one example (pure SGD), mini-batch SGD averages it over a small batch, typically tens to thousands of examples. That gives a noisy but unbiased estimate of the full gradient at a fraction of the cost, and it maps well onto GPU parallelism.
Batch size trades off several things. Larger batches reduce gradient noise and use hardware more efficiently, so throughput rises, but each epoch has fewer updates, memory grows, and very large batches can generalize slightly worse unless the learning rate and warmup are retuned. A common heuristic is to scale the learning rate roughly linearly with batch size, with warmup, until it stops working. Smaller batches add noise that can act as a regularizer. If the batch you want does not fit in memory, gradient accumulation sums gradients over several micro-batches before one optimizer step.
Likely follow-up: Does gradient accumulation give exactly the same result as a larger batch when the model uses batch norm?
Momentum keeps a velocity that accumulates past gradients: in PyTorch's form, v = mu * v + g and w = w - lr * v, with mu typically 0.9. It acts like an exponential moving average of the gradient direction.
It helps in two ways. In long, narrow valleys, plain SGD zig-zags across the steep direction and crawls along the flat one, because its learning rate is capped by the steepest curvature. With momentum, the oscillating components partly cancel while the consistent component builds up, so progress along the valley is much faster, and the stable learning-rate range is wider. It also smooths mini-batch noise.
Nesterov momentum evaluates the gradient at the look-ahead position, which corrects overshoot earlier; PyTorch implements an equivalent form with nesterov=True. SGD with momentum is still the standard recipe for many vision models and is cheaper in memory than Adam (one state tensor instead of two).
Likely follow-up: What happens if momentum is set too close to 1?
Adam keeps two moving averages per parameter: the mean gradient m (momentum, beta1 = 0.9) and the mean squared gradient v (beta2 = 0.999). Both start at zero, so it applies bias correction, dividing by 1 - beta^t, before updating w -= lr * m_hat / (sqrt(v_hat) + eps). The division normalizes each parameter's step to roughly lr, which makes Adam robust to gradient scale and good for sparse or unevenly scaled gradients.
In PyTorch, Adam(weight_decay=...) adds wd * w to the gradient, which is L2 regularization. Inside Adam that term is also divided by sqrt(v_hat), so weights with large gradients are barely decayed and weights with small gradients are decayed heavily. AdamW decouples the decay: it shrinks the weights directly by lr * wd * w and runs the Adam step on the loss gradient alone. That behaves like true weight decay and generalizes better, which is why AdamW (default weight_decay=0.01) is the standard for transformers.
Likely follow-up: Why do people exclude biases and LayerNorm weights from weight decay? · How much optimizer memory does Adam need per parameter?
The learning rate is usually the most important hyperparameter, and the best value changes during training. Early on, a large rate makes fast progress; later, it keeps the optimizer bouncing around the minimum, so decaying it lets training settle and usually improves the final loss.
Warmup ramps the rate linearly from near zero over the first few hundred or thousand steps. At the start, weights are random, activations and gradients can be badly scaled, and Adam's second-moment estimates are based on very few samples, so full-size steps can cause loss spikes or divergence. Transformers in particular are trained with warmup.
Cosine decay then lowers the rate along half a cosine curve from the peak to a small floor by the final step. It is smooth and has one fewer hyperparameter than step decay. Alternatives include step decay, linear decay (common for fine-tuning), one-cycle schedules and warmup-stable-decay for LLM pretraining. Pick the peak with a short sweep or a learning-rate range test.
Likely follow-up: If you resume training from a checkpoint, what must you restore besides the weights?
Batch norm normalizes each feature (each channel for convolutions) using the mean and variance computed across the current mini-batch, then applies a learned scale gamma and shift beta so the layer can still represent any distribution. It allows higher learning rates, reduces sensitivity to initialization and adds a little regularizing noise.
Because batch statistics depend on which examples happen to share the batch, they cannot be used at inference, where you may have a single example. So during training the layer also keeps running averages of the mean and variance, and in eval mode it uses those fixed values instead. In PyTorch that switch is model.train() versus model.eval(); forgetting eval() makes predictions depend on batch-mates and usually hurts accuracy.
Batch norm struggles with tiny batches (noisy statistics), variable-length sequences and distributed training without synchronized statistics, which is why transformers use layer norm instead.
Likely follow-up: What would you do when fine-tuning a CNN with batch size 2?
Layer norm normalizes each example independently across its feature dimension: for a token vector, subtract the mean of its own features, divide by their standard deviation, then apply learned gamma and beta. No statistics cross the batch, so training and inference behave identically and batch size does not matter.
That suits transformers: sequences have variable length and padding, batches can be small per device, and autoregressive generation processes one token at a time, all of which make batch statistics awkward. Layer norm also works the same whether a token is in a batch of one or a thousand.
Two refinements come up. Pre-LN places the norm before each attention and MLP sublayer, inside the residual branch; it trains more stably than the original post-LN layout and needs less warmup. RMSNorm drops the mean subtraction and beta, scaling only by the root mean square, which is cheaper and is used in many recent LLMs such as Llama.
Likely follow-up: Where exactly does the norm sit in a pre-LN transformer block?
During training, dropout sets each activation to zero with probability p, so the network cannot rely on any single unit and co-adapted features are discouraged. It behaves roughly like training an ensemble of thinned subnetworks that share weights.
At inference all units are active, so the expected activation would be larger than during training. Modern frameworks use inverted dropout: during training they divide the surviving activations by 1 - p, keeping the expected value unchanged, and at evaluation dropout is simply the identity. That is why model.eval() matters: it turns dropout off.
Typical rates are 0.1 in transformers and 0.2 to 0.5 in fully connected layers. Dropout is less common in convolutional layers (spatial dropout drops whole channels instead) and is often disabled entirely when pretraining large language models on huge datasets, where overfitting is not the bottleneck. Combining dropout right before batch norm can cause a train and test mismatch in variance.
Likely follow-up: What is Monte Carlo dropout used for?
Detect it from the curves: training loss keeps falling while validation loss bottoms out and rises, or the train and validation metric gap keeps widening. Check that the validation set is representative and free of leakage before trusting either number.
Remedies, roughly in order of payoff:
Deep networks can memorize random labels, yet large models often generalize well; over-parameterization alone is not the problem, and test error can even fall again past the interpolation point (double descent). So judge by validation behaviour, not parameter count.
Likely follow-up: Your validation loss rises but validation accuracy keeps improving. What is happening?
A convolution slides a small kernel (for example 3x3) across the image and computes a dot product at each position. Two properties make it efficient. Local connectivity: each output looks at a small neighbourhood, matching how nearby pixels are correlated. Weight sharing: the same kernel is used everywhere, so the layer is translation-equivariant and has far fewer parameters than a dense layer, which would need a weight for every input pixel and output unit.
The output size per spatial dimension is floor((n + 2p - k) / s) + 1 for input n, kernel k, padding p and stride s. A 3x3 kernel with padding 1 and stride 1 keeps the size ("same" padding).
Parameters are k * k * C_in * C_out + C_out, independent of the image size. A 3x3 convolution from 64 to 128 channels has 73,856 parameters whether the input is 32x32 or 1024x1024.
def conv_out(n, k, s=1, p=0, d=1):
return (n + 2 * p - d * (k - 1) - 1) // s + 1
print(conv_out(224, 7, s=2, p=3)) # 112
print(conv_out(32, 3, p=1)) # 32 ("same" padding)
print(3 * 3 * 64 * 128 + 128) # 73856 parametersLikely follow-up: What does a 1x1 convolution do?
The receptive field is the region of the input image that can influence one unit. Stacking convolutions grows it: two 3x3 layers see 5x5, three see 7x7, with fewer parameters and more nonlinearities than one 7x7 layer, which is the VGG argument.
Downsampling multiplies the growth. Each layer adds (k - 1) times the product of all previous strides, so after a stride-2 layer or a 2x2 pooling step, every later 3x3 convolution grows the field twice as fast. Dilation spaces the kernel taps apart, enlarging the field without extra parameters or downsampling, which helps in segmentation.
Max pooling keeps the strongest response in each window and gives some tolerance to small shifts; average pooling smooths. Global average pooling at the end collapses each channel to one number, which removes the dependence on input resolution before the classifier. The effective receptive field is usually smaller than the theoretical one, because central pixels contribute most.
Likely follow-up: Why do many modern CNNs replace pooling with strided convolutions?
A residual block computes y = x + F(x) instead of y = F(x). Two things follow.
First, gradient flow: the derivative of y with respect to x is I + dF/dx, so the gradient always has an identity path straight back through every block, regardless of what the layers do. It cannot vanish just because many layer Jacobians are small.
Second, optimization: before ResNet, adding layers to a plain network made even the training error worse, which is a degradation problem, not overfitting. With a skip connection, a block that is not helpful can learn F(x) close to 0 and behave like the identity, so a deeper network is at least as easy to fit as a shallower one.
The same idea is the residual stream in every transformer block. When the shapes differ, the skip path uses a 1x1 convolution or projection. Initializing the last layer of each residual branch near zero makes very deep models start close to the identity.
Likely follow-up: How is a transformer block a residual network?
A vanilla RNN updates a hidden state h_t = tanh(W h_(t-1) + U x_t + b) and is trained with backpropagation through time. The gradient from step t back to step t - k multiplies through the recurrent Jacobian k times, so it vanishes or explodes exponentially and the network struggles to learn long-range dependencies.
An LSTM adds a cell state updated additively: c_t = f_t * c_(t-1) + i_t * g_t. Sigmoid gates decide what to forget, what to write and what to output. When the forget gate is near 1, information and gradient flow along the cell state with little decay, much like a residual connection through time.
A GRU merges the cell and hidden state and uses two gates (update and reset). It has about three quarters of the LSTM's parameters and often performs similarly. Exploding gradients are still handled with gradient clipping, and long sequences with truncated BPTT. For most tasks, transformers have replaced both because they parallelize across time.
Likely follow-up: Where do recurrent models still make sense today?
d_k?midEach token's vector is projected into a query, a key and a value with learned matrices. The attention scores are the dot products of every query with every key, Q K^T, scaled by 1/sqrt(d_k), passed through a softmax over the keys, and used to take a weighted average of the values: softmax(Q K^T / sqrt(d_k)) V. Each output is therefore a mixture of all tokens' values, weighted by relevance, and every token can attend to every other in one step.
The scaling matters because if query and key components have unit variance, their dot product has variance d_k. With d_k = 64 the scores have a standard deviation around 8, the softmax becomes nearly one-hot, and its gradients become tiny. Dividing by sqrt(d_k) brings the variance back to about 1.
In a decoder, a causal mask sets scores for future positions to minus infinity before the softmax, so a token cannot see what comes after it.
import math
q = [1.0, 0.0]
keys = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
values = [[1.0, 0.0], [0.0, 1.0], [0.5, 0.5]]
scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(len(q)) for k in keys]
exps = [math.exp(s - max(scores)) for s in scores]
weights = [e / sum(exps) for e in exps]
out = [sum(w * v[j] for w, v in zip(weights, values)) for j in range(2)]
print([round(w, 4) for w in weights]) # [0.4011, 0.1978, 0.4011]
print([round(x, 4) for x in out]) # [0.6017, 0.3983]Likely follow-up: What is the difference between self-attention and cross-attention?
Multi-head attention splits the model dimension d_model into h heads of size d_model / h. Each head has its own query, key and value projections and computes attention independently; the head outputs are concatenated and passed through an output projection W_O.
The benefit is that different heads can attend to different things at once: one head may track the previous token, another a syntactic dependency, another a long-range reference. A single head produces one softmax distribution per query, so it must average those relationships into one pattern.
It costs essentially nothing extra. The Q, K and V projections together are still d_model x d_model matrices, just reshaped into heads, plus the same output projection, so the parameter count and FLOPs match a single head of full width. Variants change the key and value side: multi-query attention shares one K and V head across all query heads, and grouped-query attention shares each K and V head among a group, which shrinks the KV cache at inference.
Likely follow-up: Why does grouped-query attention mainly help inference rather than training?
Self-attention is permutation-equivariant: it treats its input as a set, so shuffling the tokens just shuffles the outputs. Without position information, "dog bites man" and "man bites dog" look the same. Positional encodings inject order.
Main approaches:
The choice affects how well a model handles sequences longer than it saw in training.
Likely follow-up: Why can a model with learned absolute positions not read sequences longer than its maximum length?
In autoregressive generation, each new token attends to all previous tokens. Recomputing every previous token's keys and values at each step would make generation quadratic. The KV cache stores the keys and values for every layer as tokens are processed, so each step only computes the new token's query, key and value and attends against the cache.
Its size is 2 (K and V) * layers * kv_heads * head_dim * sequence_length * batch * bytes_per_element. For an illustrative 8B-class configuration with 32 layers, 8 KV heads, head dimension 128 and fp16, that is 128 KiB per token, so 32k tokens take 4 GiB per sequence. At long contexts and large batches, the cache, not the weights, limits throughput.
Mitigations: grouped-query or multi-query attention (fewer KV heads), KV cache quantization, paged allocation as in vLLM to avoid fragmentation, sliding-window attention, and prefix caching to share common prompt prefixes across requests.
Likely follow-up: Why is decoding memory-bandwidth bound while prefill is compute bound?
For a sequence of n tokens with model width d, the score matrix Q K^T is n x n, so attention costs O(n^2 d) compute and, naively, O(n^2) memory per head to hold the scores. Doubling the context quadruples that part of the cost, while the MLP layers grow only linearly in n.
Ways to make long context practical:
n x n matrix to GPU memory. Compute is still quadratic, but memory becomes linear and wall-clock time drops a lot because attention is memory-bandwidth bound. PyTorch exposes fused kernels through scaled_dot_product_attention.Likely follow-up: Why does FlashAttention speed things up if it does the same number of FLOPs?
An embedding layer is a learnable lookup table: a matrix with one row per item (token, user, product) and d columns. Looking up an id returns its row, which is mathematically the same as multiplying a one-hot vector by the matrix, but far cheaper. In PyTorch it is nn.Embedding(num_items, d).
The rows start random and are trained by backpropagation along with the rest of the model. Only the rows used in a batch receive gradients, which is why embedding gradients are sparse and why adaptive optimizers like Adam help rare items. Because the model is rewarded for predicting well, items that behave similarly end up with nearby vectors, so the embedding space captures similarity.
In language models the input embedding and the output projection are often tied (they share weights), saving parameters. Embeddings replace one-hot encodings for high-cardinality categorical features, and pretrained embeddings can be reused for search, clustering and recommendation.
Likely follow-up: How would you handle an id that was never seen in training?
Start from a model pretrained on a large dataset, because its early layers learn general features (edges and textures in vision, syntax and semantics in language) that transfer to new tasks.
The choice depends on data size and similarity:
Fine-tuning uses a much smaller learning rate than pretraining (for BERT-style models, around 1e-5 to 5e-5) with warmup, and must reuse the pretrained preprocessing, such as tokenizer and normalization statistics. For batch norm models with small batches, keep the norm layers in eval mode. For very large models, parameter-efficient methods such as LoRA fine-tune a small number of added weights instead.
Likely follow-up: What is catastrophic forgetting and how do you limit it?
LoRA (low-rank adaptation) freezes the pretrained weight W and learns a low-rank update: the layer computes W x + (alpha / r) * B A x, where A is r x d_in, B is d_out x r and the rank r is small, such as 8 or 16. B starts at zero, so training begins exactly at the pretrained model.
For a 4096 by 4096 projection, the full matrix has about 16.8 million parameters, while a rank-8 adapter has 65,536, about 0.4 percent. Only the adapters need gradients and optimizer state, so memory for Adam's moments and gradients falls by orders of magnitude; the frozen weights still have to fit. QLoRA goes further by storing the frozen base model in 4-bit precision while training adapters in higher precision.
After training, B A can be merged into W, so inference has no extra latency, or adapters can be swapped per task on one base model. Typical targets are the attention projections and often the MLP layers; rank and alpha are the main hyperparameters.
def lora_params(d_in, d_out, r):
full = d_in * d_out
lora = r * (d_in + d_out) # A: r x d_in, B: d_out x r
return full, lora, f"{100 * lora / full:.2f}%"
print(lora_params(4096, 4096, 8)) # (16777216, 65536, '0.39%')Likely follow-up: When would LoRA underperform full fine-tuning?
Mixed precision runs most matrix multiplications and activations in a 16-bit format, which halves activation memory and uses tensor cores for much higher throughput, while keeping numerically sensitive parts in float32: the master copy of the weights, optimizer state, reductions and many loss computations. In PyTorch, torch.autocast chooses the precision per operation.
The two 16-bit formats differ in where they spend their bits. fp16 has 5 exponent bits and 10 mantissa bits: good precision but a small range, so small gradients underflow to zero. It therefore needs loss scaling: multiply the loss before backward(), unscale the gradients before the step, and skip steps that overflow. torch.amp.GradScaler does this dynamically. bf16 has 8 exponent bits, the same range as float32, and only 7 mantissa bits. It rarely underflows, so it usually needs no loss scaling, which is why it is the default on hardware that supports it. Its lower precision is fine for training but is a reason to keep accumulations in float32.
Likely follow-up: What does a GradScaler do when it detects an inf gradient?
Training memory has four parts: weights, gradients, optimizer state and activations saved for the backward pass, plus temporary buffers. With Adam and mixed precision, a common estimate is 16 bytes per parameter before activations: 2 for 16-bit weights, 2 for gradients, 4 for float32 master weights and 8 for Adam's two moments. A 7B-parameter model therefore needs about 112 GB for those alone, while inference in bf16 needs about 14 GB for weights plus the KV cache.
Activations scale with batch size, sequence length (quadratically for naive attention), width and depth. Fixes, from cheapest:
Profile first: torch.cuda.max_memory_allocated() shows the peak.
Likely follow-up: Why does activation checkpointing slow training down, and by roughly how much?
I start with the cheapest checks that rule out the most.
C balanced classes, initial cross-entropy should be near ln(C). Far off suggests bad init or a broken loss.optimizer.zero_grad() each step, parameters actually passed to the optimizer, model.train() set, no accidental detach().Change one thing at a time with a fixed seed so each result is attributable.
Likely follow-up: The loss becomes NaN after a few thousand steps. What do you check first?
No questions match that filter.
Prefer multiple choice? All 20 Deep Learning MCQs with answers →