Deep Learning MCQs multiple-choice questions with answers & explanations
All 20 Deep Learning quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.
Official reference: Dive into Deep Learning
- 1.easy
Three linear layers are stacked with no activation between them. What does this print, and what does it show?
def layer(x): return 2 * x + 1 # a linear layer with no activation def net(x): return layer(layer(layer(x))) print([net(x) for x in (0, 1, 2)])- A
[7, 15, 23]: the stack is just the line8x + 7 - B
[7, 15, 23]: depth made the function nonlinear - C
[1, 3, 5]: only the last layer counts - D
[3, 9, 27]: layers multiply their inputs
Show answer
Answer: A (
[7, 15, 23]: the stack is just the line8x + 7)Composing linear (affine) maps gives another affine map, here
8x + 7, so the outputs rise by a constant 8. Depth adds expressive power only with a nonlinearity between the layers. - A
- 2.easy
What does this print?
import math s = lambda z: 1 / (1 + math.exp(-z)) print(max(round(s(z) * (1 - s(z)), 4) for z in [-4, -1, 0, 1, 4]))- A
1.0 - B
0.5 - C
0.25 - D
0.1966
Show answer
Answer: C (
0.25)The sigmoid derivative
s(z)(1 - s(z))peaks atz = 0, wheres = 0.5, giving 0.25. Since every sigmoid layer multiplies the gradient by at most 0.25, deep sigmoid stacks suffer from vanishing gradients.0.1966is the value atz = 1, not the maximum. - A
- 3.easy
For
L = relu(w * x), what does this print?def relu(z): return max(0.0, z) x, w = 3.0, -2.0 z = w * x grad_w = (1.0 if z > 0 else 0.0) * x # dL/dw print(relu(z), grad_w)- A
0.0 0.0 - B
0.0 3.0 - C
-6.0 3.0 - D
0.0 -2.0
Show answer
Answer: A (
0.0 0.0)The pre-activation is negative, so ReLU outputs 0 and its local derivative is 0, which blocks the gradient to
wentirely. If that happens for every input, the unit is "dead" and never recovers;3.0would be the gradient only ifzwere positive. - A
- 4.mid
What does this print?
import math def softmax(z): e = [math.exp(v - max(z)) for v in z] return [round(v / sum(e), 4) for v in e] print(softmax([1.0, 2.0, 3.0]) == softmax([101.0, 102.0, 103.0]))- A
True - B
False - CAn
OverflowError
Show answer
Answer: A (
True)Softmax is unchanged when the same constant is added to every logit, because it cancels in the ratio. That property is exactly what the max-subtraction trick relies on, which is also why
exp(103)never has to be computed here and nothing overflows. - A
- 5.mid
This computes the gradient of softmax cross-entropy with respect to the logits. What does it print?
import math logits, target = [0.0, 0.0], 0 p = [math.exp(z) / sum(math.exp(v) for v in logits) for z in logits] grad = [pi - (1.0 if i == target else 0.0) for i, pi in enumerate(p)] print(grad)- A
[0.5, 0.5] - B
[-0.5, 0.5] - C
[-1.0, 0.0] - D
[0.5, -0.5]
Show answer
Answer: B (
[-0.5, 0.5])For softmax followed by cross-entropy the logit gradient is
p - y:[0.5 - 1, 0.5 - 0]. Gradient descent then raises the target logit and lowers the other.[0.5, -0.5]has the signs reversed. - A
- 6.mid
An untrained 10-class classifier outputs equal logits. What does this print?
import math num_classes = 10 logits = [0.0] * num_classes loss = -math.log(math.exp(logits[3]) / sum(math.exp(z) for z in logits)) print(round(loss, 3))- A
0.1 - B
1.0 - C
2.303 - D
10.0
Show answer
Answer: C (
2.303)A uniform prediction gives each class probability 0.1, and
-ln(0.1) = ln(10), about 2.303. Checking that the first-step loss is nearln(C)is a quick sanity check; a loss stuck at that value means the model is still predicting uniformly. - A
- 7.easy
This implements inverted dropout in training mode with
p = 0.5. The mask keeps the first two units. What does it print?import random rng = random.Random(0) p = 0.5 x = [1.0, 1.0, 1.0, 1.0] mask = [1.0 if rng.random() >= p else 0.0 for _ in x] print([xi * m / (1 - p) for xi, m in zip(x, mask)])- A
[1.0, 1.0, 0.0, 0.0] - B
[2.0, 2.0, 0.0, 0.0] - C
[0.5, 0.5, 0.0, 0.0] - D
[1.0, 1.0, 1.0, 1.0]
Show answer
Answer: B (
[2.0, 2.0, 0.0, 0.0])Inverted dropout divides surviving activations by
1 - p, so kept units become 2.0 and the expected value stays 1.0. At evaluation dropout is then the identity, with no rescaling. Without the division you would get[1.0, 1.0, 0.0, 0.0]and a train and test mismatch. - A
- 8.easy
What does this print?
def conv_out(n, k, s=1, p=0): return (n + 2 * p - k) // s + 1 print(conv_out(32, 5), conv_out(224, 3, s=2, p=1))- A
28 112 - B
32 112 - C
28 111 - D
27 224
Show answer
Answer: A (
28 112)A 5x5 kernel with no padding loses 4 pixels:
32 - 5 + 1 = 28. A 3x3 kernel with padding 1 and stride 2 halves the size:(224 + 2 - 3) // 2 + 1 = 112. Keeping 32 would need padding 2. - A
- 9.mid
How many trainable parameters does
nn.Conv2d(3, 16, kernel_size=3)have, for any input image size?c_in, c_out, k = 3, 16, 3 print(k * k * c_in * c_out + c_out)- A
144 - B
432 - C
448 - DIt depends on the image resolution
Show answer
Answer: C (
448)Each of the 16 filters has
3 * 3 * 3 = 27weights plus one bias:16 * 28 = 448. Weight sharing makes the count independent of the image size;432forgets the biases. - A
- 10.mid
What receptive field does a unit have after three stacked 3x3 convolutions with stride 1?
def receptive_field(layers): r, jump = 1, 1 for k, s in layers: r += (k - 1) * jump jump *= s return r print(receptive_field([(3, 1), (3, 1), (3, 1)]))- A
3 - B
7 - C
9 - D
27
Show answer
Answer: B (
7)Each stride-1 3x3 layer adds 2 pixels:
1 + 2 + 2 + 2 = 7. Three 3x3 layers see the same 7x7 region as one 7x7 layer with fewer parameters and more nonlinearities. Strides earlier in the stack would make later layers grow the field faster. - A
- 11.hard
This computes Adam's first update for a parameter whose first gradient is 1000. What does it print?
import math lr, b1, b2, eps = 0.01, 0.9, 0.999, 1e-8 g = 1000.0 m = (1 - b1) * g v = (1 - b2) * g * g m_hat, v_hat = m / (1 - b1), v / (1 - b2) print(round(lr * m_hat / (math.sqrt(v_hat) + eps), 6))- A
10.0 - B
0.01 - C
0.001 - D
1000.0
Show answer
Answer: B (
0.01)After bias correction,
m_hat = gandsqrt(v_hat) = |g|, so the first step is aboutlr * sign(g), regardless of the gradient's size. SGD would steplr * g = 10.0. Adam's scale-invariance is its strength, and the full-size early steps are one reason for warmup. - A
- 12.mid
With PyTorch-style momentum and a constant gradient of 1, what does this print?
mu, v = 0.9, 0.0 for g in [1.0, 1.0, 1.0]: v = mu * v + g print(round(v, 2))- A
1.0 - B
2.71 - C
3.0 - D
0.27
Show answer
Answer: B (
2.71)The velocity goes 1, 1.9, 2.71 and approaches
1 / (1 - 0.9) = 10under a constant gradient. Consistent directions therefore get up to ten times the step, while alternating gradients cancel.0.27would be the dampened(1 - mu)formulation. - A
- 13.mid
A CNN with batch norm scores well in validation during training, but in production its prediction for the same image changes depending on which other images share the request batch. What is the most likely cause?
- AThe model was not switched to
model.eval(), so batch norm uses per-batch statistics - BDropout is still active
- CThe learning rate is too high
- DBatch norm always behaves this way at inference
Show answer
Answer: A (The model was not switched to
model.eval(), so batch norm uses per-batch statistics)In training mode batch norm normalizes with the current batch's mean and variance, so outputs depend on batch-mates.
model.eval()switches it to the stored running statistics. Active dropout would cause random variation even for an identical batch, not batch-dependent outputs. - AThe model was not switched to
- 14.mid
For a transformer activation of shape
(batch, seq_len, d_model), which values doesnn.LayerNorm(d_model)average over to compute one mean?- AAll batch elements at the same position and feature
- BAll tokens in the sequence for one feature
- CEvery value in the tensor
- DThe
d_modelfeatures of a single token
Show answer
Answer: D (The
d_modelfeatures of a single token)Layer norm normalizes each token vector over its own features, so no statistics cross the batch or the sequence. Averaging across the batch for each feature is what batch norm does, which is why layer norm behaves identically in training and inference.
- 15.hard
Queries and keys have 64 unit-variance components. This measures the variance of their dot products. What does it print?
import random rng = random.Random(0) d, trials = 64, 20000 dots = [] for _ in range(trials): q = [rng.gauss(0, 1) for _ in range(d)] k = [rng.gauss(0, 1) for _ in range(d)] dots.append(sum(a * b for a, b in zip(q, k))) var = sum(x * x for x in dots) / trials print(round(var), round(var / d, 1))- A
1 0.0 - B
8 0.1 - C
63 1.0 - D
4096 64.0
Show answer
Answer: C (
63 1.0)A sum of 64 independent unit-variance products has variance about 64, so raw scores have standard deviation about 8 and the softmax saturates. Dividing the scores by
sqrt(d_k) = 8divides the variance by 64, back to about 1, which is the reason for the scaling in attention. - A
- 16.mid
A causal mask is applied before the softmax. What attention weights does the first token get?
import math scores = [[2.0, 1.0, 0.5], [0.3, 1.2, 0.8], [1.0, 1.0, 1.0]] masked = [[s if j <= i else -math.inf for j, s in enumerate(row)] for i, row in enumerate(scores)] e = [math.exp(s) for s in masked[0]] print([round(x / sum(e), 2) for x in e])- A
[0.63, 0.23, 0.14] - B
[1.0, 0.0, 0.0] - C
[0.33, 0.33, 0.33] - D
[nan, nan, nan]
Show answer
Answer: B (
[1.0, 0.0, 0.0])The first token may attend only to itself; masked scores of
-infbecomeexp(-inf) = 0, so all the weight lands on position 0.[0.63, 0.23, 0.14]is the unmasked softmax, which would let the token see the future. NaN appears only if every score in a row is masked. - A
- 17.hard
A toy transformer caches keys and values in fp16. How many bytes does this KV cache use?
layers, kv_heads, head_dim, seq_len, batch, bytes_per = 2, 2, 4, 10, 1, 2 print(2 * layers * kv_heads * head_dim * seq_len * batch * bytes_per)- A
160 - B
320 - C
640 - D
1280
Show answer
Answer: C (
640)The leading 2 counts keys and values:
2 * 2 * 2 * 4 * 10 * 1 * 2 = 640. The cache grows linearly with sequence length and batch, which is why grouped-query attention (fewer KV heads) and cache quantization matter at long contexts.320forgets that both K and V are stored. - A
- 18.mid
A rank-4 LoRA adapter is added to a frozen 1024 by 1024 weight matrix. What does this print?
d_in, d_out, r = 1024, 1024, 4 print(r * (d_in + d_out), d_in * d_out)- A
8192 1048576 - B
4096 1048576 - C
16 1048576 - D
1048576 8192
Show answer
Answer: A (
8192 1048576)LoRA trains
A(r x d_in) andB(d_out x r):4 * 2048 = 8192parameters, under 1 percent of the 1,048,576 frozen ones. Only those need gradients and optimizer state.4096counts just one of the two matrices. - A
- 19.hard
Python's
structformateis IEEE half precision (fp16). What does this print?import struct def to_fp16(x): return struct.unpack('e', struct.pack('e', x))[0] print(to_fp16(1e-8), to_fp16(0.1))- A
1e-08 0.1 - B
1e-08 0.0999755859375 - CAn
OverflowError - D
0.0 0.0999755859375
Show answer
Answer: D (
0.0 0.0999755859375)fp16's smallest positive value is about
6e-8, so1e-8underflows to zero, and 0.1 is rounded to the nearest representable value. Small gradients vanishing this way is why fp16 training uses loss scaling; bf16 has float32's exponent range and rarely needs it. - A
- 20.hard
How many parameters does PyTorch's
nn.LSTM(input_size=10, hidden_size=20)(one layer) have?input_size, hidden_size = 10, 20 # 4 gates, each with W_ih, W_hh, b_ih and b_hh print(4 * (hidden_size * input_size + hidden_size * hidden_size + 2 * hidden_size))- A
640 - B
2480 - C
2560 - D
1920
Show answer
Answer: C (
2560)Each of the four gates has a 20x10 input matrix, a 20x20 recurrent matrix and two 20-element biases in PyTorch:
4 * 640 = 2560. The textbook formula with a single bias per gate gives 2480, and a GRU has three gate blocks instead of four (1920). - A
No questions match these filters. Try a different subtopic or clear the filters.