pencils ready ✎

Deep Learning MCQs multiple-choice questions with answers & explanations

All 20 Deep Learning quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.

20 questions
  1. 1.

    Three linear layers are stacked with no activation between them. What does this print, and what does it show?

    easy
    def layer(x):
        return 2 * x + 1          # a linear layer with no activation
    
    def net(x):
        return layer(layer(layer(x)))
    
    print([net(x) for x in (0, 1, 2)])
    1. A[7, 15, 23]: the stack is just the line 8x + 7
    2. B[7, 15, 23]: depth made the function nonlinear
    3. C[1, 3, 5]: only the last layer counts
    4. D[3, 9, 27]: layers multiply their inputs
    Show answer

    Answer: A ([7, 15, 23]: the stack is just the line 8x + 7)

    Composing linear (affine) maps gives another affine map, here 8x + 7, so the outputs rise by a constant 8. Depth adds expressive power only with a nonlinearity between the layers.

  2. 2.

    What does this print?

    easy
    import math
    s = lambda z: 1 / (1 + math.exp(-z))
    print(max(round(s(z) * (1 - s(z)), 4) for z in [-4, -1, 0, 1, 4]))
    1. A1.0
    2. B0.5
    3. C0.25
    4. D0.1966
    Show answer

    Answer: C (0.25)

    The sigmoid derivative s(z)(1 - s(z)) peaks at z = 0, where s = 0.5, giving 0.25. Since every sigmoid layer multiplies the gradient by at most 0.25, deep sigmoid stacks suffer from vanishing gradients. 0.1966 is the value at z = 1, not the maximum.

  3. 3.

    For L = relu(w * x), what does this print?

    easy
    def relu(z):
        return max(0.0, z)
    
    x, w = 3.0, -2.0
    z = w * x
    grad_w = (1.0 if z > 0 else 0.0) * x     # dL/dw
    print(relu(z), grad_w)
    1. A0.0 0.0
    2. B0.0 3.0
    3. C-6.0 3.0
    4. D0.0 -2.0
    Show answer

    Answer: A (0.0 0.0)

    The pre-activation is negative, so ReLU outputs 0 and its local derivative is 0, which blocks the gradient to w entirely. If that happens for every input, the unit is "dead" and never recovers; 3.0 would be the gradient only if z were positive.

  4. 4.

    What does this print?

    mid
    import math
    def softmax(z):
        e = [math.exp(v - max(z)) for v in z]
        return [round(v / sum(e), 4) for v in e]
    print(softmax([1.0, 2.0, 3.0]) == softmax([101.0, 102.0, 103.0]))
    1. ATrue
    2. BFalse
    3. CAn OverflowError
    Show answer

    Answer: A (True)

    Softmax is unchanged when the same constant is added to every logit, because it cancels in the ratio. That property is exactly what the max-subtraction trick relies on, which is also why exp(103) never has to be computed here and nothing overflows.

  5. 5.

    This computes the gradient of softmax cross-entropy with respect to the logits. What does it print?

    mid
    import math
    logits, target = [0.0, 0.0], 0
    p = [math.exp(z) / sum(math.exp(v) for v in logits) for z in logits]
    grad = [pi - (1.0 if i == target else 0.0) for i, pi in enumerate(p)]
    print(grad)
    1. A[0.5, 0.5]
    2. B[-0.5, 0.5]
    3. C[-1.0, 0.0]
    4. D[0.5, -0.5]
    Show answer

    Answer: B ([-0.5, 0.5])

    For softmax followed by cross-entropy the logit gradient is p - y: [0.5 - 1, 0.5 - 0]. Gradient descent then raises the target logit and lowers the other. [0.5, -0.5] has the signs reversed.

  6. 6.

    An untrained 10-class classifier outputs equal logits. What does this print?

    mid
    import math
    num_classes = 10
    logits = [0.0] * num_classes
    loss = -math.log(math.exp(logits[3]) / sum(math.exp(z) for z in logits))
    print(round(loss, 3))
    1. A0.1
    2. B1.0
    3. C2.303
    4. D10.0
    Show answer

    Answer: C (2.303)

    A uniform prediction gives each class probability 0.1, and -ln(0.1) = ln(10), about 2.303. Checking that the first-step loss is near ln(C) is a quick sanity check; a loss stuck at that value means the model is still predicting uniformly.

  7. 7.

    This implements inverted dropout in training mode with p = 0.5. The mask keeps the first two units. What does it print?

    easy
    import random
    rng = random.Random(0)
    p = 0.5
    x = [1.0, 1.0, 1.0, 1.0]
    mask = [1.0 if rng.random() >= p else 0.0 for _ in x]
    print([xi * m / (1 - p) for xi, m in zip(x, mask)])
    1. A[1.0, 1.0, 0.0, 0.0]
    2. B[2.0, 2.0, 0.0, 0.0]
    3. C[0.5, 0.5, 0.0, 0.0]
    4. D[1.0, 1.0, 1.0, 1.0]
    Show answer

    Answer: B ([2.0, 2.0, 0.0, 0.0])

    Inverted dropout divides surviving activations by 1 - p, so kept units become 2.0 and the expected value stays 1.0. At evaluation dropout is then the identity, with no rescaling. Without the division you would get [1.0, 1.0, 0.0, 0.0] and a train and test mismatch.

  8. 8.

    What does this print?

    easy
    def conv_out(n, k, s=1, p=0):
        return (n + 2 * p - k) // s + 1
    
    print(conv_out(32, 5), conv_out(224, 3, s=2, p=1))
    1. A28 112
    2. B32 112
    3. C28 111
    4. D27 224
    Show answer

    Answer: A (28 112)

    A 5x5 kernel with no padding loses 4 pixels: 32 - 5 + 1 = 28. A 3x3 kernel with padding 1 and stride 2 halves the size: (224 + 2 - 3) // 2 + 1 = 112. Keeping 32 would need padding 2.

  9. 9.

    How many trainable parameters does nn.Conv2d(3, 16, kernel_size=3) have, for any input image size?

    mid
    c_in, c_out, k = 3, 16, 3
    print(k * k * c_in * c_out + c_out)
    1. A144
    2. B432
    3. C448
    4. DIt depends on the image resolution
    Show answer

    Answer: C (448)

    Each of the 16 filters has 3 * 3 * 3 = 27 weights plus one bias: 16 * 28 = 448. Weight sharing makes the count independent of the image size; 432 forgets the biases.

  10. 10.

    What receptive field does a unit have after three stacked 3x3 convolutions with stride 1?

    mid
    def receptive_field(layers):
        r, jump = 1, 1
        for k, s in layers:
            r += (k - 1) * jump
            jump *= s
        return r
    
    print(receptive_field([(3, 1), (3, 1), (3, 1)]))
    1. A3
    2. B7
    3. C9
    4. D27
    Show answer

    Answer: B (7)

    Each stride-1 3x3 layer adds 2 pixels: 1 + 2 + 2 + 2 = 7. Three 3x3 layers see the same 7x7 region as one 7x7 layer with fewer parameters and more nonlinearities. Strides earlier in the stack would make later layers grow the field faster.

  11. 11.

    This computes Adam's first update for a parameter whose first gradient is 1000. What does it print?

    hard
    import math
    lr, b1, b2, eps = 0.01, 0.9, 0.999, 1e-8
    g = 1000.0
    m = (1 - b1) * g
    v = (1 - b2) * g * g
    m_hat, v_hat = m / (1 - b1), v / (1 - b2)
    print(round(lr * m_hat / (math.sqrt(v_hat) + eps), 6))
    1. A10.0
    2. B0.01
    3. C0.001
    4. D1000.0
    Show answer

    Answer: B (0.01)

    After bias correction, m_hat = g and sqrt(v_hat) = |g|, so the first step is about lr * sign(g), regardless of the gradient's size. SGD would step lr * g = 10.0. Adam's scale-invariance is its strength, and the full-size early steps are one reason for warmup.

  12. 12.

    With PyTorch-style momentum and a constant gradient of 1, what does this print?

    mid
    mu, v = 0.9, 0.0
    for g in [1.0, 1.0, 1.0]:
        v = mu * v + g
    print(round(v, 2))
    1. A1.0
    2. B2.71
    3. C3.0
    4. D0.27
    Show answer

    Answer: B (2.71)

    The velocity goes 1, 1.9, 2.71 and approaches 1 / (1 - 0.9) = 10 under a constant gradient. Consistent directions therefore get up to ten times the step, while alternating gradients cancel. 0.27 would be the dampened (1 - mu) formulation.

  13. 13.

    A CNN with batch norm scores well in validation during training, but in production its prediction for the same image changes depending on which other images share the request batch. What is the most likely cause?

    mid
    1. AThe model was not switched to model.eval(), so batch norm uses per-batch statistics
    2. BDropout is still active
    3. CThe learning rate is too high
    4. DBatch norm always behaves this way at inference
    Show answer

    Answer: A (The model was not switched to model.eval(), so batch norm uses per-batch statistics)

    In training mode batch norm normalizes with the current batch's mean and variance, so outputs depend on batch-mates. model.eval() switches it to the stored running statistics. Active dropout would cause random variation even for an identical batch, not batch-dependent outputs.

  14. 14.

    For a transformer activation of shape (batch, seq_len, d_model), which values does nn.LayerNorm(d_model) average over to compute one mean?

    mid
    1. AAll batch elements at the same position and feature
    2. BAll tokens in the sequence for one feature
    3. CEvery value in the tensor
    4. DThe d_model features of a single token
    Show answer

    Answer: D (The d_model features of a single token)

    Layer norm normalizes each token vector over its own features, so no statistics cross the batch or the sequence. Averaging across the batch for each feature is what batch norm does, which is why layer norm behaves identically in training and inference.

  15. 15.

    Queries and keys have 64 unit-variance components. This measures the variance of their dot products. What does it print?

    hard
    import random
    rng = random.Random(0)
    d, trials = 64, 20000
    dots = []
    for _ in range(trials):
        q = [rng.gauss(0, 1) for _ in range(d)]
        k = [rng.gauss(0, 1) for _ in range(d)]
        dots.append(sum(a * b for a, b in zip(q, k)))
    var = sum(x * x for x in dots) / trials
    print(round(var), round(var / d, 1))
    1. A1 0.0
    2. B8 0.1
    3. C63 1.0
    4. D4096 64.0
    Show answer

    Answer: C (63 1.0)

    A sum of 64 independent unit-variance products has variance about 64, so raw scores have standard deviation about 8 and the softmax saturates. Dividing the scores by sqrt(d_k) = 8 divides the variance by 64, back to about 1, which is the reason for the scaling in attention.

  16. 16.

    A causal mask is applied before the softmax. What attention weights does the first token get?

    mid
    import math
    scores = [[2.0, 1.0, 0.5],
              [0.3, 1.2, 0.8],
              [1.0, 1.0, 1.0]]
    masked = [[s if j <= i else -math.inf for j, s in enumerate(row)] for i, row in enumerate(scores)]
    e = [math.exp(s) for s in masked[0]]
    print([round(x / sum(e), 2) for x in e])
    1. A[0.63, 0.23, 0.14]
    2. B[1.0, 0.0, 0.0]
    3. C[0.33, 0.33, 0.33]
    4. D[nan, nan, nan]
    Show answer

    Answer: B ([1.0, 0.0, 0.0])

    The first token may attend only to itself; masked scores of -inf become exp(-inf) = 0, so all the weight lands on position 0. [0.63, 0.23, 0.14] is the unmasked softmax, which would let the token see the future. NaN appears only if every score in a row is masked.

  17. 17.

    A toy transformer caches keys and values in fp16. How many bytes does this KV cache use?

    hard
    layers, kv_heads, head_dim, seq_len, batch, bytes_per = 2, 2, 4, 10, 1, 2
    print(2 * layers * kv_heads * head_dim * seq_len * batch * bytes_per)
    1. A160
    2. B320
    3. C640
    4. D1280
    Show answer

    Answer: C (640)

    The leading 2 counts keys and values: 2 * 2 * 2 * 4 * 10 * 1 * 2 = 640. The cache grows linearly with sequence length and batch, which is why grouped-query attention (fewer KV heads) and cache quantization matter at long contexts. 320 forgets that both K and V are stored.

  18. 18.

    A rank-4 LoRA adapter is added to a frozen 1024 by 1024 weight matrix. What does this print?

    mid
    d_in, d_out, r = 1024, 1024, 4
    print(r * (d_in + d_out), d_in * d_out)
    1. A8192 1048576
    2. B4096 1048576
    3. C16 1048576
    4. D1048576 8192
    Show answer

    Answer: A (8192 1048576)

    LoRA trains A (r x d_in) and B (d_out x r): 4 * 2048 = 8192 parameters, under 1 percent of the 1,048,576 frozen ones. Only those need gradients and optimizer state. 4096 counts just one of the two matrices.

  19. 19.

    Python's struct format e is IEEE half precision (fp16). What does this print?

    hard
    import struct
    
    def to_fp16(x):
        return struct.unpack('e', struct.pack('e', x))[0]
    
    print(to_fp16(1e-8), to_fp16(0.1))
    1. A1e-08 0.1
    2. B1e-08 0.0999755859375
    3. CAn OverflowError
    4. D0.0 0.0999755859375
    Show answer

    Answer: D (0.0 0.0999755859375)

    fp16's smallest positive value is about 6e-8, so 1e-8 underflows to zero, and 0.1 is rounded to the nearest representable value. Small gradients vanishing this way is why fp16 training uses loss scaling; bf16 has float32's exponent range and rarely needs it.

  20. 20.

    How many parameters does PyTorch's nn.LSTM(input_size=10, hidden_size=20) (one layer) have?

    hard
    input_size, hidden_size = 10, 20
    # 4 gates, each with W_ih, W_hh, b_ih and b_hh
    print(4 * (hidden_size * input_size + hidden_size * hidden_size + 2 * hidden_size))
    1. A640
    2. B2480
    3. C2560
    4. D1920
    Show answer

    Answer: C (2560)

    Each of the four gates has a 20x10 input matrix, a 20x20 recurrent matrix and two 20-element biases in PyTorch: 4 * 640 = 2560. The textbook formula with a single bias per gate gives 2480, and a GRU has three gate blocks instead of four (1920).

esc