“Walk me through what a convolutional layer computes” is a staple of computer vision and general deep learning interviews. Interviewers use it to separate candidates who have only stacked Conv2d layers from those who can predict a tensor shape, count parameters on a whiteboard and explain why a network sees enough of the image to recognise an object. The follow-ups are almost always arithmetic: what shape comes out, how many weights does this layer hold, and how far can one unit see.
Before you start
You should know what a dot product is and picture a neural network layer as “weights times inputs plus a bias, then a nonlinearity”. The examples are plain Python 3 with the standard library. The PyTorch snippets in the worked scenario target PyTorch 2.x and are illustrative: they were not executed for this article.
The short answer
A convolutional layer slides a small learned kernel (for example 3x3 across all input channels) over the input and, at every position, takes the dot product of the kernel with the patch underneath, adds a bias and writes one number into the output map. Libraries actually compute cross-correlation (no kernel flip), which makes no difference because the kernel is learned. For input size n, kernel k, padding p and stride s, the output size is floor((n + 2p - k) / s) + 1, and the layer has k*k*C_in*C_out + C_out parameters regardless of image size. Because the same weights are reused everywhere, convolution is translation equivariant, and stacking layers grows the receptive field so deeper units see larger regions.
How it works
Place a 3x3 window on the top-left corner of the input, multiply the nine values by the matching kernel weights and sum: that is output pixel (0, 0). Slide one column right and repeat. With several input channels the kernel has a 3x3 slice per channel and the dot product runs over all of them; each output channel has its own kernel. The single-channel version, without a kernel flip:
def conv2d(image, kernel, stride=1, padding=0):
"""Cross-correlation, as deep learning libraries compute it (no kernel flip)."""
if padding:
w = len(image[0]) + 2 * padding
zero_row = [0] * w
image = ([zero_row] * padding
+ [[0] * padding + row + [0] * padding for row in image]
+ [zero_row] * padding)
k = len(kernel)
out_h = (len(image) - k) // stride + 1
out_w = (len(image[0]) - k) // stride + 1
out = []
for i in range(out_h):
row = []
for j in range(out_w):
total = 0
for a in range(k):
for b in range(k):
total += image[i * stride + a][j * stride + b] * kernel[a][b]
row.append(total)
out.append(row)
return out
image = [
[0, 0, 0, 9, 9],
[0, 0, 0, 9, 9],
[0, 0, 0, 9, 9],
[0, 0, 0, 9, 9],
[0, 0, 0, 9, 9],
]
vertical_edge = [
[-1, 0, 1],
[-1, 0, 1],
[-1, 0, 1],
]
for row in conv2d(image, vertical_edge):
print(row)
# [0, 27, 27]
# [0, 27, 27]
# [0, 27, 27]The input is dark on the left and bright on the right. The kernel subtracts the left column of each patch from the right column, so it fires (27) wherever the window straddles the dark-to-bright boundary and stays at 0 over the flat region. A learned filter is exactly this: a pattern detector whose weights gradient descent chooses.
Two properties follow from the sliding. Weight sharing: the same nine numbers are used at every position, so the parameter count does not depend on image size. Translation equivariance: shift the input and the output map shifts by the same amount, because each output only depends on the local patch. Pooling and the final classifier then add a degree of translation invariance, which is a different property: the prediction stays the same when the object moves.
Step-by-step walkthrough
Step 1: Change padding and stride, then shift the input
for row in conv2d(image, vertical_edge, padding=1):
print(row)
# [0, 0, 18, 18, -18]
# [0, 0, 27, 27, -27]
# [0, 0, 27, 27, -27]
# [0, 0, 27, 27, -27]
# [0, 0, 18, 18, -18]
for row in conv2d(image, vertical_edge, stride=2, padding=1):
print(row)
# [0, 18, -18]
# [0, 27, -27]
# [0, 18, -18]
flipped = [row[::-1] for row in vertical_edge[::-1]] # true mathematical convolution
print(conv2d(image, flipped)[0]) # [0, -27, -27]
shifted = [[0] + row[:-1] for row in image] # move the content one column right
print(conv2d(shifted, vertical_edge)[0]) # [0, 0, 27]Padding 1 keeps the 5x5 size, which is why “same” padding for a 3x3 kernel is 1. The negative last column is a fake bright-to-dark edge created by the zero padding, a real border artefact in feature maps. Stride 2 samples every other position and roughly halves the size. Flipping the kernel (textbook convolution) only negates this antisymmetric kernel; since the network learns the weights, it would simply learn the flipped version. The shifted input moves the response one column right, which is equivariance in action.
Step 2: Compute output shapes without running the model
def conv_out(n, k, s=1, p=0, d=1):
"""PyTorch's formula; with d=1 it is floor((n + 2p - k) / s) + 1."""
return (n + 2 * p - d * (k - 1) - 1) // s + 1
print(conv_out(5, 3)) # 3
print(conv_out(5, 3, p=1)) # 5
print(conv_out(224, 7, s=2, p=3)) # 112 (ResNet stem convolution)
print(conv_out(112, 3, s=2, p=1)) # 56 (ResNet stem max pool)
print(conv_out(8, 3, s=2)) # 3 (floor drops the leftover column)
print(conv_out(32, 3, d=2)) # 28 (dilation 2 makes a 3x3 kernel span 5)Dilation spaces the kernel taps apart, so a 3x3 kernel with dilation d covers d*(k-1) + 1 pixels with still only nine weights. The floor matters: when the stride does not divide evenly, the last rows and columns are silently ignored, which is why odd input sizes produce surprising shapes.
Step 3: Count parameters and see why 1x1 convolutions exist
def conv_params(k, c_in, c_out, bias=True):
return k * k * c_in * c_out + (c_out if bias else 0)
def dense_params(n_in, n_out, bias=True):
return n_in * n_out + (n_out if bias else 0)
print(conv_params(3, 3, 64)) # 1792
print(conv_params(3, 64, 128)) # 73856
print(dense_params(32 * 32 * 3, 32 * 32 * 64)) # 201392128
direct = conv_params(3, 256, 256)
bottleneck = conv_params(1, 256, 64) + conv_params(3, 64, 64) + conv_params(1, 64, 256)
print(direct, bottleneck, round(direct / bottleneck, 1)) # 590080 70016 8.4
print(3 * 3 * 3 * 64 * 64, 7 * 7 * 64 * 64) # 110592 200704A 3x3 conv from 3 to 64 channels has 1,792 parameters; a dense layer mapping a 32x32x3 input to the same 32x32x64 output needs over 201 million, one weight per input-output pair.
A 1x1 convolution mixes channels at each pixel without looking at neighbours: it is a dense layer applied per position. ResNet’s bottleneck block uses one to shrink 256 channels to 64, runs the 3x3 conv on the thin tensor, and expands back, using about 8.4 times fewer weights than a direct 3x3 conv. The last line is the VGG argument: three stacked 3x3 layers see a 7x7 region with fewer weights than one 7x7 layer, and add two extra nonlinearities.
Step 4: Pool to downsample
def max_pool(x, k=2, s=2):
return [[max(x[i + a][j + b] for a in range(k) for b in range(k))
for j in range(0, len(x[0]) - k + 1, s)]
for i in range(0, len(x) - k + 1, s)]
def avg_pool(x, k=2, s=2):
return [[sum(x[i + a][j + b] for a in range(k) for b in range(k)) / (k * k)
for j in range(0, len(x[0]) - k + 1, s)]
for i in range(0, len(x) - k + 1, s)]
fmap = [[1, 3, 2, 0], [4, 6, 1, 1], [0, 2, 9, 5], [1, 1, 3, 7]]
print(max_pool(fmap)) # [[6, 2], [2, 9]]
print(avg_pool(fmap)) # [[3.5, 1.0], [1.0, 6.0]]
print(sum(map(sum, fmap)) / 16) # 2.875 (global average pooling)Pooling has no learned parameters. Max pooling keeps the strongest activation in each window (“was this feature present anywhere here?”), average pooling keeps the mean. Global average pooling collapses each channel’s whole map to one number, producing a C-length vector whatever the spatial size.
Step 5: Track receptive field across layers
def receptive_field(layers):
"""layers: list of (kernel, stride, dilation). Returns (r, jump) after each layer."""
r, jump = 1, 1
history = []
for k, s, d in layers:
k_eff = d * (k - 1) + 1
r = r + (k_eff - 1) * jump # r_l = r_{l-1} + (k_l - 1) * product of earlier strides
jump = jump * s
history.append((r, jump))
return history
print(receptive_field([(3, 1, 1)] * 3)) # [(3, 1), (5, 1), (7, 1)]
print(receptive_field([(3, 1, 1), (3, 1, 2), (3, 1, 4)])) # [(3, 1), (7, 1), (15, 1)]
print(receptive_field([(3, 1, 1), (3, 1, 1), (2, 2, 1)] * 3)[-1]) # (36, 8)The receptive field is the input region that can influence one output unit. Each layer adds k - 1 steps, and every earlier stride multiplies the size of a step in input pixels (the jump). That is why downsampling early grows the receptive field quickly, and why dilations of 1, 2, 4 grow it exponentially (3, 7, 15) without losing resolution. The theoretical receptive field is an upper bound: in practice pixels near its centre have much more influence than those at the edge.
Worked scenario
A team trains a classifier on 224x224 images with three blocks of 3x3 conv (padding 1), ReLU and 2x2 max pool, then flattens into a linear layer. The illustrative PyTorch 2.x code:
import torch
from torch import nn
def block(c_in, c_out):
return nn.Sequential(nn.Conv2d(c_in, c_out, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2))
class SmallNet(nn.Module):
def __init__(self, num_classes=10):
super().__init__()
self.features = nn.Sequential(block(3, 32), block(32, 64), block(64, 64))
self.classifier = nn.Linear(64 * 28 * 28, num_classes) # hard-coded for 224x224
def forward(self, x):
return self.classifier(torch.flatten(self.features(x), 1))Later they feed 256x256 crops to capture more detail, and a batch of 8 fails with RuntimeError: mat1 and mat2 shapes cannot be multiplied (8x65536 and 50176x10). The shape arithmetic explains it exactly:
def flatten_size(n, channels=64, blocks=3):
for _ in range(blocks):
n = conv_out(n, 3, s=1, p=1) # 3x3 conv, padding 1 keeps the size
n = conv_out(n, 2, s=2) # 2x2 max pool halves it (floor)
return channels * n * n, n
for size in (224, 256, 250):
print(size, flatten_size(size))
# 224 (50176, 28)
# 256 (65536, 32)
# 250 (61504, 31)The convolutions accept any size; only the flatten-then-linear step bakes in 28x28. Resizing back to 224 works but discards the extra detail. The robust fix is to make the head size-independent with adaptive or global average pooling:
self.pool = nn.AdaptiveAvgPool2d(1) # (N, 64, H, W) -> (N, 64, 1, 1) for any H, W
self.classifier = nn.Linear(64, num_classes)
def forward(self, x):
return self.classifier(torch.flatten(self.pool(self.features(x)), 1))The head also shrinks from 501,770 parameters to 650. One caution: a model trained at 224 may lose some accuracy at 256 because object scales shift, so fine-tune or evaluate at the new resolution.
Common mistake
- “Convolution flips the kernel, so frameworks are wrong.” Frameworks compute cross-correlation and call it convolution. The learned weights absorb the flip, so the distinction only matters when you port hand-designed kernels.
- Forgetting the input channels. A 3x3 conv from 64 to 128 channels has 3x3x64 weights per filter, not 3x3. Say the full
k*k*C_in*C_out + C_out. - “Parameters grow with image size.” Only activations, memory and compute do.
- “Pooling makes CNNs translation invariant.” Convolution is equivariant; pooling gives only limited local invariance, and strided downsampling can make outputs change when the input shifts by one pixel.
- Adding receptive fields directly. Two 3x3 layers after a stride-2 layer add
2 * 2each, not 2. Track the product of strides.
Verify the behavior
Check the formulas against brute force, including a receptive field computed by walking one output back to the input indices it reads:
def test_formula_matches_brute_force():
for n in range(3, 12):
for k in (1, 2, 3):
for s in (1, 2, 3):
for p in (0, 1):
img = [[1] * n for _ in range(n)]
ker = [[1] * k for _ in range(k)]
assert len(conv2d(img, ker, stride=s, padding=p)) == conv_out(n, k, s, p)
def brute_force_rf(layers):
positions = {0}
for k, s, d in reversed(layers):
positions = {i * s + d * j for i in positions for j in range(k)}
return max(positions) - min(positions) + 1
def test_receptive_field_recurrence():
for layers in ([(3, 1, 1)] * 3, [(7, 2, 1), (3, 2, 1), (3, 1, 1), (3, 1, 1)],
[(3, 1, 1), (3, 1, 2), (3, 1, 4)], [(3, 1, 1), (2, 2, 1)] * 4):
assert receptive_field(layers)[-1][0] == brute_force_rf(layers)
test_formula_matches_brute_force(); test_receptive_field_recurrence()
print("ok")Follow-up questions
Why do ResNets use skip connections? A residual block computes x + F(x), so it learns a correction to the identity and the gradient has a direct additive path through the + x. That made 50 to 150+ layer networks trainable, where equally deep plain stacks trained worse than shallower ones.
What is a depthwise separable convolution? A per-channel spatial conv (groups=C_in) followed by a 1x1 conv that mixes channels. It costs roughly k*k*C_in + C_in*C_out weights instead of k*k*C_in*C_out, which is why MobileNet uses it.
How do you count FLOPs for a conv layer? Multiply-accumulates are about H_out * W_out * C_out * k * k * C_in. Unlike parameters, this scales with resolution.
Interview exercise
A network applies, in order: a 7x7 conv with stride 2, a 3x3 max pool with stride 2, then two 3x3 convs with stride 1 (all with appropriate padding). The input is 224x224. What is the spatial size after the stem, and what is the receptive field of one unit after the last conv? How many parameters does each 3x3 conv have if both have 64 input and 64 output channels with bias?
Answer and reasoning
Sizes come from the output formula: the 7x7 stride-2 conv with padding 3 gives floor((224 + 6 - 7) / 2) + 1 = 112, and the 3x3 stride-2 pool with padding 1 gives 56. The two stride-1 convs with padding 1 keep 56x56. For the receptive field, track r and the jump: the 7x7 layer gives r = 7, jump 2; the pool adds (3 - 1) * 2 = 4, so r = 11, jump 4; each 3x3 conv adds (3 - 1) * 4 = 8, giving 19 and then 27. That matches the [(7, 2), (11, 4), (19, 4), (27, 4)] the calculator prints for this stack. Each conv has 3 * 3 * 64 * 64 + 64 = 36,928 parameters. Interviewers listen for two points: padding changes sizes but not the receptive field, and strides multiply every later layer’s contribution.