1 · Images, data, and locality

From pixels to a trained image classifier.

Convolutional
neural networks

Read an image. Build a local detector. Train the whole network. Then look inside it.

Input & geometryLearned weightsResponses & predictions
Six learned filters from the original ML course LeNet notebook
1 · Images, data, and locality

The route through this lecture

One sequence, following the ML course lecture.

Start withWork throughFinish with
Images and dataFilters · padding · stride · poolingA complete LeNet
RGB and feature mapsShapes · parameters · learned weightsA real MNIST walkthrough
The DL extensionGradients · receptive fields · efficient blocksTransfer and representation analysis
1 · Images, data, and locality

A category has many appearances.

A useful training set must expose variation within a class.

Real Oxford-IIIT Pet image of a cat

Pose, lighting, scale, background, and occlusion change the pixels.

The class label can stay the same.

A model needs examples of the variation we expect it to handle.
1 · Images, data, and locality

ImageNet made data part of the story.

Categories, many examples, and a shared evaluation task.

WordNeta hierarchy of concepts
→
Imagesmany views per concept
→
Labelsa supervised prediction task
→
Benchmarka common comparison
Why would five pictures of a cat be a weak training set?

They cover very few poses, backgrounds, and lighting conditions. More examples can expose variation, though quantity alone does not fix sampling bias.

1 · Images, data, and locality

What exactly is being predicted?

One image enters. A score for every candidate class leaves.

RGB imageheight × width × 3
→
Networklearned transformations
→
1,000 logitsILSVRC classification
→
Top predictionsrank the class scores

Top-1: is the highest-scoring class correct?

Top-5: is the correct class among the five highest?

Dataset, labels, split, and metric define the experiment.
1 · Images, data, and locality

LeNet → AlexNet → deeper networks

The operation stayed simple; the scale and architectures changed.

1998 · LeNet

Learned convolutional features for document recognition.

2012 · AlexNet

Large-scale ImageNet training, ReLU, GPUs, and regularization.

What we will unpack

Local filters, shared parameters, nonlinearities, and learned hierarchies.

1 · Images, data, and locality

An image is an array of numbers.

Spatial axes and channels describe different things.

An actual RGB pixel from the DL course photograph
Grayscale: H × W × 1
RGB: H × W × 3

A channel is a plane of values. Three channels at one location describe one colour pixel.

1 · Images, data, and locality

Could we connect every pixel to 100 hidden units?

First count the weights of the obvious dense layer.

1 · Images, data, and locality

Flattening does not delete the pixels.

It changes the indexing. The architecture determines which relationships are easy to learn.

Image
123456789
[1, 2, 3, 4, 5, 6, 7, 8, 9]

All nine values remain. A generic dense layer gives each position its own weights.

The concern is the missing architectural constraint, not lost numerical information.
1 · Images, data, and locality

An ear is useful wherever it appears.

Should the detector need a new set of weights at every location?

Cat feature at one spatial location
The same cat feature shifted within the canvas
What would we like to reuse?

The local calculation that detects a pattern. Its response map should move when the pattern moves.

1 · Images, data, and locality

A dense layer can learn this.
Why make it learn the same thing twice?

A useful edge may appear at many locations. Make nearby pixels interact, then reuse the same weights at every location.

32 × 32 grayscale → 32 × 32 responsesConnectionsParameters, with biases
DenseEvery input → every output1,049,600
Locally connected, 3 × 3A different local rule at each position10,240
Convolution, 3 × 3One shared local rule; padding 110
Locality reduces connections. Sharing ties weights. They are two separate assumptions.

Flattening is reversible; it does not destroy pixel values. A dense layer simply does not enforce the spatial structure; it does not assume statistical independence of the pixels. Locally connected count includes 9 weights + 1 bias at every padded output.

1 · Images, data, and locality

Local patterns can be composed.

A larger pattern can depend on several smaller ones.

Pixelslocal contrast
→
Early featuresedges and textures
→
Later featurescombinations and parts
→
Classifierevidence for a label
This is a useful interpretation, not a promise that every channel will become a named “ear detector”.
2 · Build a local detector

Begin with the ML course’s 6×6 image.

Three bright columns meet three dark columns.

X · 6×6
111000111000111000111000111000111000

Which part of this image is an edge?

A patch of constant brightness should give little evidence for an edge.

We want a response to a change across a local neighbourhood.
2 · Build a local detector

Compare the left side with the right side.

A 3×3 kernel can ask one local question.

K · shared 3×3 weights
10-110-110-1
left column − right column

The middle column gets weight zero. Three rows repeat the comparison.

Would a uniformly bright patch respond?
3 − 3 = 0
2 · Build a local detector

First patch: all three columns are bright.

Predict the first output value.

Top-left patch
111000111000111000111000111000111000
What is y[0,0]?
(1−1) + (1−1) + (1−1) = 0
2 · Build a local detector

Move one column to the right.

The same weights now straddle the boundary.

Patch starting at column 1
111000111000111000111000111000111000
What is y[0,1]?
(1−0) + (1−0) + (1−0) = 3
2 · Build a local detector

One local calculation makes a feature map.

Four legal starts along each axis give a 4×4 result.

Input 6×6
111000111000111000111000111000111000
Output 4×4
0330033003300330
Nine kernel weights produce sixteen responses. The responses are not sixteen learned parameters.
2 · Build a local detector

One patch becomes one response.

Choose an output cell to move the kernel. Click input pixels to flip 0 ↔ 1.

Each output uses the same nine weights. Only the input patch changes.
2 · Build a local detector

Reverse the contrast. Reverse the response.

Dark → bright and bright → dark carry different signs.

Opposite contrast
000111000111000111
One output row
0-3-30
What happens if we negate the kernel instead?

Every response changes sign. A later learned layer can use either orientation.

2 · Build a local detector

The exact same calculation in PyTorch.

Name the batch and channel axes even for a single image.

x = torch.zeros(1, 1, 6, 6)
x[:, :, :, :3] = 1
k = torch.tensor([[1., 0., -1.]] * 3)
k = k.reshape(1, 1, 3, 3)
y = F.conv2d(x, k)
# y.shape: [1, 1, 4, 4]
# every output row: [0, 3, 3, 0]
A handwritten calculation and a library call should agree.
2 · Build a local detector

The kernel is not flipped.

Deep-learning libraries usually call this “convolution”; the forward operation is cross-correlation.

Y[r,c] = b + Σᵢ Σⱼ K[i,j] X[r+i,c+j]

Cross-correlation

Use K in the orientation shown in the worksheet.

Mathematical convolution

Reverse K along both spatial axes before sliding.

For a learned kernel, the model can learn either orientation. For a specified edge filter, the convention changes the answer.

Indices here use stride 1, no padding, and one channel. Bias is one scalar per output channel, shared over positions.

2 · Build a local detector

Change the question by changing the weights.

The computation stays “multiply, sum, move”.

An averaging filter uses nine weights of 1/9. It suppresses local variation instead of emphasizing it.
2 · Build a local detector

Return to the beach and the buildings.

The original ML tutorial photographs, with the filtering computed live.

2 · Build a local detector

Who chooses the filter?

So far, we did. In a CNN, training does.

Image-processing example

Choose K → compute Y

Useful for understanding a local operator.

Supervised CNN

Initialize K → loss → ∂L/∂K → update K

Useful patterns are selected jointly with the classifier.

3 · Padding and stride

Count the legal starting positions.

A length-3 window has four valid starts in a length-6 row.

012345
First start: 0 · last start: n − k
Number of starts: n − k + 1
For a 6×6 input and a 3×3 kernel: 4×4 output.
3 · Padding and stride

What if we keep applying valid 5×5 filters?

Use the 32×32 example from the ML lecture.

3 · Padding and stride

Does every input pixel participate equally?

Count how many output patches contain each input pixel.

3 · Padding and stride

Add a border before sliding.

Zero padding makes additional kernel starts legal.

3×3 image with p=1 border
0000001110011100111000000
Input width becomes n + 2p
Output width = n + 2p − k + 1
3 · Padding and stride

When does the output keep the same size?

Solve the stride-1 equation.

n + 2p − k + 1 = n
p = (k − 1) / 2
KernelSymmetric paddingOutput at stride 1
3×31n×n
5×52n×n
4×4not an integerneeds asymmetric padding for exact same size
3 · Padding and stride

Stride chooses which starts we visit.

A length-3 kernel on length 8, stride 2.

Legal starts: 0, 2, 4
Why is start 6 invalid?

It would need positions 6, 7, and 8; the final input index is 7.

⌊(8 − 3) / 2⌋ + 1 = 3
3 · Padding and stride

Geometry changes where we ask.

Change one control at a time. The outlined input cells are the actual sampled taps.

Zero padding invents boundary values. Dilation spaces the taps; it does not add kernel weights. Output cells can be selected to inspect each placement.

3 · Padding and stride

Count moves, then include the first position.

Kernel spankₑ = d(k − 1) + 1
→
Available traveln + 2p − kₑ
→
Legal startsfloor(travel / s) + 1
nₒᵤₜ = ⌊(n + 2p − d(k−1) − 1) / s⌋ + 1
Predict: n = 6, k = 3, p = 1, s = 2, d = 1. Is the output width 3 or 4?
Reveal the reasoning

3. The padded width is 8. Kernel starts 0, 2, 4 fit; start 6 would reach past the boundary. ⌊(6 + 2 − 3)/2⌋ + 1 = 3.

Apply independently to height and width. For s = 1 and odd kₑ, p = (kₑ − 1)/2 preserves size. “Same” is a padding policy, not a universal p = 1 rule.

3 · Padding and stride

Predict all three output shapes.

Keep channels separate from spatial arithmetic.

Input H×WKernel / padding / strideYour prediction
7×73 / 0 / 2?
7×73 / 1 / 2?
32×325 / 2 / 1?
3×3 · 4×4 · 32×32
Stride and padding change the number of responses. They do not create additional kernel weights.
3 · Padding and stride

Dilation spaces the samples inside a kernel.

A 3×3 kernel can span 5×5 while keeping nine weights.

Sampled offsets · dilation 2
1010100000101010000010101
Effective width = d(k−1)+1
= 2(3−1)+1 = 5
Bounding span and the number of sampled points are different quantities.
4 · Pooling, colour, and feature maps

Replace a local window by its maximum.

The first pooling example: 4×4 → 2×2.

Input
1320421501722346
With a 2×2 window and stride 2, what is the result?
Max pool
4537
No learned weights. The window size and stride are hyperparameters.
4 · Pooling, colour, and feature maps

Replace the same window by its average.

A different summary retains different information.

Same input
1320421501722346
What is the top-left average?
(1 + 3 + 4 + 2) / 4 = 2.5
Full result: [[2.5, 2], [1.5, 4.75]]
4 · Pooling, colour, and feature maps

ReLU gates responses. Pooling summarizes a window.

Compare max and average pooling with and without ReLU. This signed feature map is illustrative.

ReLU changes values, not shape. This 2 × 2, stride-2 pool changes spatial size, not channel count.

Pool each channel independently. Global average pooling instead averages the entire spatial map into one value per channel.

4 · Pooling, colour, and feature maps

What information did max pooling remove?

Two different patches can give the same pooled value.

Patch A
4000
Patch B
0004
max(A) = max(B) = 4
Some local movements become indistinguishable. A movement across a window boundary can still change the output.
4 · Pooling, colour, and feature maps

RGB is three aligned planes.

A colour image is not three unrelated photographs.

The same spatial address selects one value from each input channel.
4 · Pooling, colour, and feature maps

A single filter spans all input channels.

For RGB and a 3×3 spatial window: one filter has 27 weights.

R patch3×3 · Kᵣ
→
G patch3×3 · Kɢ
→
B patch3×3 · Kʙ
→
One responsesum all 27 products + b
y = ⟨Xᵣ,Kᵣ⟩ + ⟨Xɢ,Kɢ⟩ + ⟨Xʙ,Kʙ⟩ + b
4 · Pooling, colour, and feature maps

Three input slices. One output value.

Inspect one RGB patch. Change a channel’s kernel scale and watch its contribution change.

Sum across input channels. Stack across output filters.

Hand-chosen 2 × 2 RGB patch and weights; each channel has its own spatial weights. One bias is added after the channel sum.

4 · Pooling, colour, and feature maps

One filter gives one output channel.

A second full-depth filter asks a different question.

3 input channels
K₀: 3×3×3 → response map 0
K₁: 3×3×3 → response map 1
K₂: 3×3×3 → response map 2
… K₁₅ → response map 15
Sum over input channels. Stack across output filters.
4 · Pooling, colour, and feature maps

First add a bias. Then apply a nonlinearity.

The ML lecture’s final building block completes the layer.

z = convolution(X, K) + b
a = ReLU(z) = max(0, z)
z = [−3, −1, 0, 2, 5]
a = [0, 0, 0, 2, 5]
ReLU preserves the tensor shape. One output channel shares one bias across its spatial positions.
4 · Pooling, colour, and feature maps

Why put nonlinearities between the layers?

Composition of affine maps is still an affine map.

W₂(W₁x + b₁) + b₂
= (W₂W₁)x + (W₂b₁ + b₂)
ReLU makes the composition more expressive than one affine transformation.

The same principle applies when the matrices have convolutional structure.

4 · Pooling, colour, and feature maps

“Sixteen filters” means sixteen whole volumes.

X: [N, 3, 32, 32]
W: [16, 3, 3, 3]   b: [16]
Y: [N, 16, 32, 32]   (p=1, s=1)

Parameters

16 × (3 × 3 × 3 + 1) = 448

Responses per image

16 × 32 × 32 = 16,384

The input-channel axis is reduced inside each dot product. The output-channel axis indexes different learned questions.

PyTorch convention: NCHW. For groups = 1, every filter reads all input channels. Deeper feature channels need not correspond to colors.

4 · Pooling, colour, and feature maps

Count the bank, not the image.

Input [N, 3, 32, 32]; sixteen 3×3 filters; padding 1.

How many weights, biases, and responses per image?
QuantityCalculationValue
Weights16×3×3×3432
Biasesone per output channel16
Parameters432+16448
Responses16×32×3216,384
N changes the batch, not the learned parameter count.
5 · Rebuild the LeNet exercise

Now assemble a complete network.

The LeNet-style network from the ML lecture.

Input1×32×32
→
Conv + pool6×14×14
→
Conv + pool16×5×5
→
Flatten400
→
MLP120 → 84 → 10
We will account for every shape and every learned parameter.
5 · Rebuild the LeNet exercise

Q1 · What is the input?

Read the axes before reading the architecture.

One image32 high × 32 wide
→
One channelgrayscale
→
One batch elementN = 1
What shape would PyTorch receive?
[1, 1, 32, 32]
5 · Rebuild the LeNet exercise

Q2 · How does 32 become 28?

The first convolution is valid, stride 1.

Solve for the kernel width.
28 = 32 − k + 1
k = 5
Input1×32×32
→
Six 5×5 filterseach spans 1 channel
→
Output6×28×28
5 · Rebuild the LeNet exercise

Q3 · How many parameters are in Conv1?

Six filters; one input channel; one bias per filter.

Count weights and biases separately.
6 × 1 × 5 × 5 + 6 = 156
Outputs: 6 × 28 × 28 = 4,704
4,704 activations, 156 parameters. The same filter is reused at 784 positions.
5 · Rebuild the LeNet exercise

Q4 · How does 28 become 14?

A 2×2 pool with stride 2 works independently in all six channels.

Before pool6×28×28
→
2×2 / stride 2zero learned parameters
→
After pool6×14×14
Does pooling reduce the channel count?

No. Each channel is pooled separately; six feature maps remain.

5 · Rebuild the LeNet exercise

Q5 · How does 6×14×14 become 16×10×10?

Each of the sixteen filters spans all six input channels.

Spatial output: 14 − 5 + 1 = 10
What is the weight tensor?
[16, 6, 5, 5]
16 × 6 × 5 × 5 + 16 = 2,416 parameters
5 · Rebuild the LeNet exercise

Q6 · Pool the second set of feature maps.

A 2×2 pool with stride 2 changes space again.

Before pool16×10×10
→
Max poolingzero parameters
→
After pool16×5×5
How many numbers remain per image?
16 × 5 × 5 = 400
5 · Rebuild the LeNet exercise

Q7 · How do the feature maps meet an MLP?

Flatten the final activation tensor into one vector per image.

[N, 16, 5, 5] → [N, 400]
400 featuresflattened activations
→
Linear + ReLU120
→
Linear + ReLU84
→
Linear10 logits
Flattening introduces no learned parameters. The following linear layer does.
5 · Rebuild the LeNet exercise

Where are most of the parameters?

Count the dense head one layer at a time.

LayerWeights + biasesParameters
400 → 120400×120 + 12048,120
120 → 84120×84 + 8410,164
84 → 1084×10 + 10850
The first dense layer alone is much larger than both convolutional layers together.
5 · Rebuild the LeNet exercise

Revisit the LeNet exercise.

Switch input size. Predict which parameter count changes before you run it.

Valid 5 × 5 convolutions; 2 × 2, stride-2 pools; biases included. This is the fully connected LeNet-style variant in the course materials.

5 · Rebuild the LeNet exercise

The notebook uses 28×28 MNIST directly.

The same kernels produce a different flatten dimension.

28input
→
24 → 12conv1 → pool
→
8 → 4conv2 → pool
→
16×4×4256 features
FC1: 256×120 + 120 = 30,840
Total: 44,426 parameters
Copying Linear(400,120) into this notebook would cause a shape error.
5 · Rebuild the LeNet exercise

Define the trainable parts.

Five small module definitions implement the notebook network.

self.conv1 = nn.Conv2d(1, 6, kernel_size=5)
self.conv2 = nn.Conv2d(6, 16, kernel_size=5)
self.fc1 = nn.Linear(16 * 4 * 4, 120)
self.fc2 = nn.Linear(120, 84)
self.fc3 = nn.Linear(84, 10)
Check: sum(p.numel() for p in model.parameters()) == 44_426
5 · Rebuild the LeNet exercise

Write the forward pass in the same order.

Each line corresponds to one step we already calculated.

x = F.max_pool2d(F.relu(self.conv1(x)), 2)  # 6×12×12
x = F.max_pool2d(F.relu(self.conv2(x)), 2)  # 16×4×4
x = x.flatten(1)                         # 256
x = F.relu(self.fc1(x))                  # 120
x = F.relu(self.fc2(x))                  # 84
logits = self.fc3(x)                     # 10
6 · Train and open up the MNIST network

Make the training experiment explicit.

Reproduce the ML notebook’s data split and model.

SplitOfficial MNIST indicesPurpose
Trainingtrain[0:5000]learn parameters
Validationtrain[50000:51000]inspect generalization during development
Testall 10,000 official test imagesevaluate once after the fixed run

Adam · learning rate 0.001 · batch size 64 · 10 epochs · seed 0

The new measurements are a separate reproducible run. The original notebook outputs remain linked for comparison.
6 · Train and open up the MNIST network

One minibatch has four axes.

Sixty-four grayscale images, each 28×28.

images.shape = [64, 1, 28, 28]
labels.shape = [64]
for images, labels in train_loader:
    logits = model(images)     # [64, 10]
    loss = F.cross_entropy(logits, labels)
Each label is an integer class index from 0 to 9.
6 · Train and open up the MNIST network

The training loop is the one you already know.

Convolution changes the architecture, not the gradient-descent recipe.

model.train()
for images, labels in train_loader:
    optimizer.zero_grad(set_to_none=True)
    logits = model(images)
    loss = F.cross_entropy(logits, labels)
    loss.backward()
    optimizer.step()
The loss updates convolution kernels, convolution biases, and all dense-layer parameters.
6 · Train and open up the MNIST network

Did training make progress on unseen validation images?

Measured cross-entropy from the reproduced experiment.

6 · Train and open up the MNIST network

Read the test result in its experimental context.

The fixed 10-epoch model is evaluated on the official test split.

9,610 / 10,000 correct = 96.10%

What this supports

This small LeNet generalizes to many held-out MNIST digits.

What it does not establish

Robustness to arbitrary handwriting, photographs, or distribution shifts.

Now inspect both successful predictions and errors.
6 · Train and open up the MNIST network

Choose a digit. Inspect its prediction.

Every probability comes from a real forward pass in this browser.

6 · Train and open up the MNIST network

What did the first six filters learn?

Compare random initialization with the trained kernel bank.

6 · Train and open up the MNIST network

Open one calculation inside a learned filter.

At spatial position [8,8], multiply 25 pixels by 25 learned weights.

6 · Train and open up the MNIST network

Conv1 produces six 24×24 response maps.

These are the signed values before ReLU.

6 · Train and open up the MNIST network

ReLU keeps the positive responses.

No shape change: six 24×24 maps remain.

6 · Train and open up the MNIST network

Pool each channel independently.

The six 24×24 maps become six 12×12 maps.

6 · Train and open up the MNIST network

Conv2 combines information across the six channels.

Sixteen filters produce sixteen 8×8 maps.

6 · Train and open up the MNIST network

The second ReLU changes values, not axes.

Sixteen 8×8 feature maps become nonnegative.

6 · Train and open up the MNIST network

After the second pool: sixteen 4×4 maps.

We have reached the feature tensor that feeds the dense head.

6 · Train and open up the MNIST network

Flatten preserves those 256 numbers.

Name the index ranges rather than treating flatten as a learned operation.

6 · Train and open up the MNIST network

The dense head learns combinations of features.

256 activations → 120 hidden units → 84 hidden units.

6 · Train and open up the MNIST network

Ten scores, then one probability distribution.

The output layer has 84×10 + 10 = 850 parameters.

6 · Train and open up the MNIST network

Compare with the original notebook’s feature-map walkthrough.

The same sequence was already a strong part of the ML course.

Original ML notebook conv1 activation gallery
Original ML notebook pool2 activation gallery

Original notebook · Reproduction source in this repository

7 · Follow the shared gradients

Reuse the kernel twice.

The DL lecture’s exact one-dimensional example.

x = [2, 1, 3] · w = [0.5, −1] · b = 0
Compute the two outputs.
y₀ = 2×0.5 + 1×(−1) = 0
y₁ = 1×0.5 + 3×(−1) = −2.5
7 · Follow the shared gradients

The loss sends one upstream derivative to each output.

Suppose ∂L/∂y = [1, 2].

g₀ = 1 · g₁ = 2
Contribution from y₀?
∂L/∂w from y₀ = 1 × [2,1] = [2,1]

The same two weights also affect y₁. Their work is not finished.

7 · Follow the shared gradients

One shared weight has several paths to the loss.

The contributions add.

from y₀: 1 × [2,1] = [2,1]
from y₁: 2 × [1,3] = [2,6]
∂L/∂w = [2,1] + [2,6] = [4,7]
∂L/∂b = 1 + 2 = 3
7 · Follow the shared gradients

One kernel. A sum of gradient contributions.

Use x = [2, 1, 3], w = [0.5, −1], b = 0. The rest of the network sends upstream gradients g = [1, 2].

∂L/∂wⱼ = Σᵣ (∂L/∂yᵣ) xᵣ₊ⱼ. Sum across positions, then across examples according to the loss reduction.

Input gradients accumulate too: overlapping windows send multiple contributions to the same input pixel.

7 · Follow the shared gradients

The overlapping input also receives two contributions.

The middle value x₁ appears in both windows.

∂L/∂x₀ = 1×0.5 = 0.5
∂L/∂x₁ = 1×(−1) + 2×0.5 = 0
∂L/∂x₂ = 2×(−1) = −2
Zero gradient at x₁ comes from cancellation, not from being unused.
7 · Follow the shared gradients

Use the gradient for one update.

Start at w = [0.5, −1], b = 0. Choose learning rate 0.1.

∂L/∂w = [4,7] · ∂L/∂b = 3
w ← [0.5,−1] − 0.1[4,7] = [0.1,−1.7]
b ← 0 − 0.1×3 = −0.3
The updated kernel will again be reused at every spatial location.
7 · Follow the shared gradients

The two-dimensional rule is the same sum.

For stride 1 and no padding, with upstream map G:

∂L/∂K[i,j] = Σᵣ Σ꜀ G[r,c] X[r+i,c+j]
∂L/∂b = Σᵣ Σ꜀ G[r,c]

Add another sum over examples in the batch. With input channels, each kernel slice uses its corresponding input plane.

One kernel coefficient is updated using evidence from every place where it was reused.
7 · Follow the shared gradients

Make a whole training step small enough to inspect.

A second example complements MNIST: six 5×5 bar images.

5×5 imagehorizontal or vertical
→
2 learned filters3×3 + bias + ReLU
→
Global means2 features
→
Linear head2 logits
2×(9+1) + (2×2+2) = 26 parameters
This tiny model exposes the mechanics. MNIST supplies the larger held-out experiment.
7 · Follow the shared gradients

Can two filters separate horizontal from vertical?

Six 5 × 5 training images: one bar, two orientations, three positions. Watch the weights and prediction change.

Actual full-batch gradient descent in JavaScript, cross-entropy averaged over six images, learning rate 0.4, seed 7. All metrics are training metrics. This tiny exercise does not measure generalization to photographs.

7 · Follow the shared gradients

Turn a response map into one feature.

This inspector shares the training widget’s current model.

One spatial mean per channel gives two features for the classifier.
7 · Follow the shared gradients

A response is not a probability.

The classifier weights combine both features. Then softmax turns two logits into a distribution.

The training objective averages this image-level cross-entropy over all six examples.

Uses the current trained model and the same selected example as the preceding trace. Displayed products are rounded; calculations use full precision.

7 · Follow the shared gradients

A complete model needs very little code.

model = nn.Sequential(
    nn.Conv2d(1, 2, 3),   # [N,2,3,3]
    nn.ReLU(),
    nn.AdaptiveAvgPool2d(1),
    nn.Flatten(1),        # [N,2]
    nn.Linear(2, 2),      # logits
)

optimizer.zero_grad()
logits = model(x)
loss = F.cross_entropy(logits, y)
loss.backward()
optimizer.step()

The browser uses this architecture: 26 parameters.

20 in the convolution; 6 in the classifier. ReLU and averaging have no learned parameters.

Cross-entropy expects raw logits. Use softmax only when displaying probabilities.

The architecture matches; reproducing the browser’s exact values also requires copying its weights, examples, and optimizer settings.

8 · Spatial reasoning and complete accounting

Make the numbers small enough to follow.

Real pet-photo crop divided into five horizontal bands

The existing DL lecture averages five horizontal bands from a 45 × 45 crop, then thresholds each band at 128.

Band means: 113.83, 112.37,
145.50, 145.78, 109.32

Binary rows: 0, 0, 1, 1, 0

This deliberately simplified image carries the arithmetic. It is not the photograph’s full-resolution pixel matrix.

Keep one input and one kernel in view while the operation unfolds.
8 · Spatial reasoning and complete accounting

What does an edge filter respond to?

Select a filter. The right image is computed from the left image’s grayscale pixels.

The photo comes from the existing Oxford-IIIT Pet teaching evidence. These are hand-designed filters, not learned features. Signed responses use a fixed scale: blue is negative, red is positive, white is zero.

8 · Spatial reasoning and complete accounting

Move the input. What should happen to the feature map?

Reuse of the same local rule gives a translation-equivariant operation under suitable conditions.

f(shift(X)) = shift(f(X))

Equivariance

The representation moves in the corresponding way.

Invariance

The final result stays the same.

Stride, finite boundaries, and downsampling must be part of the claim.
8 · Spatial reasoning and complete accounting

Shift the input. Does the response follow?

Compare f(shift(X)) with shift(f(X)). Circular boundaries isolate the stride effect.

At stride 1 with circular boundaries, integer circular shifts commute exactly. At stride 2, only shifts aligned to the stride have an integer output-shift counterpart. Finite zero boundaries can also break equality.

8 · Spatial reasoning and complete accounting

Pooling can change when a feature crosses a boundary.

Max pool, width 2, stride 2:

[0, 1 | 2, 0] → [1, 2]
[0, 0 | 1, 2] → [0, 2]

The second input is the first shifted right by one, with zero fill. Its pooled map is not a simple integer shift of the first output.

Equivariance means the output transforms with the input. Invariance means the output stays the same.

Global averaging is invariant to permutations of the positions being averaged. A whole CNN is invariant only when its preceding feature transformation and boundaries support that claim. Low-pass filtering before subsampling can reduce aliasing; it is not a blanket guarantee.

8 · Spatial reasoning and complete accounting

Which input pixels can affect this output?

Track both receptive-field width r and the output spacing j.

Start: r₀ = 1, j₀ = 1
rₗ = rₗ₋₁ + (kₗ−1)dₗ jₗ₋₁
jₗ = jₗ₋₁ sₗ
A later kernel steps across locations that may already be several input pixels apart.
8 · Spatial reasoning and complete accounting

Conv → pool → conv: follow the same recurrence.

The DL course’s three-layer example.

LayerKernel / strideReceptive field rOutput jump j
Input—11
Conv3 / 131
Pool2 / 242
Conv3 / 182
Why does the last 3×3 layer add four pixels of reach?
(3−1) × previous jump 2 = 4
8 · Spatial reasoning and complete accounting

Depth grows the receptive field.

Select a stack. The highlighted pixels can influence one interior output, away from boundaries.

The theoretical receptive field counts possible influence. Actual influence depends on weights and activations; the effective receptive field can be much smaller.

8 · Spatial reasoning and complete accounting

The same field of view can hide different functions.

One 5 × 5 layer

25C² weights
5 × 5 receptive field

3 × 3 → ReLU → 3 × 3

18C² weights
5 × 5 receptive field

Assume C input, hidden, and output channels; stride 1; dilation 1; ignore biases.

The stacked block uses fewer weights and adds a nonlinearity. It is not numerically equivalent to an arbitrary 5 × 5 convolution.

Without ReLU, a stack is still linear (affine with biases), but factorization restricts the possible kernels, and finite padding can complicate boundary equivalence. The next ledgers compare a flatten-based head with a global-average head.

8 · Spatial reasoning and complete accounting

A large matrix with repeated entries.

[y₀] [a b 0 0] [x₀]
[y₁] = [0 a b 0] [x₁]
[y₂] [0 0 a b] [x₂]
[x₃]

A 1-D valid kernel [a, b], no bias.

Zeros encode locality. Repeated a and b encode weight sharing.

In im2col, collect patches into columns and multiply a flattened filter by those columns.

Libraries may choose other algorithms; materializing all patches can use substantial memory.

The convolution layer is linear without bias, affine with bias. The CNN needs nonlinearities to build a richer function.
8 · Spatial reasoning and complete accounting

Build a second complete classifier.

This is the DL course’s 32×32 RGB example.

RGB3×32×32
→
Conv116×32×32
→
Pool16×16×16
→
Conv232×16×16
→
GAP → Linear32 → 10

Both convolutions: 3×3, padding 1, stride 1. Pool: 2×2, stride 2.

8 · Spatial reasoning and complete accounting

Conv1: capacity and work are different counts.

3 input channels → 16 output channels, on a 32×32 grid.

Parameters = 16×(3×3×3+1) = 448
MACs = 32×32×16×(3×3×3) = 442,368
Activations = 16×32×32 = 16,384
8 · Spatial reasoning and complete accounting

Conv2: less space, more channels.

After pooling: 16 input channels, 32 output channels, 16×16 output.

Parameters = 32×(3×3×16+1) = 4,640
MACs = 16×16×32×(3×3×16) = 1,179,648
A smaller spatial grid can still cost more when channel mixing grows.
8 · Spatial reasoning and complete accounting

Global average pooling gives one value per channel.

Each 16×16 map becomes its spatial mean.

32×16×16feature maps
→
Average over H,Wno learned parameters
→
32 valuesone per feature type
→
Linear → 1032×10 + 10 = 330 parameters
The same head accepts different spatial resolutions. It trades away detailed spatial location.
8 · Spatial reasoning and complete accounting

Write the ledger before training.

LayerOutput C × H × WParametersMACs / image
Input3 × 32 × 3200
Conv 3×3, 3→16, p116 × 32 × 32448442,368
ReLU + max pool 2, s216 × 16 × 160—
Conv 3×3, 16→32, p132 × 16 × 164,6401,179,648
ReLU + global average32 × 1 × 10—
Linear 32→1010 logits330320
Total5,4181,622,336

Counts include learned biases; MAC total covers only convolution and linear layers. Pooling, averaging, nonlinearities, bias additions, and softmax also cost work.

8 · Spatial reasoning and complete accounting

Shared weights still do repeated work.

Change resolution or channel counts. Predict which totals will move before touching a control.

Parameters depend on kernel and channels. Spatial resolution multiplies both computation and activation storage.

One image, square 3 × 3 convolution, stride 1, padding 1, groups 1, with bias. MACs count one multiply-accumulate as one unit; exclude biases and activation functions. Float32 storage here is only the output tensor, not total training memory.

8 · Spatial reasoning and complete accounting

Double both spatial dimensions.

Keep channels, kernels, strides, and padding fixed.

What happens to parameters, activations, and convolution MACs?
QuantityEffect
Parametersunchanged
Spatial activations4× as many
Convolution MACs4× as many
GAP-to-linear MACsunchanged: 320
Exact total: 4×(442,368+1,179,648)+320 = 6,488,384 MACs
9 · From LeNet to modern CNN blocks

What changed when CNNs scaled up?

AlexNet kept the same basic sequence and made it much larger.

ImageRGB input
→
Convolution blockslearn local feature banks
→
Nonlinearities & poolingcompose and reduce
→
Dense headclass logits

ReLU, GPU training, data augmentation, and dropout helped make a larger supervised system practical.

The architectural building blocks are now familiar. The experimental scale is different.

9 · From LeNet to modern CNN blocks

VGG: repeat small filters.

Two 3×3 layers can reach a 5×5 neighbourhood.

One 5×5, 64 → 128
25×64×128 = 204,800 weights
3×3, 64 → 64 → 128
9×64×64 + 9×64×128
= 110,592 weights
For these widths: about 46% fewer weights, plus another nonlinear stage.
9 · From LeNet to modern CNN blocks

A 1 × 1 filter is a shared channel mixer.

x at one position = [2, 1, 3]
w = [1, −2, 0.5]

y = 2 − 2 + 1.5 = 1.5

Apply this same affine map independently at every spatial position.

3 → 16 channels needs 64 parameters with biases, even though the spatial footprint is one pixel.

It changes the channel representation. It does not directly read neighboring spatial positions.

“1 × 1” describes spatial extent, not the number of multiplications.
9 · From LeNet to modern CNN blocks

Make the expensive spatial operation narrow.

Keep 28×28 resolution; map 64 channels to 128.

1×164 → 16
→
3×316 → 16
→
1×116 → 128
Standard 3×3: 28²×9×64×128 = 57,802,752 MACs
A bottleneck reduces channel width before doing the expensive spatial mixing.
9 · From LeNet to modern CNN blocks

Did the bottleneck meet the target?

Change the intermediate width and recalculate all three layers.

9 · From LeNet to modern CNN blocks

Try several spatial scales in parallel.

Inception combines different local questions at the same resolution.

28×28×64
1×1 → 32 channels
1×1 → 3×3 → 48 channels
1×1 → 5×5 → 16 channels
pool → 1×1 → 32 channels
Concatenate: 32 + 48 + 16 + 32 = 128 channels
All branch heights and widths must match. Concatenation adds channels.
9 · From LeNet to modern CNN blocks

Ask the next block to learn a correction.

Preserve a direct route for the current representation.

xF(x)+yidentity shortcut
y = x + F(x)
9 · From LeNet to modern CNN blocks

Try the shortcut with one number.

Let x = 2 and F(x) = ax.

9 · From LeNet to modern CNN blocks

Keep a path for the existing representation.

y = x + F(x)

∂L/∂x = ∂L/∂y · (I + ∂F/∂x)

If F(x) starts small, the block begins near the identity. Learning a useful correction can be easier than relearning a whole mapping.

Addition needs matching shapes. If channels or resolution change, use an appropriate projection P(x): y = P(x) + F(x).

The direct path gives gradients another route. It does not guarantee that gradients can never vanish or explode.

This is the residual sum before any optional post-addition nonlinearity. The identity derivative shown applies to a shape-matched identity skip.

9 · From LeNet to modern CNN blocks

Depthwise + pointwise factors the operation.

32 channels
→
3 × 3 depthwiseone spatial kernel per input channel
→
1 × 1 pointwisemix 32 channels into 64

Standard 3 × 3, 32 → 64

9 × 32 × 64 = 18,432 weights

Depthwise + pointwise

9 × 32 + 32 × 64 = 2,336 weights
About 7.89× fewer weights and MACs at equal output resolution, excluding bias and other operations.

This factorization restricts the kernel family. It is not an exact replacement for every standard convolution, and lower MACs do not guarantee proportional latency savings.

9 · From LeNet to modern CNN blocks

Look inside the real pretrained backbone.

The same pet photograph passes through ResNet-18.

Actual pretrained ResNet-18 activations from the DL course evidence
Spatial resolution decreases while the number of channels increases.
10 · Transfer, representations, and checks

Start from a representation that has learned useful structure.

Your imagesmatching preprocessing
→
Pretrained backbonefreeze, then optionally fine-tune
→
New task headfit to your labels

First comparison

Train a new head on frozen features. Choose hyperparameters using validation data. Compare with a simple baseline.

Then fine-tune

Unfreeze deliberately, usually with a smaller backbone learning rate. Keep the final test set sealed.

Freezing gradients, setting evaluation mode, and disabling autograd are different operations.

requires_grad_(False) freezes parameters; eval() changes modules such as BatchNorm and dropout; no_grad() disables graph recording. Freezing alone does not freeze BatchNorm running statistics.

10 · Transfer, representations, and checks

A frozen backbone can become a feature extractor.

This is the ML lecture’s “store activations” transfer-learning recipe.

Training imagesfixed preprocessing
→
Frozen backboneevaluation mode
→
Stored featuresone vector per image
→
New classifiertrain on your labels
Caching is valid only while the feature computation stays fixed.
10 · Transfer, representations, and checks

The pretrained weights come with a preprocessing contract.

Resize, crop, range, and normalization affect the representation.

DL course preprocessing contract for the pretrained pet classifier
Use the preprocessing associated with the selected pretrained weights.
10 · Transfer, representations, and checks

Freezing, evaluation mode, and no_grad do different jobs.

Three independent controls are easy to confuse.

ControlWhat it changes
requires_grad_(False)whether parameter gradients are computed
model.eval()behaviour of dropout and BatchNorm
torch.no_grad()whether this forward pass records an autograd graph
A frozen backbone left in train mode can still update BatchNorm running statistics.
10 · Transfer, representations, and checks

Compare the actual pet transfer-learning recipes.

Existing DL course evidence: six breeds, 72 train / 36 validation images.

Measured DL course pet transfer curves

Head only: 33/36 · late stage: 27/36 · all layers: 34/36 at the selected validation checkpoints.

10 · Transfer, representations, and checks

Can we see structure in the learned representation?

Project the same 1,000 test images into two dimensions.

10 · Transfer, representations, and checks

A two-dimensional plot is not the full classifier.

Separation can be suggestive without being a complete explanation.

PCA preserves directions of high variance. It can hide distinctions used by the ten-class head.

Colouring by known labels helps interpretation, but visual clusters do not prove generalization or causal features.

Connect a representation plot to held-out predictions and concrete failure cases.
10 · Transfer, representations, and checks

Make the failure small enough to inspect.

SymptomFirst useful check
Shapes fail at the classifierPrint NCHW after every spatial layer; compute the flatten dimension.
Loss never improvesTry to overfit a tiny batch; inspect labels, gradient norms, and parameter updates.
Validation changes with batch orderCheck eval mode, BatchNorm statistics, and stochastic transforms.
A one-pixel shift flips a predictionInspect stride phase, padding, and pooling; compare before/after downsampling.
Transfer learning behaves strangelyCheck input range, color order, normalization, and the chosen pretrained weights.

A successful tiny-batch overfit is a plumbing check. It is not evidence that the model generalizes.

10 · Transfer, representations, and checks

Keep the original tutorials within reach.

Each notebook answers a particular question.

10 · Transfer, representations, and checks

One last calculation: explain each axis.

Input [8,3,32,32], Conv2d(3,16,3,padding=1), then 2×2 pool.

What are the weight tensor, parameter count, and final activation shape?
Weights: [16,3,3,3]
Parameters: 16×3×3×3 + 16 = 448
After pool: [8,16,16,16]
Batch, channels, spatial dimensions, and learned parameters are four different counts.
10 · Transfer, representations, and checks

Can you explain the network without naming an architecture?

Why do shared weights receive a sum of gradients?

The same parameter affects many spatial outputs. The chain rule adds its contribution along every path to the loss.

Does a larger image change model capacity or cost?

With convolution and global pooling, learned parameter count stays fixed. Spatial compute and activation storage grow. A flatten-based fixed-size head would change this answer.

Why is the browser’s high training accuracy insufficient?

It trained on six clean synthetic images and reports those same examples. Generalization requires independent evaluation on a declared target distribution.

Runnable PyTorch parity companion · Existing convolution notebook · Sources and teaching choices