← Convolutional neural networks
ES 667 · source companion

Sources and reproducibility

A complete interactive adaptation of the ML course’s CNN lecture and tutorials, expanded with the DL course’s numerical examples.

128 frames in one deck. The ML sequence supplies the backbone: data and locality, filters, padding and stride, pooling, RGB, LeNet, MNIST, transfer, and representation visualization. The attention lectures inform progressive reveals, persistent objects, and a shared numerical model.

The lecture route

ChapterFramesStart
1 · Images, data, and locality12Open chapter
2 · Build a local detector12Open chapter
3 · Padding and stride10Open chapter
4 · Pooling, colour, and feature maps12Open chapter
5 · Rebuild the LeNet exercise13Open chapter
6 · Train and open up the MNIST network18Open chapter
7 · Follow the shared gradients12Open chapter
8 · Spatial reasoning and complete accounting17Open chapter
9 · From LeNet to modern CNN blocks11Open chapter
10 · Transfer, representations, and checks11Open chapter

Every frame mapped to its original source · Recording controls and run instructions

The original course material

MNIST: a real reproduced experiment

The architecture and split match the ML teaching notebook: Conv 1→6, 5×5; ReLU/max pool; Conv 6→16, 5×5; ReLU/max pool; flatten 256; Linear 120, 84, 10. Total: 44,426 parameters. The notebook uses 28×28 inputs; the lecture’s 32×32 exercise instead has 61,706 parameters.

Training: official train indices 0–4,999. Validation: official train indices 50,000–50,999. Test: all 10,000 official test examples. Adam, learning rate 0.001, batch 64, seed 0, ten fixed epochs. No test-based checkpoint selection. The new run achieved 9,610 correct test predictions. The old notebook also reports 96.1%; the weights and curves here are a separate run.

Saved epochs 0, 1, 3, and 10 are selectable. Every displayed activation and probability is recalculated in JavaScript from the selected weights. This selection replays saved training checkpoints; it does not train LeNet in the browser. Four deliberately chosen mistakes accompany one test example per digit. The PCA views use the same first 1,000 test images; labels only colour points. Each representation has its own fitted PCA axes.

Full experiment evidence · Reproduction script · Independent browser/PyTorch parity check

Other numerical evidence

The 26-parameter bar classifier really trains in the browser. Its six synthetic images are training examples, and its displayed accuracy measures training fit only. The live forward trace and classifier use the same current weights. Its gradients are checked by finite differences and a PyTorch companion.

The pet transfer plots and pretrained ResNet activations are copied from the existing DL evidence. The transfer study uses six breeds, 72 train / 36 validation / 36 sealed test images. Head-only, late-stage, and full tuning use different recipes. A one-image validation advantage is not a universal result. Original results and selection contract.

Conventions and refinements

Library convolution is cross-correlation: the kernel is not flipped. The old SciPy tutorial used mathematical convolution, so directional-filter signs can differ. Parameter counts include biases unless the frame says weights only. MAC counts exclude bias additions, nonlinearities, pooling, normalization, and memory traffic.

Flattening retains values; generic MLPs do not require independent inputs. The repeated valid 5×5 sequence 32→28→…→4 never reaches 1. Float32 uses four bytes. Convolutional feature maps are equivariant only under stated boundary and stride conditions; pooling does not guarantee invariant classification. A visualized Conv2 kernel slice is not its entire six-channel filter.

Source images and saved notebook plots are copied unchanged; browser transforms are stated on the relevant slides. Oxford-IIIT Pet images and derivatives retain the source-recorded CC BY-SA 4.0 attribution; copyright remains with original image owners. Detailed figure provenance.

Primary external teaching references

Review used available lecture slides, official notes, and papers; it did not involve watching each lecture video end to end. All worksheets, prose, diagrams, and browser code here are newly authored adaptations of the course material, not copied lecture figures.

ImageNet · official overview

Supports the distinction between the broad labelled collection and the classification benchmark.

Krizhevsky, Sutskever, Hinton · AlexNet (2012)

Original large-scale ImageNet CNN paper.

Dive into Deep Learning · LeNet

Additional explanation of the convolutional classifier and implementation.