Sources and reproducibility
A complete interactive adaptation of the ML course’s CNN lecture and tutorials, expanded with the DL course’s numerical examples.
128 frames in one deck. The ML sequence supplies the backbone: data and locality, filters, padding and stride, pooling, RGB, LeNet, MNIST, transfer, and representation visualization. The attention lectures inform progressive reveals, persistent objects, and a shared numerical model.
The lecture route
| Chapter | Frames | Start |
|---|---|---|
| 1 · Images, data, and locality | 12 | Open chapter |
| 2 · Build a local detector | 12 | Open chapter |
| 3 · Padding and stride | 10 | Open chapter |
| 4 · Pooling, colour, and feature maps | 12 | Open chapter |
| 5 · Rebuild the LeNet exercise | 13 | Open chapter |
| 6 · Train and open up the MNIST network | 18 | Open chapter |
| 7 · Follow the shared gradients | 12 | Open chapter |
| 8 · Spatial reasoning and complete accounting | 17 | Open chapter |
| 9 · From LeNet to modern CNN blocks | 11 | Open chapter |
| 10 · Transfer, representations, and checks | 11 | Open chapter |
Every frame mapped to its original source · Recording controls and run instructions
The original course material
ImageNet and WordNet, camera parameter counts, locality, edge filters, padding, stride, pooling, RGB, the full LeNet exercise, training, transfer, and feature visualization.
The exact 6×6 binary image, 3×3 vertical filter, and real-image filtering.
Sliding windows, explicit dot products, and animated patch movement.
Padding and stride, MNIST/CIFAR examples, red/green/blue planes, the original beach and buildings, sharpening, and directional filtering.
The 28×28 LeNet-style network, 5,000-example training split, ten epochs of Adam, evaluation, learned filters, all intermediate activations, and flattening. Saved notebook figures are retained alongside a newly reproduced inspectable model.
Exact pet-band input and shared-gradient calculation; equivariance conditions; receptive-field recursion; the complete 5,418-parameter and 1,622,336-MAC classifier.
Small filters, channel mixing, bottleneck accounting, Inception, scalar residuals, depthwise operations, real ResNet activations, and measured pet transfer experiments.
Reference for teaching rhythm and interaction: predict, calculate, inspect, then extend the same model. This is an independent CNN runtime.
MNIST: a real reproduced experiment
The architecture and split match the ML teaching notebook: Conv 1→6, 5×5; ReLU/max pool; Conv 6→16, 5×5; ReLU/max pool; flatten 256; Linear 120, 84, 10. Total: 44,426 parameters. The notebook uses 28×28 inputs; the lecture’s 32×32 exercise instead has 61,706 parameters.
Training: official train indices 0–4,999. Validation: official train indices 50,000–50,999. Test: all 10,000 official test examples. Adam, learning rate 0.001, batch 64, seed 0, ten fixed epochs. No test-based checkpoint selection. The new run achieved 9,610 correct test predictions. The old notebook also reports 96.1%; the weights and curves here are a separate run.
Saved epochs 0, 1, 3, and 10 are selectable. Every displayed activation and probability is recalculated in JavaScript from the selected weights. This selection replays saved training checkpoints; it does not train LeNet in the browser. Four deliberately chosen mistakes accompany one test example per digit. The PCA views use the same first 1,000 test images; labels only colour points. Each representation has its own fitted PCA axes.
Full experiment evidence · Reproduction script · Independent browser/PyTorch parity check
Other numerical evidence
The 26-parameter bar classifier really trains in the browser. Its six synthetic images are training examples, and its displayed accuracy measures training fit only. The live forward trace and classifier use the same current weights. Its gradients are checked by finite differences and a PyTorch companion.
The pet transfer plots and pretrained ResNet activations are copied from the existing DL evidence. The transfer study uses six breeds, 72 train / 36 validation / 36 sealed test images. Head-only, late-stage, and full tuning use different recipes. A one-image validation advantage is not a universal result. Original results and selection contract.
Conventions and refinements
Library convolution is cross-correlation: the kernel is not flipped. The old SciPy tutorial used mathematical convolution, so directional-filter signs can differ. Parameter counts include biases unless the frame says weights only. MAC counts exclude bias additions, nonlinearities, pooling, normalization, and memory traffic.
Flattening retains values; generic MLPs do not require independent inputs. The repeated valid 5×5 sequence 32→28→…→4 never reaches 1. Float32 uses four bytes. Convolutional feature maps are equivariant only under stated boundary and stride conditions; pooling does not guarantee invariant classification. A visualized Conv2 kernel slice is not its entire six-channel filter.
Source images and saved notebook plots are copied unchanged; browser transforms are stated on the relevant slides. Oxford-IIIT Pet images and derivatives retain the source-recorded CC BY-SA 4.0 attribution; copyright remains with original image owners. Detailed figure provenance.
Primary external teaching references
Reference for connecting affine layers to convolution and for following the tensor transformation through a network. This version makes the local dot products editable. Official lecture video.
Reference for explicit shapes, parameter sharing, and layer-by-layer numerical reasoning. Used alongside the existing ML lecture’s architecture worksheet.
Reference for the practical transition from image features to a pretrained backbone and a new classifier. The browser toy is explicitly separated from real transfer evidence.
The locality and sharing derivation motivates treating convolution as an architectural constraint. The channel chapter supports the distinction between summing input slices and stacking output filters.
Reference for receptive-field size, effective stride, and the distinction between possible and effective influence. The lesson adds a live support map showing dilation holes.
Motivates testing the failure of shift invariance and discussing low-pass filtering before subsampling. No claim that pooling automatically guarantees invariance.
A supporting reference for local connectivity, weight counts, and the channel-spanning filter. Here, legal output placements use the general floor formula.
Further interactive reading: inspect how a complete CNN connects individual operations. The present lesson’s browser experiment is a separate original six-example model.
Implementation contract for cross-correlation, NCHW, weight shape, groups, stride, padding, and dilation. See the companion for executable parity checks.
Review used available lecture slides, official notes, and papers; it did not involve watching each lecture video end to end. All worksheets, prose, diagrams, and browser code here are newly authored adaptations of the course material, not copied lecture figures.
Supports the distinction between the broad labelled collection and the classification benchmark.
Original large-scale ImageNet CNN paper.
Additional explanation of the convolutional classifier and implementation.