Lecture decks

These 25 canonical PDFs are the course sequence. Each PDF is a compact, one-slide-per-page Typst handout; Present retains the progressive classroom builds where available. [Q], [V], [D], and [I] mark short checkpoints, visuals, derivations, and optional interactives. Start with L1 if you are new to the sequence, or use this page as the authoritative map of the course.

Module 1 · Foundationslosses → models → calculus → backprop
1Where Deep Learning Losses Come Froma probabilistic view — distributions → likelihood → loss, priors → MAP → regularizationPDF·Present·Main Colab·Robust Colab
2From Linear Models to Neural Networkslinear & logistic regression, neurons, MLPs, universal approximationPDF·Present·Linear GD Colab·XOR Colab·Classify Lab·Approximation Lab
3Calculus Toolkitderivatives, gradients, Jacobians & Hessians — the geometry behind optimizationPDF·Present·Quadratic GD Colab·Jacobian/VJP/Hessian Colab
4Computation Graphs & Backpropagationupstream × local, gradients accumulate — a fully worked graph + autodiffPDF·Present·Complete Autograd Colab·Autograd Lab·Gradient-flow Lab
Module 2 · Optimization & Traininguse the gradients well
5Optimization for Deep Learninggradient estimates → momentum → AdaGrad / RMSProp → AdamW, with warmup + cosinePDF·Present·GD Colab·Optimizer Colab
6Making Deep Networks Trainableactivation gates → matched initialization → normalization → residual pathsPDF·Present·Autopsy Colab·Variance Colab·Controls Colab
7Generalization & Regularizationdiagnose the gap → stop well → test one parameter, representation, data, or target interventionPDF·Present·Curves Colab·Mechanisms Colab
Module 3 · Computer Visionconvolution → backbones → dense prediction
8Convolutional Neural Networkslocality + sharing → output geometry → equivariance / pooling → receptive fields → a tiny CNNPDF·Present·Convolution Colab
9Modern CNN Pipelines & Transfer Learningspend context and channels wisely → preserve identity → factor compute → reuse representationsPDF·Present·Design Moves Colab
10Localization & Object Detectionbox geometry → dense candidates + training targets → score + NMS → matching + APPDF·Present·Geometry + AP Colab·NMS Colab
11Semantic & Instance Segmentationpixel evidence → spatial context + detail → overlap-aware metrics → separate object identitiesPDF·Present·Segmentation Colab
Module 4 · Sequences & Languagetokens → embeddings → next-token
12Next-Token Predictioncausal windows → embeddings → one complete MLP language model → autoregressive decodingPDF·Present·Language Model Colab
13Sequence Models I — RNNs, LSTMs, GRUsrecurrent state, backprop through time, vanishing gradients & clipping, gated cellsPDF
14Convolutional Sequence Modelscausal & dilated convolutions, receptive fields, TCNs, WaveNetPDF
15Sequence Models II — Encoder–Decoder & Attentionseq2seq, teacher forcing, beam search, and attention over source statesPDF
Module 5 · Transformersattention is all you need
16Transformers I — Self-Attentionscaled dot-product, causal masking, multi-head, positional encoding, the blockPDF
17Transformers II — BERT, GPT & T5encoder/decoder/enc–decoder families, pretraining objectives, finetuning, decodingPDF
Module 6 · Representation & Multimodallearn features → align images & text
18Self-Supervised & Representation Learningcontrastive / SimCLR & InfoNCE, masked reconstruction / MAE, teacher–student / BYOL & DINOPDF
19Vision Transformers & Multimodal Modelsimage patches as tokens, ViT, Swin, CLIP contrastive alignment, BLIP-2 / LLaVAPDF
Module 7 · Generative Modelinglatents → adversaries → diffusion
20Autoencoders & VAEsthe latent bottleneck, the ELBO, the reparameterization trick, posterior collapsePDF
21Generative Adversarial Networksthe minimax game, the optimal discriminator, non-saturating loss, mode collapse, WGANPDF
22Diffusion Models I — Theoryforward / reverse, predict-the-noise objective, sampling, the U-Net + time conditioningPDF
23Diffusion Models II — Practicelatent diffusion, cross-attention conditioning, classifier-free guidance, DDIM, Stable DiffusionPDF
Module 8 · Systems & Frontiermake it fast → where it's going
24Efficient Inferencememory-bound decode, KV-cache & MQA/GQA, quantization, MoE, FlashAttention, speculative decodingPDF
25Frontier Topics & Course Synthesisscaling, agents & reasoning, interpretability — and full circle to Lecture 1PDF

All decks are authored in Typst. Earlier Marp and planning material is retained in the repository for provenance, but it is not part of the public course path.