ES 667
  • Home
  • Lectures
  • Grading
  • Deadlines
  • FAQ
  • Materials

Lecture library

Notes, slides, practice and videos for every lecture.

Full lecture playlist 3-minute playlist
Foundations 1–4Optimization 5–9Generalization 10Attention 11–15Vision 16–20
Notes Slides Cheat sheet Notebook Recording 3-min summary

Foundations

Lectures 1–4
Lecture 01

Where Deep Learning Losses Come From I

Distributions, likelihood, and loss.

NotesSlidesCheat sheet

3-minute summaries

  • Where do loss functions come from?2:54
Lecture 02

Where Deep Learning Losses Come From II

MLE, regression, and classification.

NotesSlidesCheat sheet

3-minute summaries

  • Where MSE and cross-entropy come from2:53
Lecture 03

Likelihood to Loss in Practice

MAP, regularization, and robust observation models.

NotesLecture slidesCheat sheet

Practice notebooks

  • Likelihood notebook
  • Robust regression

3-minute summaries

  • Outliers: how the noise model picks the loss2:55
  • Priors: where regularization comes from2:54
Lecture 04

From Linear Models to Neural Networks

Neurons, activations, multilayer networks, XOR, and approximation.

NotesSlidesCheat sheet

Practice notebooks

  • Linear gradient descent
  • XOR

3-minute summaries

  • From weighted sums to neurons: solving XOR2:59
  • ReLU hinges and universal approximation2:53

Gradients and optimization

Lectures 5–9
Lecture 05

Optimization I — Gradient Estimates & Step Sizes

Batches, stochastic gradients, and optimizer geometry.

NotesSlidesCheat sheet

Full recording

  • Optimization for Deep Learning31:38

3-minute summaries

  • SGD vs minibatch: how much data per step?3:00
  • Why gradient descent zigzags: step sizes2:52
Lecture 06

Backpropagation & Autograd from Scratch

Computation graphs, reverse mode, and a scalar engine.

NotesSlidesCheat sheet

Practice notebooks

  • Colab notebook

Full recording

  • Backpropagation and Autograd from Scratch19:44

3-minute summaries

  • Backprop by hand: one local rule3:03
  • Build a tiny autograd engine from scratch2:59
Lecture 07

Optimization II — Momentum, RMSProp & Adam

Moving averages and adaptive coordinate scaling.

NotesSlidesCheat sheet

Full recordings

  • Momentum in Deep Learning29:50
  • Momentum, RMSProp and Adam23:30

3-minute summaries

  • Momentum: averaging away the zigzag3:03
  • RMSProp and Adam: a step size per coordinate3:14
Lecture 08

Vector, Matrix & Affine Autograd

Gradients, Jacobians, dense layers, and batched VJPs.

Calculus notesAutograd notesCalculus slidesAutograd slidesCheat sheet

Practice notebooks

  • Colab notebook

3-minute summaries

  • Backprop through layers: Jacobians to batches3:04
Lecture 09

Valid Outputs by Design & Learning-Rate Schedules

Constrained parameters, warmup, decay, and cosine.

Valid outputs notesLR schedule notesValid outputs slidesLR schedule slidesCheat sheet

3-minute summaries

  • Valid outputs by design: softplus to softmax2:59
  • Learning-rate schedules: warmup, decay, cosine2:53

Generalization

Lectures 10
Lecture 10

Generalization & Regularization

Learning curves, weight decay, early stopping, dropout, augmentation, Mixup, and label smoothing.

NotesSlidesCheat sheet

Practice notebooks

  • Learning curves
  • Dropout and regularization

3-minute summaries

  • Overfitting, weight decay and early stopping3:08
  • Dropout, augmentation, Mixup and label smoothing3:08

Attention and language

Lectures 11–15
Lecture 11

From characters to next-token prediction

Tokenization, learned embeddings, a hidden-layer MLP, training, and generation.

NotesSlidesCheat sheet

Full recording

  • Language Modeling: From Characters to Next-Token Prediction1:06:54

3-minute summaries

  • Next-token prediction from scratch3:01
Lecture 12

Self-attention, from first principles

NotesSlidesCheat sheet

Full recording

  • Self-Attention from First Principles1:11:18

3-minute summaries

  • Why attention?2:32
  • Query, key, value3:05
Lecture 13

Self-attention, continued

Embedding updates, next-token prediction, and positional encoding.

NotesSlidesCheat sheet

3-minute summaries

  • From attention to prediction3:12
  • Position and order2:56
Lecture 14

Self-attention and multi-head attention

Self-attention notesMulti-head attention notesSelf-attention slidesMulti-head attention slides
Lecture 15

From Attention to Applications

Encoder, decoder-only and encoder–decoder Transformers; token and sentence classification; causal self-attention and cross-attention.

Interactive notesPDF handoutSlides

Full recording

  • From Attention to Applications: Encoder, Decoder and Cross-Attention51:43

Vision and language

Lectures 16–20
Lecture 16

Vision Transformer for image classification

Patch embeddings, full attention, embedding updates and CLS readout; CNN comparisons and image-covering experiments. 55 slides including the cover.

NotesSlidesPDF · 55 pages
Lecture 17

CLIP: From Fixed Labels to Language-Defined Vision

Applications, two encoders, a step-by-step contrastive-loss calculation, zero-shot classification and linear probing.

HTML lectureSlidesPDF · 164 pages
Lecture 18

From CLIP to Vision-Language Models

Visual prefixes, cross-attention and learned summaries; architecture diagrams, code walkthroughs, training and generation.

HTML lectureSlidesPDF · 155 pages
Lecture 19

Object Detection: From One Label to a Set of Objects

Spatial candidates, target assignment, worked losses, decoding, NMS, evaluation and DETR-style set prediction.

HTML lectureSlidesPDF · 133 pages

Interactive practice

  • Detector lab: score thresholds and NMS
Lecture 20

Beyond Boxes: Dense, Structured, and Promptable Vision

Semantic, instance and panoptic segmentation; keypoints, depth, promptable masks and language grounding.

HTML lectureSlidesPDF · 60 pages

Interactive practice

  • Beyond Boxes lab: masks, pose, depth and grounding

No lectures match. Try another topic or resource type.

Notes are the complete PDF handouts or interactive lectures. Slides open the classroom presentation. Cheat sheets are two-page reviews.

Prof. Nipun Batra · IIT Gandhinagar

 

Course materials