Where Deep Learning Losses Come From I
Distributions, likelihood, and loss.
Notes, slides, practice and videos for every lecture.
Distributions, likelihood, and loss.
MLE, regression, and classification.
MAP, regularization, and robust observation models.
Neurons, activations, multilayer networks, XOR, and approximation.
Batches, stochastic gradients, and optimizer geometry.
Computation graphs, reverse mode, and a scalar engine.
Moving averages and adaptive coordinate scaling.
Gradients, Jacobians, dense layers, and batched VJPs.
Constrained parameters, warmup, decay, and cosine.
Learning curves, weight decay, early stopping, dropout, augmentation, Mixup, and label smoothing.
Tokenization, learned embeddings, a hidden-layer MLP, training, and generation.
Embedding updates, next-token prediction, and positional encoding.
Encoder, decoder-only and encoder–decoder Transformers; token and sentence classification; causal self-attention and cross-attention.
Patch embeddings, full attention, embedding updates and CLS readout; CNN comparisons and image-covering experiments. 55 slides including the cover.
Applications, two encoders, a step-by-step contrastive-loss calculation, zero-shot classification and linear probing.
Visual prefixes, cross-attention and learned summaries; architecture diagrams, code walkthroughs, training and generation.
Spatial candidates, target assignment, worked losses, decoding, NMS, evaluation and DETR-style set prediction.
Semantic, instance and panoptic segmentation; keypoints, depth, promptable masks and language grounding.
No lectures match. Try another topic or resource type.
Notes are the complete PDF handouts or interactive lectures. Slides open the classroom presentation. Cheat sheets are two-page reviews.