Deep learning · Dense and promptable vision
Beyond Boxes
Dense, Structured, and Promptable Vision
The question determines the structure of the answer.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
01What can a box tell us, and what does it leave unspecified?
Teaching notes
Students have already studied object detection, CLIP, and VLM fundamentals. This lecture asks what spatial structure an answer must contain, then combines language grounding with promptable masks. All street overlays are authored illustrations. No model prediction is claimed.
00 · One image, many questions
One image, many questions
What should a vision model output when a box is not enough?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
02What should a vision model output when a box is not enough?
Teaching notes
Section 00. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
00 · One image, many questions
One image, six questions
Same input. Different target structure.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
03Which of these questions needs one answer for every pixel?
Teaching notes
Read the question first, then the target. Classification can mean a single scene label or a vector of category-presence labels. The street contains several categories, so a single object-class softmax would not list everything present. Detection supplies a variable number of records. The examples illustrate task definitions rather than a benchmark annotation policy. Begin with the image alone. Reveal one question and its visual target at a time. The persistent family labels distinguish global outputs, sets, dense fields, and query-conditioned outputs; these are useful organizational views rather than mutually exclusive model classes.
01 · Label every pixel
Label every pixel
Which pixels are actually dog, road, or person?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
04Which pixels are actually dog, road, or person?
Teaching notes
Section 01. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
01 · Label every pixel
A box does not describe the silhouette
A box gives approximate extent. Segmentation specifies pixel membership.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
05Could we estimate the dog’s visible area from its bounding box?
Teaching notes
Reveal the mask after discussing the box. The teal outline is a manually authored, coarse silhouette. It is neither official ground truth nor a model prediction. Visible-region segmentation excludes occluded parts unless the task explicitly asks for an amodal mask.
01 · Label every pixel
Semantic segmentation labels the image grid
Each valid pixel receives one class label under the chosen annotation scheme.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
06Would two people receive different class labels?
Teaching notes
Use the road and sky first, then the people. Only selected regions are coloured so the source image remains legible. This is not an exhaustive semantic annotation: buildings, vegetation, bicycle, and a distant pedestrian are deliberately left unpainted. Colours encode class identity in this illustration.
01 · Label every pixel
One pixel has one target and C competing scores
The network produces scores; supervision supplies the correct class.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
07What is the difference between yᵢⱼ and zᵢⱼ?
Teaching notes
Start with the selected pixel on the foreground person in the street. Reveal its target class, then the feature vector at that location, then C logits. The magnified pixel patch is an authored semantic illustration around that same location; it is not a dataset annotation. The classifier sees contextual features, not only one RGB triplet.
01 · Label every pixel
The pixel classifier shares its weights
Softmax normalizes over classes separately at each location.
Authored schematic / numerical teaching example.
Long et al. (2015), Fully Convolutional Networks, §4
08Should probabilities across all image pixels sum to one?
Teaching notes
The same learned classifier operates at every spatial location. Softmax normalizes over C classes at each pixel separately. The illustrative logits log(0.10), log(0.70), and log(0.20) give probabilities exactly 0.10, 0.70, and 0.20. This keeps one numerical example throughout the pixel trace.
01 · Label every pixel
Cross-entropy asks for the target class probability
A confident mistake receives a larger penalty than a correct high-probability answer.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
09Which of the three probabilities enters the loss for this pixel?
Teaching notes
Hold the target person and probabilities 0.10, 0.70, 0.20 fixed. Reveal the negative log calculation, then draw its 0.357 value back onto the selected location in the schematic loss grid. Other cells remain blank until the image-loss slide. The calculation is illustrative and uses natural logarithms.
01 · Label every pixel
The image loss averages valid pixel losses
Semantic segmentation applies shared classification over many spatial locations.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
10If a dataset marks a boundary as unknown, should we force it to be background?
Teaching notes
Show one loss per valid pixel. The highlighted pixel retains the same 0.357 value from the previous slide. A grey crossed cell is ignored. Reveal the mean over V only after the students see the spatial loss field. Raw logits and integer targets are the usual cross-entropy inputs; argmax is not part of training.
01 · Label every pixel
A spatial pipeline produces one classifier input per pixel
The encoder supplies context; the decoder returns features to the target grid.
Authored schematic / numerical teaching example.
Long et al. (2015), Fully Convolutional Networks, §4
11What would global average pooling discard?
Teaching notes
The encoder computes contextual features on a smaller grid. A spatial decoder increases resolution and combines features. A final class head maps each location to C logits. This is a schematic, not a particular implementation: outputs may first be produced at a lower resolution and then resized to the target grid.
02 · Recover spatial detail
How do we recover spatial detail?
Where could the missing boundary information come from?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
12Where could the missing boundary information come from?
Teaching notes
Section 02. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
02 · Recover spatial detail
Upsampling increases resolution, not evidence
Several different fine boundaries can produce the same coarse summary.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
13Can interpolation recover the exact original shape from these four values?
Teaching notes
This is an explicit toy compression, not a claim that a CNN uses average pooling at every stage. The 8 × 8 binary mask is averaged within four 4 × 4 blocks. The coarse values are 0.5, 0, 0.5, and 0.125. Nearest-neighbour upsampling repeats them. A learned decoder can exploit training priors, but it cannot guarantee recovery of arbitrary detail discarded from its inputs.
02 · Recover spatial detail
Where could the missing detail come from?
A skip connection brings fine encoder features into the decoder.
Authored schematic / numerical teaching example.
Ronneberger et al. (2015), U-Net, §2
14What information does the long connection add?
Teaching notes
First trace the deep path through upsampling. Then reveal the skip from the fine encoder stage. Match the spatial dimensions before fusion. U-Net concatenates the encoder and decoder feature channels and applies convolutions. This is a feature concatenation, unlike an additive residual connection. The network still has to learn how to use the features. Do not name the connection on its first build. First ask where detail could come from, then reveal the fine feature branch, then name the skip and the fusion operation.
02 · Recover spatial detail
U-Net repeats this combination across scales
At each decoder scale, aligned encoder features help refine the prediction.
Authored schematic / numerical teaching example.
Ronneberger et al. (2015), U-Net, §2
15Which arrows provide context, and which preserve finer detail?
Teaching notes
Reveal the skip paths after following the U-shaped computation. This drawing omits channel counts, convolution blocks, and the final class head for clarity. Original U-Net used valid convolutions and cropped skip tensors before concatenation; modern variants often use padding. The conceptual claim is aligned feature fusion, not exact original tensor dimensions.
02 · Recover spatial detail
FCN also combines coarse and fine predictions
The common problem is precise localization from a hierarchy of features.
Authored schematic / numerical teaching example.
Long et al. (2015), Fully Convolutional Networks, §4
16What does this FCN diagram fuse?
Teaching notes
Long et al. showed how a classification network could produce dense predictions using convolutional score maps and learned upsampling. The skip variants combine coarse semantic scores with finer layers. The drawing summarizes score fusion and omits the final upsampling step. This is a short connection, not an architecture survey.
02 · Recover spatial detail
Inference turns class scores into a label map
The decoded output has one class ID per pixel.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
17Where did the C dimension go?
Teaching notes
The probabilities can be retained for confidence analysis, but the hard semantic prediction is an integer map. Argmax of logits gives the same winning class as argmax of softmax probabilities. Training uses differentiable scores; hard labels are convenient for display and discrete overlap metrics.
02 · Recover spatial detail
Mask IoU measures foreground agreement
Five pixels agree on foreground; four more belong to only one mask.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
18What are the intersection and union sizes?
Teaching notes
Use the exact same 4 × 4 masks throughout the builds. Ground truth has 6 foreground pixels; prediction has 8. First ask students to count overlap, then reveal errors, then show the quotient. Background agreement is outside the foreground union. These are synthetic labels and predictions, not measured model outputs. Use green for TP, coral for FP, and amber for FN. Build 3 shows counts; build 4 calculates the quotient.
02 · Recover spatial detail
Background can dominate pixel accuracy
A high accuracy can coexist with complete failure on a small foreground region.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
19Is an all-background prediction useful for finding these six pixels?
Teaching notes
The miniature 10 × 10 image has exactly six foreground pixels. Explain why accuracy is dominated by the 94 background pixels. Mean IoU averages class-wise overlap, subject to the benchmark’s absent-class convention. Report the class of interest and define which pixels and classes participate.
03 · Which instance?
Which pixels belong to which instance?
The class is person. Which person?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
20The class is person. Which person?
Teaching notes
Section 03. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
03 · Which instance?
A semantic class can contain several individuals
Semantic labels give the class. Instance IDs give the individual.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
21Could connected components always recover the people from a semantic map?
Teaching notes
The selected foreground pedestrians and cyclist are all people. The semantic illustration gives them the same colour. The reveal assigns different IDs through different colours. It does not change their shared class. A distant pedestrian is not annotated in this selected-region illustration. Avoid interpreting the coloured sample as an exhaustive instance set.
03 · Which instance?
Each instance has its own binary mask
The new field in the familiar object record is the binary mask.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
22Does mask number 1 always mean the same person across different images?
Teaching notes
An instance record can include a class, a box, and an image-sized binary mask. Boxes are useful for the detector-based architecture we will use. Some methods predict masks without a separate box stage. Unlike semantic class channels, the instance dimension N is variable. This reconnects briefly with Part I’s set target without repeating candidate assignment. Show the old class-and-box record first, then append the mask M_i. The records remain an unordered variable-size set.
03 · Which instance?
What extra detector output would describe shape?
A mask branch extends the detector’s object record.
Authored schematic / numerical teaching example.
He et al. (2017), Mask R-CNN, §3
23The detector already predicts a class and a box. What is missing?
Teaching notes
Begin with the region feature and class/box outputs already familiar from object detection. Ask for the additional output, then reveal a local mask. Only after this motivation show how region features are aligned and how Mask R-CNN implements the branches.
03 · Which instance?
A detector can provide features for one region
A region feature grid preserves spatial layout inside a proposed object.
Authored schematic / numerical teaching example.
He et al. (2017), Mask R-CNN, §3
24Why does the mask branch need a grid instead of a single pooled vector?
Teaching notes
A detector supplies proposed regions over a shared image feature map. RoIAlign samples features using bilinear interpolation without the coordinate rounding of RoIPool. It produces fixed spatial dimensions for different region sizes. The schematic follows one dog proposal. Feature extraction is shared across regions.
03 · Which instance?
Mask R-CNN adds a mask branch to region detection
Shared region features support classification, box refinement, and a local mask.
Authored schematic / numerical teaching example.
He et al. (2017), Mask R-CNN, §3
25Which branch decides the shape of the object?
Teaching notes
Reveal the class branch, the box branch, then the mask branch. The mask branch is a convolutional spatial predictor running in parallel with the class and box heads. It is not obtained by rasterizing the predicted box. The full system also includes the region proposal network and detector losses, already familiar from the object detection lecture.
03 · Which instance?
A local mask is placed back into the image
The mask grid predicts shape. The box supplies image location.
Course image and authored illustrative targets · No measured model predictions.
He et al. (2017), Mask R-CNN, §3
26Must every proposed region directly predict an H × W array?
Teaching notes
The displayed 28 × 28 binary mask is rasterized from the same authored dog polygon and box used by the lab; it is not a measured model output. Original Mask R-CNN commonly uses a 28 × 28 mask grid. At inference, select the predicted class mask, resize it to the detected box, and threshold to obtain a binary mask. Align resizing and coordinate conventions with the implementation.
03 · Which instance?
The mask loss learns binary membership
The mask branch answers foreground or background within the matched region.
Authored schematic / numerical teaching example.
He et al. (2017), Mask R-CNN, §3
27Does the mask head need a softmax across the C class channels?
Teaching notes
Mask R-CNN separates class prediction from mask prediction. For each positive matched region, the target class chooses the channel for the mask loss. At inference the predicted class chooses it. Background regions do not receive an object mask loss. Loss terms shown are schematic unit-weight sums; the paper also includes proposal training. Keep class-specific masks distinct from one semantic softmax per pixel. Keep the main explanation to the three losses on a matched positive region; class-channel mechanics belong in these notes.
03 · Which instance?
Panoptic segmentation combines things and stuff
Countable objects receive instance IDs; background regions receive semantic labels.
Course image and authored illustrative targets · No measured model predictions.
Kirillov et al. (2019), Panoptic Segmentation, §3
28Does the road need one object ID for every separate visible patch?
Teaching notes
One slide is sufficient. A panoptic output assigns each valid pixel a semantic class and an instance identifier for things, using a single non-overlapping partition. Stuff does not require object instance identity. The selected-region overlay illustrates the distinction; it is not a complete benchmark panoptic map. The things/stuff taxonomy depends on the dataset.
04 · Pose and depth
Other structured outputs: pose and depth
The output can describe geometry instead of a class or mask.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
29The output can describe geometry instead of a class or mask.
Teaching notes
Section 04. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
04 · Pose and depth
A person’s pose requires internal geometry
Named joints describe an arrangement that a bounding box cannot capture.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
30Can two people have similar boxes but different poses?
Teaching notes
Use the recurring street for context, then compare the two schematic skeletons. Amber circles are joints and blue segments show known anatomical connections. The photo overlay is authored and approximate. Keypoint annotations are image coordinates, not automatically a 3D body model.
04 · Pose and depth
Pose can be predicted as coordinates
Each joint has a defined identity, coordinate system, and annotation-validity flag.
Authored schematic / numerical teaching example.
Toshev & Szegedy (2014), DeepPose, §3
31Should an unlabelled wrist be trained toward coordinate (0, 0)?
Teaching notes
DeepPose is a direct coordinate-regression example. The numerical crop calculation is illustrative, not measured from the street image or skeleton. Normalized x is multiplied by crop width and normalized y by crop height, then the crop offset can be added to recover image coordinates. Visibility and validity are related but not identical. A common coordinate loss is squared error over supervised joints.
04 · Pose and depth
A heatmap represents possible joint locations
The output contains one spatial map for each joint type.
Authored schematic / numerical teaching example.
Newell et al. (2016), Stacked Hourglass Networks, §3
32What does the K dimension mean here?
Teaching notes
Targets often place a Gaussian bump around an annotated coordinate; a common objective is squared error between predicted and target heatmaps. The displayed maps are illustrative. The hard peak can be decoded with argmax, with optional subpixel refinement or a soft-argmax method. Heatmap dimensions h × w may be smaller than the image and require coordinate rescaling. In a top-down pipeline, predict K maps per detected person.
04 · Pose and depth
Dense prediction can be continuous
Depth assigns a number to each valid pixel instead of choosing a class.
Authored schematic / numerical teaching example.
Eigen et al. (2014), Depth Map Prediction, §3
33Would a softmax over road, person, and sky produce a distance?
Teaching notes
The grey map only illustrates near/far ordering. It is not a predicted or physically measured depth map and does not show realistic surface variation. Darker means nearer here; other visualizations reverse the convention. Depth often means camera-axis z distance rather than Euclidean range along a ray. The dataset must define it. Sky may not have a finite valid depth label.
04 · Pose and depth
A depth loss compares numerical values
A dense regression target needs units and a valid-pixel mask.
Authored schematic / numerical teaching example.
Eigen et al. (2014), Depth Map Prediction, §3
34Which positions should contribute if a depth sensor returns no measurement?
Teaching notes
The L1 expression is a simple teaching loss, not a claim that the cited Eigen system uses only L1. Depth methods may operate in log depth or inverse depth and may include scale-invariant, gradient, or geometric objectives. The purpose is to distinguish categorical per-pixel cross-entropy from continuous per-pixel regression.
04 · Pose and depth
Depth values need a scale convention
Relative depth gives ordering; metric depth requires physical scale.
Authored schematic / numerical teaching example.
Ranftl et al. (2022), Robust Monocular Depth Estimation, §3
35Can a visually plausible depth map automatically be read in metres?
Teaching notes
The ray diagram is a geometric illustration. A monocular model can learn useful scale priors from data, object sizes, and camera information. Those priors do not make absolute scale observable in every arbitrary scene. Explain the ambiguity before promising distances in metres. The diagram shows the same angular size for two different size-distance combinations. Relative predictions can capture ordering or shape while leaving scale unknown. Some methods, including MiDaS-style training, handle scale and shift ambiguity in an inverse-depth representation. Do not generalize that all relative methods have exactly the same invariance. Metric evaluation needs the camera convention, physical units, and any permitted alignment stated explicitly.
04 · Pose and depth
The answer can have several kinds of structure
What if the output also depends on what the user is asking for?
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
36How do keypoints fit beside a class, a set of masks, and a dense field?
Teaching notes
Use four persistent families: global, set, dense, query-conditioned. The first three describe output organization. The fourth adds conditioning and can produce a set or dense mask. Keypoints are structured geometry: K named coordinates per person, and a set across people. Query conditioning is an additional design choice, not a separate tensor layout.
04 · Pose and depth
Now the tensor shapes have a visual meaning
A label map has H × W entries. Its classifier produces C scores at each entry.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
37Does semantic segmentation require C ground-truth numbers per pixel?
Teaching notes
Consolidate the output notation after the visual examples. H and W refer to the chosen output grid, C to semantic classes, N to the number of annotated instances, and K to joint types. The table uses channels last and omits batch size. Detection and instance segmentation have variable target sets even if their networks allocate fixed candidate capacity. The mask implementation extends Part I’s object records. This synthesis appears only after students have seen all unprompted target types visually. Query-conditioned mask notation is introduced in the next section.
05 · Promptable vision
What if the user selects the region?
The image contains many regions. Which one is being requested?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
38The image contains many regions. Which one is being requested?
Teaching notes
Section 05. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
05 · Promptable vision
Which region do I want a mask for?
Image + query produces a spatial answer.
Course image and authored illustrative targets · No measured model predictions.
39Must a segmentation model return every object when we only want the dog?
Teaching notes
Show the point first, then the illustrative dog mask. The click conveys a region of interest without assigning a semantic class name. Both the prompt and the mask are authored teaching elements. This starts a different interface: image plus query produces a spatial answer.
05 · Promptable vision
Spatial prompts express different constraints
A point, box, or previous mask gives the model information about the requested region.
Course image and authored illustrative targets · No measured model predictions.
40What could a negative point add after a positive point?
Teaching notes
The standard released original SAM interface supports positive and negative points, boxes, and mask inputs. The right panel represents a prior mask; the actual dense mask prompt is low resolution. These inputs can be combined. No class name is required for this interface. The examples are authored, not live SAM runs. Reveal positive point, negative point, box, and previous mask in that order. The negative point lies on the road beside the dog; it marks a region to exclude.
05 · Promptable vision
A promptable model combines two sources of information
A mask decoder uses visual features together with the encoded prompt.
Authored schematic / numerical teaching example.
Kirillov et al. (2023), Segment Anything, §3 and Appendix A
41Which part changes when we click somewhere else in the same image?
Teaching notes
Build the image branch first, then the prompt branch, then the mask decoder. The original SAM uses a ViT image encoder. Sparse points and boxes have positional and type embeddings; dense mask prompts are encoded with convolutions. Its lightweight decoder uses self-attention and two-way cross-attention between prompts/output tokens and image features, then produces mask logits and estimated mask quality. This is a high-level view; details of upscaling and dynamic mask prediction are omitted. Reveal image encoder, cached embedding, prompt encoder, mask decoder, candidate masks, and predicted quality in six steps. The model name is introduced here only after the prompt interface has been motivated.
05 · Promptable vision
One prompt can admit several valid masks
Multiple candidate masks can express ambiguity in the prompt.
Course image and authored illustrative targets · No measured model predictions.
Kirillov et al. (2023), Segment Anything, §3 and Appendix A
42Does a click on the dog’s head specify only the head or the whole dog?
Teaching notes
The point is placed on the dog’s head. The alternatives are conceptual, not SAM predictions. Original SAM can produce multiple masks for ambiguous prompts, together with predicted IoU quality scores. This is an estimated quality score, not an observed ground-truth IoU or a semantic class probability. Additional prompts can clarify the intended extent.
05 · Promptable vision
Image features can be reused across prompts
Interactive refinement changes the query while keeping the image representation fixed.
Authored schematic / numerical teaching example.
Kirillov et al. (2023), Segment Anything, §3 and Appendix A
43Why separate the image encoder from the prompt-dependent computation?
Teaching notes
Trace the cached embedding and the new prompt into the decoder. Positive and negative corrections or a prior mask can refine the region. This is inference, not a gradient update to the model weights. SAM predicts masks and their quality; it does not inherently name the semantic category of the mask. Connect this to reusing expensive image representations in CLIP retrieval or cached visual context in a VLM. Reuse is architectural, not a claim that these models share the same embedding space.
06 · Language grounding
What if the query is language?
How can words become a spatial prompt?
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
44How can words become a spatial prompt?
Teaching notes
Section 06. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
06 · Language grounding
How can words select a region?
Language describes the referent; grounding must locate it.
Course image and authored illustrative targets · No measured model predictions.
45Where should the phrase “the dog beside the cyclist” act in the computation?
Teaching notes
CLIP and VLM fundamentals are prior knowledge. Do not announce VLMs as the next lecture. The standard original SAM spatial interface does not directly parse this sentence. Introduce a separate language-to-region stage. The phrase-to-dog association here is an authored demonstration and does not establish relational understanding.
06 · Language grounding
Language can supply the class representation
CLIP-like comparison, now associated with spatial regions.
Authored schematic / numerical teaching example.
Radford et al. (2021), Learning Transferable Visual Models, §2
46What changes when a text encoder supplies the comparison vector?
Teaching notes
Recall the fixed-vocabulary region classifier from detection. Replace its fixed class vector with a supplied text representation in a compatible learned space. The dot-product expression is conceptual; normalized features give cosine similarity and practical systems may add scaling or bias. Grounding DINO is not just crop-by-crop CLIP. Relational phrases require contextual cross-modal mechanisms and appropriate training, not only independent noun similarity.
06 · Language grounding
Phrase grounding asks which region the words refer to
A grounded box gives a spatial referent for the phrase.
Course image and authored illustrative targets · No measured model predictions.
47How is this different from assigning one label to the whole image?
Teaching notes
Use the recurring campus scene and the phrase “the dog beside the cyclist”. Reveal the authored box around the dog. The example is a stored illustrative association, not a measured prediction. A successful demonstration on a scene with one dog would not by itself prove understanding of “beside”.
06 · Language grounding
Visual and language features must meet spatially
Phrase-region interaction produces text-grounded object boxes.
Authored schematic / numerical teaching example.
48Which output must be retained instead of pooling everything into one image score?
Teaching notes
Build the image features, language features, cross-modal interaction, then grounded boxes. Name a Grounding-DINO-style model only on the final build. The paper uses a feature enhancer, language-guided query selection, and a cross-modality decoder. This high-level drawing abstracts those modules and does not claim that a single dot product implements the entire model.
06 · Language grounding
A grounded box becomes a spatial prompt for SAM
Language defines the target. Detection locates it. SAM refines its shape.
Course image and authored illustrative targets · No measured model predictions.
IDEA Research, Grounded Segment Anything, official implementation
49Does SAM need to read the sentence in this composed system?
Teaching notes
Trace both image paths: grounding needs the image and the text; SAM needs the same image and the resulting box. The interface between the systems is spatial geometry, not an assumed shared text embedding. The output shown is illustrative, not a measured Grounded SAM result. Errors can propagate: a wrong grounded box can produce a precise mask of the wrong object. The lab reuses these exact stored boxes and masks.
07 · Choose the output
Choose the output before the architecture
The question determines what the model must preserve.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
50The question determines what the model must preserve.
Teaching notes
Section 07. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.
07 · Choose the output
Choose the question, then the target
The target determines what the model must preserve.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
51Which outputs need class identities, and which can be useful without them?
Teaching notes
Return to the same image and reveal global, instance, dense, geometry, spatial-prompt, and language-prompt questions. These are questions about one input, not an architecture catalogue. The four-family strip remains a guide to output organization and conditioning.
07 · Choose the output
Which internal variables should a person understand?
Next: Concept Bottleneck Models
Course image and authored illustrative targets · No measured model predictions.
Koh et al. (2020), Concept Bottleneck Models
52We can use language to select a region. Can we also inspect the concepts that support a decision?
Teaching notes
This is the new course-order ending. Students already know VLMs. Spatial grounding answers where to look; the next question is whether the decision itself can use inspectable semantic variables. Introduce concept bottlenecks as the next topic without claiming that a region mask alone explains a model’s reasoning.
Optional reference
Optional: optical flow needs two frames
Flow predicts a 2D displacement in the image plane for each source pixel.
Authored schematic / numerical teaching example.
A1Is one RGB frame sufficient to observe optical flow?
Teaching notes
The two blue rectangles form a synthetic example, translated rightward. For a source coordinate (x, y), flow (u, v) points toward (x + u, y + v) in the next frame. This is image-plane displacement, not necessarily 3D scene flow or object velocity in metres per second. Occlusion, motion blur, and large displacement make estimation difficult.
Optional reference
Optional: IoU and Dice on the same mask
Both metrics measure overlap, with different normalizations.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
A2Why does Dice give a different number for the same masks?
Teaching notes
This slide uses hard-mask metrics. Soft Dice replaces hard counts with probability-based sums during training; a precise smoothing and absent-class convention is needed. Do not assume that maximizing pixel accuracy, Dice, and IoU gives identical gradients. The existing segmentation notebook remains available separately and is unchanged.
Optional reference
Optional: common channels-first tensor shapes
Tensor layout is an implementation convention; the target meaning stays the same.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
A3Why does the cross-entropy target omit the C dimension?
Teaching notes
B is batch size. R is the number of sampled or retained regions, not a semantic class count. m is local mask resolution, and h × w is heatmap resolution. The lecture’s main diagrams use channels last to emphasize a score vector at each spatial position. PyTorch commonly uses channels first. The original Mask R-CNN channel count and foreground/background class conventions depend on the implementation.
Optional reference
A single view can leave absolute scale ambiguous
Image appearance alone may be compatible with several physical scales.
Course image and authored illustrative targets · No measured model predictions.
Eigen et al. (2014), Depth Map Prediction, §3
A4What happens if we double both the object size and its distance?
Teaching notes
The ray diagram is a geometric illustration. A monocular model can learn useful scale priors from data, object sizes, and camera information. Those priors do not make absolute scale observable in every arbitrary scene. Explain the ambiguity before promising distances in metres. The diagram shows the same angular size for two different size-distance combinations.
Optional reference
Which target would answer each question?
Choose the required output before choosing an architecture.
Course image and authored illustrative targets · No measured model predictions.
ES 667 Deep Learning · Nipun Batra
A5What targets would you choose for road area, person count, and an arm gesture?
Teaching notes
Let students answer before revealing the suggestions. There can be more than one valid design. If only a count is required, direct count regression is another possibility. The examples ask which of the lecture’s spatial representations supports the task. Avoid presenting the listed methods as exclusive choices.
Optional reference
Optional: primary reading
Paper links and image provenance are included in the reading notes and README.
Authored schematic / numerical teaching example.
ES 667 Deep Learning · Nipun Batra
A6Which paper should I read for the mechanism I found most interesting?
Teaching notes
All architecture citations link to primary papers. The course-generated campus street is reused read-only from lecture9. Every overlay, toy mask, probability, and diagram is authored for instruction. No new model inference or empirical accuracy claim is presented. See SOURCES.md for the complete bibliography and provenance.
Optional reference
Optional: a VLM can generate coordinate tokens
Generated coordinates still need a defined syntax, scale, and spatial evaluation.
Authored schematic / numerical teaching example.
Bai et al. (2023), Qwen-VL, §2
A7Can a generative multimodal model return a box in its output sequence?
Teaching notes
Qwen-VL is one historical example of a multimodal model with grounding outputs. The displayed <box> example is schematic, not a literal promise about every VLM or a particular tokenizer. A valid token sequence does not guarantee correct spatial grounding. Keep this optional; the main lecture uses the clearer text-to-box-to-SAM composition.
Optional reference
Relative and metric depth answer different questions
The meaning of depth values follows from the supervision and scale convention.
Authored schematic / numerical teaching example.
Ranftl et al. (2022), Robust Monocular Depth Estimation, §3
A8Can a relative-depth model’s raw output be reported as metres?
Teaching notes
Relative predictions can capture ordering or shape while leaving scale unknown. Some methods, including MiDaS-style training, handle scale and shift ambiguity in an inverse-depth representation. Do not generalize that all relative methods have exactly the same invariance. Metric evaluation needs the camera convention, physical units, and any permitted alignment stated explicitly.