Beyond Boxes
Open Beyond Boxes LabInk stays in this browser.

Deep learning · Dense and promptable vision

Beyond Boxes

Dense, Structured, and Promptable Vision

Beyond BoxesWhat should a visionmodel output whena box is not enough?One image.Several kinds of answer.

The question determines the structure of the answer.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

01

What can a box tell us, and what does it leave unspecified?

Teaching notes

Students have already studied object detection, CLIP, and VLM fundamentals. This lecture asks what spatial structure an answer must contain, then combines language grounding with promptable masks. All street overlays are authored illustrations. No model prediction is claimed.

00 · One image, many questions

One image, many questions

One image, many questionsclassboxespixel labelsinstance maskskeypointsdepthprompted region

What should a vision model output when a box is not enough?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

02

What should a vision model output when a box is not enough?

Teaching notes

Section 00. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

00 · One image, many questions

One image, six questions

One image, six questionsWhat is present?class / classesWhat is present?class / classesWhere are the objects?boxesWhat is present?class / classesWhere are the objects?boxesWhat class is every pixel?semantic fieldWhat is present?class / classesWhere are the objects?boxesWhat class is every pixel?semantic fieldWhich pixels form each person?instance masksWhat is present?class / classesWhere are the objects?boxesWhat class is every pixel?semantic fieldWhich pixels form each person?instance masksWhere are body joints?keypointsWhat is present?class / classesWhere are the objects?boxesWhat class is every pixel?semantic fieldWhich pixels form each person?instance masksWhere are body joints?keypointsHow far is each surface?depth

Same input. Different target structure.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

03

Which of these questions needs one answer for every pixel?

Teaching notes

Read the question first, then the target. Classification can mean a single scene label or a vector of category-presence labels. The street contains several categories, so a single object-class softmax would not list everything present. Detection supplies a variable number of records. The examples illustrate task definitions rather than a benchmark annotation policy. Begin with the image alone. Reveal one question and its visual target at a time. The persistent family labels distinguish global outputs, sets, dense fields, and query-conditioned outputs; these are useful organizational views rather than mutually exclusive model classes.

01 · Label every pixel

Label every pixel

Label every pixelclassboxespixel labelsinstance maskskeypointsdepthprompted region

Which pixels are actually dog, road, or person?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

04

Which pixels are actually dog, road, or person?

Teaching notes

Section 01. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

01 · Label every pixel

A box does not describe the silhouette

A box does not describe the silhouetteWhere is the dog?Which pixels belong to it?Where is the dog?Which pixels belong to it?A rectangle also contains road and empty space.A mask describes the visible shape.

A box gives approximate extent. Segmentation specifies pixel membership.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

05

Could we estimate the dog’s visible area from its bounding box?

Teaching notes

Reveal the mask after discussing the box. The teal outline is a manually authored, coarse silhouette. It is neither official ground truth nor a model prediction. Visible-region segmentation excludes occluded parts unless the task explicitly asks for an amodal mask.

01 · Label every pixel

Semantic segmentation labels the image grid

Semantic segmentation labels the image gridskyroadpersondogcarSelected regions shown.

Each valid pixel receives one class label under the chosen annotation scheme.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

06

Would two people receive different class labels?

Teaching notes

Use the road and sky first, then the people. Only selected regions are coloured so the source image remains legible. This is not an exhaustive semantic annotation: buildings, vegetation, bicycle, and a distant pedestrian are deliberately left unpainted. Colours encode class identity in this illustration.

01 · Label every pixel

One pixel has one target and C competing scores

One pixel has one target and C competing scores0110111011002222Which class isthis pixel?0110111011002222Which class isthis pixel?Target: person
yij=persony_{ij}=\text{person}
0110111011002222Which class isthis pixel?Target: person
yij=persony_{ij}=\text{person}
pixel feature
fij∈Rdf_{ij}\in\mathbb{R}^{d}
0110111011002222Which class isthis pixel?Target: person
yij=persony_{ij}=\text{person}
pixel feature
fij∈Rdf_{ij}\in\mathbb{R}^{d}
sky −2.303person −0.357road −1.609C class logits

The network produces scores; supervision supplies the correct class.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

07

What is the difference between yᵢⱼ and zᵢⱼ?

Teaching notes

Start with the selected pixel on the foreground person in the street. Reveal its target class, then the feature vector at that location, then C logits. The magnified pixel patch is an authored semantic illustration around that same location; it is not a dataset annotation. The classifier sees contextual features, not only one RGB triplet.

01 · Label every pixel

The pixel classifier shares its weights

The pixel classifier shares its weightspixel feature
fijf_{ij}
shared classifier
zij=Wfij+bz_{ij}=Wf_{ij}+b
One probability distribution per pixel.
pixel feature
fijf_{ij}
shared classifier
zij=Wfij+bz_{ij}=Wf_{ij}+b
softmax over C.10 sky.70 person.20 road
∑c=1Cpij(c)=1\sum_{c=1}^{C}p_{ij}(c)=1
One probability distribution per pixel.

Softmax normalizes over classes separately at each location.

Authored schematic / numerical teaching example.

Long et al. (2015), Fully Convolutional Networks, §4

08

Should probabilities across all image pixels sum to one?

Teaching notes

The same learned classifier operates at every spatial location. Softmax normalizes over C classes at each pixel separately. The illustrative logits log(0.10), log(0.70), and log(0.20) give probabilities exactly 0.10, 0.70, and 0.20. This keeps one numerical example throughout the pixel trace.

01 · Label every pixel

Cross-entropy asks for the target class probability

Cross-entropy asks for the target class probabilityTarget: personsky0.10person0.70road0.20Target: personsky0.10person0.70road0.20
ℓij=−log⁡(0.70)=0.357\ell_{ij}=-\log(0.70)=0.357
Use the target class probability.
Target: personsky0.10person0.70road0.20
ℓij=−log⁡(0.70)=0.357\ell_{ij}=-\log(0.70)=0.357
Use the target class probability..357Place this loss back at the selected pixel.

A confident mistake receives a larger penalty than a correct high-probability answer.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

09

Which of the three probabilities enters the loss for this pixel?

Teaching notes

Hold the target person and probabilities 0.10, 0.70, 0.20 fixed. Reveal the negative log calculation, then draw its 0.357 value back onto the selected location in the schematic loss grid. Other cells remain blank until the image-loss slide. The calculation is illustrative and uses natural logarithms.

01 · Label every pixel

The image loss averages valid pixel losses

The image loss averages valid pixel losses0.080.120.210.060.180.3570.090.170.230.140.050.110.080.320.1One loss per labelled pixelCrossed cell: ignored0.080.120.210.060.180.3570.090.170.230.140.050.110.080.320.1One loss per labelled pixelCrossed cell: ignored
L=1∣V∣∑(i,j)∈Vℓij\mathcal{L}={1\over |V|}\sum_{(i,j)\in V}\ell_{ij}
ℓij=−log⁡pij(yij)\ell_{ij}=-\log p_{ij}(y_{ij})
V contains valid annotated pixels.

Semantic segmentation applies shared classification over many spatial locations.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

10

If a dataset marks a boundary as unknown, should we force it to be background?

Teaching notes

Show one loss per valid pixel. The highlighted pixel retains the same 0.357 value from the previous slide. A grey crossed cell is ignored. Reveal the mean over V only after the students see the spatial loss field. Raw logits and integer targets are the usual cross-entropy inputs; argmax is not part of training.

01 · Label every pixel

A spatial pipeline produces one classifier input per pixel

A spatial pipeline produces one classifier input per pixelimage
I∈RH×W×3I\in\mathbb{R}^{H\times W\times3}
image
I∈RH×W×3I\in\mathbb{R}^{H\times W\times3}
encoder
F∈Rh×w×dF\in\mathbb{R}^{h\times w\times d}
coarse semantic features
image
I∈RH×W×3I\in\mathbb{R}^{H\times W\times3}
encoder
F∈Rh×w×dF\in\mathbb{R}^{h\times w\times d}
decoder
F′∈RH×W×d′F^{\prime}\in\mathbb{R}^{H\times W\times d^{\prime}}
coarse semantic featuresfeatures on the target grid
image
I∈RH×W×3I\in\mathbb{R}^{H\times W\times3}
encoder
F∈Rh×w×dF\in\mathbb{R}^{h\times w\times d}
decoder
F′∈RH×W×d′F^{\prime}\in\mathbb{R}^{H\times W\times d^{\prime}}
shared pixel classifier
Z∈RH×W×CZ\in\mathbb{R}^{H\times W\times C}
coarse semantic featuresfeatures on the target gridC logits at every pixel

The encoder supplies context; the decoder returns features to the target grid.

Authored schematic / numerical teaching example.

Long et al. (2015), Fully Convolutional Networks, §4

11

What would global average pooling discard?

Teaching notes

The encoder computes contextual features on a smaller grid. A spatial decoder increases resolution and combines features. A final class head maps each location to C logits. This is a schematic, not a particular implementation: outputs may first be produced at a lower resolution and then resized to the target grid.

02 · Recover spatial detail

How do we recover spatial detail?

How do we recover spatial detail?classboxespixel labelsinstance maskskeypointsdepthprompted region

Where could the missing boundary information come from?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

12

Where could the missing boundary information come from?

Teaching notes

Section 02. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

02 · Recover spatial detail

Upsampling increases resolution, not evidence

Upsampling increases resolution, not evidenceimageencoderdecoderpixel classifierFine maskThe four coarse values do not identify the original boundary.imageencoderdecoderpixel classifierFine maskDownsampleThe four coarse values do not identify the original boundary.imageencoderdecoderpixel classifierFine maskDownsampleUpsampleIs a largergrid enough?The four coarse values do not identify the original boundary.

Several different fine boundaries can produce the same coarse summary.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

13

Can interpolation recover the exact original shape from these four values?

Teaching notes

This is an explicit toy compression, not a claim that a CNN uses average pooling at every stage. The 8 × 8 binary mask is averaged within four 4 × 4 blocks. The coarse values are 0.5, 0, 0.5, and 0.125. Nearest-neighbour upsampling repeats them. A learned decoder can exploit training priors, but it cannot guarantee recovery of arbitrary detail discarded from its inputs.

02 · Recover spatial detail

Where could the missing detail come from?

Where could the missing detail come from?imageencoderdecoderpixel classifierdeep semanticfeaturesupsamplerefined decoder featuresimageencoderdecoderpixel classifierdeep semanticfeaturesupsamplerefined decoder featuresfine encoder featuresimageencoderdecoderpixel classifierdeep semanticfeaturesupsamplerefined decoder featuresfine encoder featuresskip connectionconcatenate, then convolve

A skip connection brings fine encoder features into the decoder.

Authored schematic / numerical teaching example.

Ronneberger et al. (2015), U-Net, §2

14

What information does the long connection add?

Teaching notes

First trace the deep path through upsampling. Then reveal the skip from the fine encoder stage. Match the spatial dimensions before fusion. U-Net concatenates the encoder and decoder feature channels and applies convolutions. This is a feature concatenation, unlike an additive residual connection. The network still has to learn how to use the features. Do not name the connection on its first build. First ask where detail could come from, then reveal the fine feature branch, then name the skip and the fusion operation.

02 · Recover spatial detail

U-Net repeats this combination across scales

U-Net repeats this combination across scalesimageencoderdecoderpixel classifier
H×WH\times W
H/2×W/2H/2\times W/2
H/4×W/4H/4\times W/4
Encoder: contextDecoder: resolution and detail
imageencoderdecoderpixel classifier
H×WH\times W
H/2×W/2H/2\times W/2
H/4×W/4H/4\times W/4
H/4×W/4H/4\times W/4
H/2×W/2H/2\times W/2
H×WH\times W
Encoder: contextDecoder: resolution and detail
imageencoderdecoderpixel classifier
H×WH\times W
H/2×W/2H/2\times W/2
H/4×W/4H/4\times W/4
H/4×W/4H/4\times W/4
H/2×W/2H/2\times W/2
H×WH\times W
copy aligned featuresEncoder: contextDecoder: resolution and detail

At each decoder scale, aligned encoder features help refine the prediction.

Authored schematic / numerical teaching example.

Ronneberger et al. (2015), U-Net, §2

15

Which arrows provide context, and which preserve finer detail?

Teaching notes

Reveal the skip paths after following the U-shaped computation. This drawing omits channel counts, convolution blocks, and the final class head for clarity. Original U-Net used valid convolutions and cropped skip tensors before concatenation; modern variants often use padding. The conceptual claim is aligned feature fusion, not exact original tensor dimensions.

02 · Recover spatial detail

FCN also combines coarse and fine predictions

FCN also combines coarse and fine predictionsimageencoderdecoderpixel classifiershallow, fine featuresdeep, coarse featuresclass score mapclass score mapupsample+fusedscoresFCN fuses score maps. U-Net concatenates features before further convolutions.

The common problem is precise localization from a hierarchy of features.

Authored schematic / numerical teaching example.

Long et al. (2015), Fully Convolutional Networks, §4

16

What does this FCN diagram fuse?

Teaching notes

Long et al. showed how a classification network could produce dense predictions using convolutional score maps and learned upsampling. The skip variants combine coarse semantic scores with finer layers. The drawing summarizes score fusion and omits the final upsampling step. This is a short connection, not an architecture survey.

02 · Recover spatial detail

Inference turns class scores into a label map

Inference turns class scores into a label mapclass logits
Z∈RH×W×CZ\in\mathbb{R}^{H\times W\times C}
[0.10,  0.70,  0.20]⟶person[0.10,\;0.70,\;0.20]\quad\longrightarrow\quad\text{person}
class logits
Z∈RH×W×CZ\in\mathbb{R}^{H\times W\times C}
argmax over classesone class ID per pixel
Y^∈{1,…,C}H×W\hat{Y}\in\{1,\ldots,C\}^{H\times W}
[0.10,  0.70,  0.20]⟶person[0.10,\;0.70,\;0.20]\quad\longrightarrow\quad\text{person}

The decoded output has one class ID per pixel.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

17

Where did the C dimension go?

Teaching notes

The probabilities can be retained for confidence analysis, but the hard semantic prediction is an integer map. Argmax of logits gives the same winning class as argmax of softmax probabilities. Training uses differentiable scores; hard labels are convenient for display and discrete overlap metrics.

02 · Recover spatial detail

Mask IoU measures foreground agreement

Mask IoU measures foreground agreementTarget G0110011001100000Prediction P0110111001010100Target G0110011001100000Prediction P0110111001010100Overlap and errorsTPTPFPTPTPTPFNFPFPTarget G0110011001100000Prediction P0110111001010100Overlap and errorsTPTPFPTPTPTPFNFPFPTP = 5FP = 3FN = 1Target G0110011001100000Prediction P0110111001010100Overlap and errorsTPTPFPTPTPTPFNFPFPTP = 5FP = 3FN = 1
IoU⁡=TPTP+FP+FN=59≈0.556\operatorname{IoU}={TP\over TP+FP+FN}={5\over9}\approx0.556

Five pixels agree on foreground; four more belong to only one mask.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

18

What are the intersection and union sizes?

Teaching notes

Use the exact same 4 × 4 masks throughout the builds. Ground truth has 6 foreground pixels; prediction has 8. First ask students to count overlap, then reveal errors, then show the quotient. Background agreement is outside the foreground union. These are synthetic labels and predictions, not measured model outputs. Use green for TP, coral for FP, and amber for FN. Build 3 shows counts; build 4 calculates the quotient.

02 · Recover spatial detail

Background can dominate pixel accuracy

Background can dominate pixel accuracy6 foreground pixelspredict backgroundeverywhereForeground overlap and class-wise IoU expose this failure.6 foreground pixelspredict backgroundeverywhere94% accuracyForeground IoU = 0Foreground overlap and class-wise IoU expose this failure.

A high accuracy can coexist with complete failure on a small foreground region.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

19

Is an all-background prediction useful for finding these six pixels?

Teaching notes

The miniature 10 × 10 image has exactly six foreground pixels. Explain why accuracy is dominated by the 94 background pixels. Mean IoU averages class-wise overlap, subject to the benchmark’s absent-class convention. Report the class of interest and define which pixels and classes participate.

03 · Which instance?

Which pixels belong to which instance?

Which pixels belong to which instance?classboxespixel labelsinstance maskskeypointsdepthprompted region

The class is person. Which person?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

20

The class is person. Which person?

Teaching notes

Section 03. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

03 · Which instance?

A semantic class can contain several individuals

A semantic class can contain several individualsAll three have class personWhich class is this pixel? Which object does this pixel belong to?All three have class personEach person also has an instance ID123Which class is this pixel? Which object does this pixel belong to?

Semantic labels give the class. Instance IDs give the individual.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

21

Could connected components always recover the people from a semantic map?

Teaching notes

The selected foreground pedestrians and cyclist are all people. The semantic illustration gives them the same colour. The reveal assigns different IDs through different colours. It does not change their shared class. A distant pedestrian is not annotated in this selected-region illustration. Avoid interpreting the coloured sample as an exhaustive instance set.

03 · Which instance?

Each instance has its own binary mask

Each instance has its own binary maskRecall one detection record
(ci,bi)(c_i,b_i)
N varies by image. The object records are unordered.
Recall one detection record
(ci,bi)(c_i,b_i)
{(ci,bi,Mi)}i=1N\{(c_i,b_i,M_i)\}_{i=1}^{N}
Mi∈{0,1}H×WM_i\in\{0,1\}^{H\times W}
N varies by image. The object records are unordered.

The new field in the familiar object record is the binary mask.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

22

Does mask number 1 always mean the same person across different images?

Teaching notes

An instance record can include a class, a box, and an image-sized binary mask. Boxes are useful for the detector-based architecture we will use. Some methods predict masks without a separate box stage. Unlike semantic class channels, the instance dimension N is variable. This reconnects briefly with Part I’s set target without repeating candidate assignment. Show the old class-and-box record first, then append the mask M_i. The records remain an unordered variable-size set.

03 · Which instance?

What extra detector output would describe shape?

What extra detector output would describe shape?region featuresclassboxregion featuresclassboxWhat output would describe the shape?region featuresclassboxmasknew output

A mask branch extends the detector’s object record.

Authored schematic / numerical teaching example.

He et al. (2017), Mask R-CNN, §3

23

The detector already predicts a class and a box. What is missing?

Teaching notes

Begin with the region feature and class/box outputs already familiar from object detection. Ask for the additional output, then reveal a local mask. Only after this motivation show how region features are aligned and how Mask R-CNN implements the branches.

03 · Which instance?

A detector can provide features for one region

A detector can provide features for one regionImage feature map + proposalSample aligned features; preserve the spatial layout within the region.Image feature map + proposalRoIAlignfixed local feature gridSample aligned features; preserve the spatial layout within the region.

A region feature grid preserves spatial layout inside a proposed object.

Authored schematic / numerical teaching example.

He et al. (2017), Mask R-CNN, §3

24

Why does the mask branch need a grid instead of a single pooled vector?

Teaching notes

A detector supplies proposed regions over a shared image feature map. RoIAlign samples features using bilinear interpolation without the coordinate rounding of RoIPool. It produces fixed spatial dimensions for different region sizes. The schematic follows one dog proposal. Feature extraction is shared across regions.

03 · Which instance?

Mask R-CNN adds a mask branch to region detection

Mask R-CNN adds a mask branch to region detectionimage features+ proposalsRoIAlignregion featuresclass scoresimage features+ proposalsRoIAlignregion featuresclass scoresbox refinementimage features+ proposalsRoIAlignregion featuresclass scoresbox refinementmask logits

Shared region features support classification, box refinement, and a local mask.

Authored schematic / numerical teaching example.

He et al. (2017), Mask R-CNN, §3

25

Which branch decides the shape of the object?

Teaching notes

Reveal the class branch, the box branch, then the mask branch. The mask branch is a convolutional spatial predictor running in parallel with the class and box heads. It is not obtained by rasterizing the predicted box. The full system also includes the region proposal network and detector losses, already familiar from the object detection lecture.

03 · Which instance?

A local mask is placed back into the image

A local mask is placed back into the imageLocal binary mask
28×2828\times28
Local binary mask
28×2828\times28
resize to box
Local binary mask
28×2828\times28
resize to boxplace in full image

The mask grid predicts shape. The box supplies image location.

Course image and authored illustrative targets · No measured model predictions.

He et al. (2017), Mask R-CNN, §3

26

Must every proposed region directly predict an H × W array?

Teaching notes

The displayed 28 × 28 binary mask is rasterized from the same authored dog polygon and box used by the lab; it is not a measured model output. Original Mask R-CNN commonly uses a 28 × 28 mask grid. At inference, select the predicted class mask, resize it to the detected box, and threshold to obtain a binary mask. Align resizing and coordinate conventions with the implementation.

03 · Which instance?

The mask loss learns binary membership

The mask loss learns binary membershipA matched positive region has three supervised outputs.classboxbinary mask
L=Lclass+Lbox+Lmask\mathcal{L}=\mathcal{L}_{\mathrm{class}}+\mathcal{L}_{\mathrm{box}}+\mathcal{L}_{\mathrm{mask}}
The mask term learns foreground membership within the matched region.

The mask branch answers foreground or background within the matched region.

Authored schematic / numerical teaching example.

He et al. (2017), Mask R-CNN, §3

27

Does the mask head need a softmax across the C class channels?

Teaching notes

Mask R-CNN separates class prediction from mask prediction. For each positive matched region, the target class chooses the channel for the mask loss. At inference the predicted class chooses it. Background regions do not receive an object mask loss. Loss terms shown are schematic unit-weight sums; the paper also includes proposal training. Keep class-specific masks distinct from one semantic softmax per pixel. Keep the main explanation to the three losses on a matched positive region; class-channel mechanics belong in these notes.

03 · Which instance?

Panoptic segmentation combines things and stuff

Panoptic segmentation combines things and stuffThingscountable objectsclass + instance IDStuffroad, sky, vegetationsemantic regionA panoptic map gives each valid pixel a class and, for things, an instance ID.

Countable objects receive instance IDs; background regions receive semantic labels.

Course image and authored illustrative targets · No measured model predictions.

Kirillov et al. (2019), Panoptic Segmentation, §3

28

Does the road need one object ID for every separate visible patch?

Teaching notes

One slide is sufficient. A panoptic output assigns each valid pixel a semantic class and an instance identifier for things, using a single non-overlapping partition. Stuff does not require object instance identity. The selected-region overlay illustrates the distinction; it is not a complete benchmark panoptic map. The things/stuff taxonomy depends on the dataset.

04 · Pose and depth

Other structured outputs: pose and depth

Other structured outputs: pose and depthclassboxespixel labelsinstance maskskeypointsdepthprompted region

The output can describe geometry instead of a class or mask.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

29

The output can describe geometry instead of a class or mask.

Teaching notes

Section 04. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

04 · Pose and depth

A person’s pose requires internal geometry

A person’s pose requires internal geometryarm downarm raisedThe same box can contain different joint arrangements.

Named joints describe an arrangement that a bounding box cannot capture.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

30

Can two people have similar boxes but different poses?

Teaching notes

Use the recurring street for context, then compare the two schematic skeletons. Amber circles are joints and blue segments show known anatomical connections. The photo overlay is authored and approximate. Keypoint annotations are image coordinates, not automatically a 3D body model.

04 · Pose and depth

Pose can be predicted as coordinates

Pose can be predicted as coordinatesK joints in a defined order1 head, 2 shoulder, …, K ankleCoordinates use the person crop’s frame.K joints in a defined order1 head, 2 shoulder, …, K ankleCoordinates use the person crop’s frame.
[(x1,y1),…,(xK,yK)]∈RK×2[(x_1,y_1),\ldots,(x_K,y_K)]\in\mathbb{R}^{K\times2}
Use validity flags for unknown joints.

Each joint has a defined identity, coordinate system, and annotation-validity flag.

Authored schematic / numerical teaching example.

Toshev & Szegedy (2014), DeepPose, §3

31

Should an unlabelled wrist be trained toward coordinate (0, 0)?

Teaching notes

DeepPose is a direct coordinate-regression example. The numerical crop calculation is illustrative, not measured from the street image or skeleton. Normalized x is multiplied by crop width and normalized y by crop height, then the crop offset can be added to recover image coordinates. Visibility and validity are related but not identical. A common coordinate loss is squared error over supervised joints.

04 · Pose and depth

A heatmap represents possible joint locations

A heatmap represents possible joint locationsheadleft wristright wrist
one person:h×w×K heatmap scores\text{one person:}\quad h\times w\times K\text{ heatmap scores}
headleft wristright wristpeak
(xk,yk)(x_k,y_k)
one person:h×w×K heatmap scores\text{one person:}\quad h\times w\times K\text{ heatmap scores}

The output contains one spatial map for each joint type.

Authored schematic / numerical teaching example.

Newell et al. (2016), Stacked Hourglass Networks, §3

32

What does the K dimension mean here?

Teaching notes

Targets often place a Gaussian bump around an annotated coordinate; a common objective is squared error between predicted and target heatmaps. The displayed maps are illustrative. The hard peak can be decoded with argmax, with optional subpixel refinement or a soft-argmax method. Heatmap dimensions h × w may be smaller than the image and require coordinate rescaling. In a top-down pipeline, predict K maps per detected person.

04 · Pose and depth

Dense prediction can be continuous

Dense prediction can be continuoussky: no finite target here
D∈RH×WD\in\mathbb{R}^{H\times W}
A real value at eachvalid spatial location.Illustrative relative ordering: darker means nearer. Sky is invalid.

Depth assigns a number to each valid pixel instead of choosing a class.

Authored schematic / numerical teaching example.

Eigen et al. (2014), Depth Map Prediction, §3

33

Would a softmax over road, person, and sky produce a distance?

Teaching notes

The grey map only illustrates near/far ordering. It is not a predicted or physically measured depth map and does not show realistic surface variation. Darker means nearer here; other visualizations reverse the convention. Depth often means camera-axis z distance rather than Euclidean range along a ray. The dataset must define it. Sky may not have a finite valid depth label.

04 · Pose and depth

A depth loss compares numerical values

A depth loss compares numerical valuessky: no finite target hereValid depth measurements define the target.This is dense regression, with a validity mask.sky: no finite target hereValid depth measurements define the target.
L=1∣V∣∑(i,j)∈V∣d^ij−dij∣\mathcal{L}={1\over|V|}\sum_{(i,j)\in V}|\hat{d}_{ij}-d_{ij}|
V excludes invalid measurements.Units follow the dataset convention.This is dense regression, with a validity mask.

A dense regression target needs units and a valid-pixel mask.

Authored schematic / numerical teaching example.

Eigen et al. (2014), Depth Map Prediction, §3

34

Which positions should contribute if a depth sensor returns no measurement?

Teaching notes

The L1 expression is a simple teaching loss, not a claim that the cited Eigen system uses only L1. Depth methods may operate in log depth or inverse depth and may include scale-invariant, gradient, or geometric objectives. The purpose is to distinguish categorical per-pixel cross-entropy from continuous per-pixel regression.

04 · Pose and depth

Depth values need a scale convention

Depth values need a scale conventioncamerasmall / nearlarge / farRelative depthnearer / farther orderingabsolute scale may be unknownScale ambiguity: changing size and distance together can preserve the image.camerasmall / nearlarge / farRelative depthnearer / farther orderingabsolute scale may be unknownMetric depthdistance in a physical unitrequires a calibrated scale conventionScale ambiguity: changing size and distance together can preserve the image.

Relative depth gives ordering; metric depth requires physical scale.

Authored schematic / numerical teaching example.

Ranftl et al. (2022), Robust Monocular Depth Estimation, §3

35

Can a visually plausible depth map automatically be read in metres?

Teaching notes

The ray diagram is a geometric illustration. A monocular model can learn useful scale priors from data, object sizes, and camera information. Those priors do not make absolute scale observable in every arbitrary scene. Explain the ambiguity before promising distances in metres. The diagram shows the same angular size for two different size-distance combinations. Relative predictions can capture ordering or shape while leaving scale unknown. Some methods, including MiDaS-style training, handle scale and shift ambiguity in an inverse-depth representation. Do not generalize that all relative methods have exactly the same invariance. Metric evaluation needs the camera convention, physical units, and any permitted alignment stated explicitly.

04 · Pose and depth

The answer can have several kinds of structure

The answer can have several kinds of structureGLOBALclass / scoreSETboxes / object masksDENSEpixel labels / depthQUERY-CONDITIONEDselected region / maskKeypoints: a structured tuple per person; a set across people.GLOBALclass / scoreSETboxes / object masksDENSEpixel labels / depthQUERY-CONDITIONEDselected region / maskKeypoints: a structured tuple per person; a set across people.A query can condition which set or dense region the model returns.

What if the output also depends on what the user is asking for?

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

36

How do keypoints fit beside a class, a set of masks, and a dense field?

Teaching notes

Use four persistent families: global, set, dense, query-conditioned. The first three describe output organization. The fourth adds conditioning and can produce a set or dense mask. Keypoints are structured geometry: K named coordinates per person, and a set across people. Query conditioning is an additional design choice, not a separate tensor layout.

04 · Pose and depth

Now the tensor shapes have a visual meaning

Now the tensor shapes have a visual meaningTASKTARGET MEANINGOUTPUT REPRESENTATIONclassificationclass ID or binary labels
C scoresC\text{ scores}
detectionunordered object records
N×(class+box)N\times(\text{class}+\text{box})
semanticone class ID per pixel
Z∈RH×W×CZ\in\mathbb{R}^{H\times W\times C}
instancesone binary mask per object
{(ci,bi,Mi)}i=1N\{(c_i,b_i,M_i)\}_{i=1}^{N}
keypointsK named coordinates
K×2 per personK\times2\ \text{per person}
depthone value per valid pixel
D∈RH×WD\in\mathbb{R}^{H\times W}

A label map has H × W entries. Its classifier produces C scores at each entry.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

37

Does semantic segmentation require C ground-truth numbers per pixel?

Teaching notes

Consolidate the output notation after the visual examples. H and W refer to the chosen output grid, C to semantic classes, N to the number of annotated instances, and K to joint types. The table uses channels last and omits batch size. Detection and instance segmentation have variable target sets even if their networks allocate fixed candidate capacity. The mask implementation extends Part I’s object records. This synthesis appears only after students have seen all unprompted target types visually. Query-conditioned mask notation is introduced in the next section.

05 · Promptable vision

What if the user selects the region?

What if the user selects the region?classboxespixel labelsinstance maskskeypointsdepthprompted region

The image contains many regions. Which one is being requested?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

38

The image contains many regions. Which one is being requested?

Teaching notes

Section 05. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

05 · Promptable vision

Which region do I want a mask for?

Which region do I want a mask for?Which region do I want?+Which region do I want?+Which region do I want?+
M=f(I,q)M=f(I,q)

Image + query produces a spatial answer.

Course image and authored illustrative targets · No measured model predictions.

Kirillov et al. (2023), §3

39

Must a segmentation model return every object when we only want the dog?

Teaching notes

Show the point first, then the illustrative dog mask. The click conveys a region of interest without assigning a semantic class name. Both the prompt and the mask are authored teaching elements. This starts a different interface: image plus query produces a spatial answer.

05 · Promptable vision

Spatial prompts express different constraints

Spatial prompts express different constraints+Positive pointinclude this area+−Positive pointinclude this areaNegative pointexclude this area+−Positive pointinclude this areaNegative pointexclude this areaBoxrestrict the region of interest+−Positive pointinclude this areaNegative pointexclude this areaBoxrestrict the region of interestPrevious maskrefine an earlier segmentation

A point, box, or previous mask gives the model information about the requested region.

Course image and authored illustrative targets · No measured model predictions.

Kirillov et al. (2023), §3

40

What could a negative point add after a positive point?

Teaching notes

The standard released original SAM interface supports positive and negative points, boxes, and mask inputs. The right panel represents a prior mask; the actual dense mask prompt is low resolution. These inputs can be combined. No class name is required for this interface. The examples are authored, not live SAM runs. Reveal positive point, negative point, box, and previous mask in that order. The negative point lies on the road beside the dog; it marks a region to exclude.

05 · Promptable vision

A promptable model combines two sources of information

A promptable model combines two sources of informationimageimage encoderSAMimageimage encodercached embeddingSAMimageimage encodercached embeddingpoint / box /mask promptprompt encoderSAMimageimage encodercached embeddingpoint / box /mask promptprompt encodermask decoderSAMimageimage encodercached embeddingpoint / box /mask promptprompt encodermask decodercandidate masksSAMimageimage encodercached embeddingpoint / box /mask promptprompt encodermask decodercandidate masksquality estimateSAM

A mask decoder uses visual features together with the encoded prompt.

Authored schematic / numerical teaching example.

Kirillov et al. (2023), Segment Anything, §3 and Appendix A

41

Which part changes when we click somewhere else in the same image?

Teaching notes

Build the image branch first, then the prompt branch, then the mask decoder. The original SAM uses a ViT image encoder. Sparse points and boxes have positional and type embeddings; dense mask prompts are encoded with convolutions. Its lightweight decoder uses self-attention and two-way cross-attention between prompts/output tokens and image features, then produces mask logits and estimated mask quality. This is a high-level view; details of upscaling and dynamic mask prediction are omitted. Reveal image encoder, cached embedding, prompt encoder, mask decoder, candidate masks, and predicted quality in six steps. The model name is introduced here only after the prompt interface has been motivated.

05 · Promptable vision

One prompt can admit several valid masks

One prompt can admit several valid masks+Does this point mean the heador the whole dog?+Does this point mean the heador the whole dog?partwhole object

Multiple candidate masks can express ambiguity in the prompt.

Course image and authored illustrative targets · No measured model predictions.

Kirillov et al. (2023), Segment Anything, §3 and Appendix A

42

Does a click on the dog’s head specify only the head or the whole dog?

Teaching notes

The point is placed on the dog’s head. The alternatives are conceptual, not SAM predictions. Original SAM can produce multiple masks for ambiguous prompts, together with predicted IoU quality scores. This is an estimated quality score, not an observed ground-truth IoU or a semantic class probability. Additional prompts can clarify the intended extent.

05 · Promptable vision

Image features can be reused across prompts

Image features can be reused across promptsimage encodercomputed oncecached imageembeddingReuse the expensive representation; change the requested region.image encodercomputed oncecached imageembeddingnew point / box / maskprompt encoder +lightweight mask decodermaskReuse the expensive representation; change the requested region.

Interactive refinement changes the query while keeping the image representation fixed.

Authored schematic / numerical teaching example.

Kirillov et al. (2023), Segment Anything, §3 and Appendix A

43

Why separate the image encoder from the prompt-dependent computation?

Teaching notes

Trace the cached embedding and the new prompt into the decoder. Positive and negative corrections or a prior mask can refine the region. This is inference, not a gradient update to the model weights. SAM predicts masks and their quality; it does not inherently name the semantic category of the mask. Connect this to reusing expensive image representations in CLIP retrieval or cached visual context in a VLM. Reuse is architectural, not a claim that these models share the same embedding space.

06 · Language grounding

What if the query is language?

What if the query is language?classboxespixel labelsinstance maskskeypointsdepthprompted region

How can words become a spatial prompt?

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

44

How can words become a spatial prompt?

Teaching notes

Section 06. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

06 · Language grounding

How can words select a region?

How can words select a region?“the dog besidethe cyclist”SAM takes spatial prompts. A grounding stage must turn this phrase into a region.“the dog besidethe cyclist”Words describe the referent.Where is that referent?SAM takes spatial prompts. A grounding stage must turn this phrase into a region.

Language describes the referent; grounding must locate it.

Course image and authored illustrative targets · No measured model predictions.

Liu et al. (2023), §3

45

Where should the phrase “the dog beside the cyclist” act in the computation?

Teaching notes

CLIP and VLM fundamentals are prior knowledge. Do not announce VLMs as the next lecture. The standard original SAM spatial interface does not directly parse this sentence. Introduce a separate language-to-region stage. The phrase-to-dog association here is an authored demonstration and does not establish relational understanding.

06 · Language grounding

Language can supply the class representation

Language can supply the class representationregion feature
rir_i
fixed classifier weightsfrom detector training
region feature
rir_i
supplied class name / phrasetext representation
tjt_j
region feature
rir_i
supplied class name / phrasetext representation
tjt_j
sij=riTtjs_{ij}=r_i^{\mathsf T}t_j
Comparison requires compatible learned representations.

CLIP-like comparison, now associated with spatial regions.

Authored schematic / numerical teaching example.

Radford et al. (2021), Learning Transferable Visual Models, §2

46

What changes when a text encoder supplies the comparison vector?

Teaching notes

Recall the fixed-vocabulary region classifier from detection. Replace its fixed class vector with a supplied text representation in a compatible learned space. The dot-product expression is conceptual; normalized features give cosine similarity and practical systems may add scaling or bias. Grounding DINO is not just crop-by-crop CLIP. Relational phrases require contextual cross-modal mechanisms and appropriate training, not only independent noun similarity.

06 · Language grounding

Phrase grounding asks which region the words refer to

Phrase grounding asks which region the words refer to“the dog beside the cyclist”“the dog beside the cyclist”Grounding returnsa location forthe referent.

A grounded box gives a spatial referent for the phrase.

Course image and authored illustrative targets · No measured model predictions.

Liu et al. (2023), §3

47

How is this different from assigning one label to the whole image?

Teaching notes

Use the recurring campus scene and the phrase “the dog beside the cyclist”. Reveal the authored box around the dog. The example is a stored illustrative association, not a measured prediction. A successful demonstration on a scene with one dog would not by itself prove understanding of “beside”.

06 · Language grounding

Visual and language features must meet spatially

Visual and language features must meet spatiallyimagevisual featuresimagevisual featurestextlanguage featuresimagevisual featurestextlanguage featurescross-modal interactionimagevisual featurestextlanguage featurescross-modal interactionboxesExample: a Grounding-DINO-style model

Phrase-region interaction produces text-grounded object boxes.

Authored schematic / numerical teaching example.

Liu et al. (2023), §3

48

Which output must be retained instead of pooling everything into one image score?

Teaching notes

Build the image features, language features, cross-modal interaction, then grounded boxes. Name a Grounding-DINO-style model only on the final build. The paper uses a feature enhancer, language-guided query selection, and a cross-modality decoder. This high-level drawing abstracts those modules and does not claim that a single dot product implements the entire model.

06 · Language grounding

A grounded box becomes a spatial prompt for SAM

A grounded box becomes a spatial prompt for SAM“the dog beside the cyclist”grounding model“the dog beside the cyclist”grounding modelbox“the dog beside the cyclist”grounding modelboxSAMsame image“the dog beside the cyclist”grounding modelboxSAMsame imagemask

Language defines the target. Detection locates it. SAM refines its shape.

Course image and authored illustrative targets · No measured model predictions.

IDEA Research, Grounded Segment Anything, official implementation

49

Does SAM need to read the sentence in this composed system?

Teaching notes

Trace both image paths: grounding needs the image and the text; SAM needs the same image and the resulting box. The interface between the systems is spatial geometry, not an assumed shared text embedding. The output shown is illustrative, not a measured Grounded SAM result. Errors can propagate: a wrong grounded box can produce a precise mask of the wrong object. The lab reuses these exact stored boxes and masks.

07 · Choose the output

Choose the output before the architecture

Choose the output before the architectureclassboxespixel labelsinstance maskskeypointsdepthprompted region

The question determines what the model must preserve.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

50

The question determines what the model must preserve.

Teaching notes

Section 07. Pause at the recurring street image. Ask the section question before returning to the mechanisms. The ladder highlights the current output; other outputs remain visible but faded.

07 · Choose the output

Choose the question, then the target

Choose the question, then the targetWhat is present?classWhat is present?classWhere are the people?boxes / instance masksWhat is present?classWhere are the people?boxes / instance masksWhat class is every pixel?semantic fieldWhat is present?classWhere are the people?boxes / instance masksWhat class is every pixel?semantic fieldWhat geometry is present?keypoints / depthWhat is present?classWhere are the people?boxes / instance masksWhat class is every pixel?semantic fieldWhat geometry is present?keypoints / depthWhich region do I mean?prompted maskWhat is present?classWhere are the people?boxes / instance masksWhat class is every pixel?semantic fieldWhat geometry is present?keypoints / depthWhich region do I mean?prompted mask“the dog beside the cyclist”grounded box, then mask

The target determines what the model must preserve.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

51

Which outputs need class identities, and which can be useful without them?

Teaching notes

Return to the same image and reveal global, instance, dense, geometry, spatial-prompt, and language-prompt questions. These are questions about one input, not an architecture catalogue. The four-family strip remains a guide to output organization and conditioning.

07 · Choose the output

Which internal variables should a person understand?

Which internal variables should a person understand?We can ask languagewhich region to look at.Can the internal decisionvariables also behuman-understandable?

Next: Concept Bottleneck Models

Course image and authored illustrative targets · No measured model predictions.

Koh et al. (2020), Concept Bottleneck Models

52

We can use language to select a region. Can we also inspect the concepts that support a decision?

Teaching notes

This is the new course-order ending. Students already know VLMs. Spatial grounding answers where to look; the next question is whether the decision itself can use inspectable semantic variables. Introduce concept bottlenecks as the next topic without claiming that a region mask alone explains a model’s reasoning.

Optional reference

Optional: optical flow needs two frames

Optional: optical flow needs two framesframe tframe t + 1at each source pixel
(u,v)∈R2(u,v)\in\mathbb{R}^{2}
H×W×2H\times W\times2
Input: TWO frames. Output: a 2D displacement vector for each source pixel.Occluded pixels may have no visible correspondence in the second frame.

Flow predicts a 2D displacement in the image plane for each source pixel.

Authored schematic / numerical teaching example.

Teed & Deng (2020), RAFT, §3

A1

Is one RGB frame sufficient to observe optical flow?

Teaching notes

The two blue rectangles form a synthetic example, translated rightward. For a source coordinate (x, y), flow (u, v) points toward (x + u, y + v) in the next frame. This is image-plane displacement, not necessarily 3D scene flow or object velocity in metres per second. Occlusion, motion blur, and large displacement make estimation difficult.

Optional reference

Optional: IoU and Dice on the same mask

Optional: IoU and Dice on the same mask0110011001100000
IoU⁡=TPTP+FP+FN=59\operatorname{IoU}={TP\over TP+FP+FN}={5\over9}
Dice⁡=2TP2TP+FP+FN=1014\operatorname{Dice}={2TP\over2TP+FP+FN}={10\over14}
Soft overlap losses use probabilities; absent classes need a convention.

Both metrics measure overlap, with different normalizations.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

A2

Why does Dice give a different number for the same masks?

Teaching notes

This slide uses hard-mask metrics. Soft Dice replaces hard counts with probability-based sums during training; a precise smoothing and absent-class convention is needed. Do not assume that maximizing pixel accuracy, Dice, and IoU gives identical gradients. The existing segmentation notebook remains available separately and is unchanged.

Optional reference

Optional: common channels-first tensor shapes

Optional: common channels-first tensor shapesRGB image
B×3×H×WB\times3\times H\times W
semantic logits
B×C×H×WB\times C\times H\times W
semantic labels
B×H×WB\times H\times W
local mask logits
R×C×m×mR\times C\times m\times m
keypoint heatmaps
R×K×h×wR\times K\times h\times w
depth
B×1×H×WB\times1\times H\times W

Tensor layout is an implementation convention; the target meaning stays the same.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

A3

Why does the cross-entropy target omit the C dimension?

Teaching notes

B is batch size. R is the number of sampled or retained regions, not a semantic class count. m is local mask resolution, and h × w is heatmap resolution. The lecture’s main diagrams use channels last to emphasize a score vector at each spatial position. PyTorch commonly uses channels first. The original Mask R-CNN channel count and foreground/background class conventions depend on the implementation.

Optional reference

A single view can leave absolute scale ambiguous

A single view can leave absolute scale ambiguouscamerasmall and nearlarge and farsame image extentScaling scene size and distance together can preserve the image.

Image appearance alone may be compatible with several physical scales.

Course image and authored illustrative targets · No measured model predictions.

Eigen et al. (2014), Depth Map Prediction, §3

A4

What happens if we double both the object size and its distance?

Teaching notes

The ray diagram is a geometric illustration. A monocular model can learn useful scale priors from data, object sizes, and camera information. Those priors do not make absolute scale observable in every arbitrary scene. Explain the ambiguity before promising distances in metres. The diagram shows the same angular size for two different size-distance combinations.

Optional reference

Which target would answer each question?

Which target would answer each question?How much of the image is road?How many people are present?Is the person’s arm raised?How much of the image is road?semantic segmentationHow many people are present?detection or instance segmentationIs the person’s arm raised?keypoints / pose

Choose the required output before choosing an architecture.

Course image and authored illustrative targets · No measured model predictions.

ES 667 Deep Learning · Nipun Batra

A5

What targets would you choose for road area, person count, and an arm gesture?

Teaching notes

Let students answer before revealing the suggestions. There can be more than one valid design. If only a count is required, direct count regression is another possibility. The examples ask which of the lecture’s spatial representations supports the task. Avoid presenting the listed methods as exclusive choices.

Optional reference

Optional: primary reading

Optional: primary readingFCN · Long, Shelhamer & Darrell (2015)U-Net · Ronneberger, Fischer & Brox (2015)Mask R-CNN · He et al. (2017)Panoptic Segmentation · Kirillov et al. (2019)DeepPose · Toshev & Szegedy (2014)Stacked Hourglass · Newell et al. (2016)Depth Map Prediction · Eigen et al. (2014)Robust Monocular Depth · Ranftl et al. (2022)Segment Anything · Kirillov et al. (2023)RAFT · Teed & Deng (2020)

Paper links and image provenance are included in the reading notes and README.

Authored schematic / numerical teaching example.

ES 667 Deep Learning · Nipun Batra

A6

Which paper should I read for the mechanism I found most interesting?

Teaching notes

All architecture citations link to primary papers. The course-generated campus street is reused read-only from lecture9. Every overlay, toy mask, probability, and diagram is authored for instruction. No new model inference or empirical accuracy claim is presented. See SOURCES.md for the complete bibliography and provenance.

Optional reference

Optional: a VLM can generate coordinate tokens

Optional: a VLM can generate coordinate tokensimage + promptgenerative VLM<box>x1 y1 x2 y2</box>Schematic coordinate syntax, not a universal model interface.A valid sequence still needs spatial validation.

Generated coordinates still need a defined syntax, scale, and spatial evaluation.

Authored schematic / numerical teaching example.

Bai et al. (2023), Qwen-VL, §2

A7

Can a generative multimodal model return a box in its output sequence?

Teaching notes

Qwen-VL is one historical example of a multimodal model with grounding outputs. The displayed <box> example is schematic, not a literal promise about every VLM or a particular tokenizer. A valid token sequence does not guarantee correct spatial grounding. Keep this optional; the main lecture uses the clearer text-to-box-to-SAM composition.

Optional reference

Relative and metric depth answer different questions

Relative and metric depth answer different questionsRelative depthMetric depthnearer / farther orderingscale may be unknownsome methods also have a shift ambiguitydistance in a physical unite.g. metres along the camera z-axisrequires a calibrated scale conventionA plausible-looking depth map does not establish metric accuracy.The training target and evaluation protocol determine what the numbers mean.

The meaning of depth values follows from the supervision and scale convention.

Authored schematic / numerical teaching example.

Ranftl et al. (2022), Robust Monocular Depth Estimation, §3

A8

Can a relative-depth model’s raw output be reported as metres?

Teaching notes

Relative predictions can capture ordering or shape while leaving scale unknown. Some methods, including MiDaS-style training, handle scale and shift ambiguity in an inverse-depth representation. Do not generalize that all relative methods have exactly the same invariance. Metric evaluation needs the camera convention, physical units, and any permitted alignment stated explicitly.

Lecture overview

Present and annotate

Right arrow, Space, N: next reveal or slide Left arrow: previous reveal or slide P: presentation / reading · O: overview Q: question · A: answer · S: teaching notes D: pen on / off · E: stroke eraser · Z: undo Escape: leave pen mode, then leave presentation Ink stays with its slide as you advance through reveals. It is saved in this browser. Use Controls → Save ink to file for a portable backup, and Load ink from file to restore it. Controls hide after a short pause. Move the pointer or press a key to show them. They remain visible while the pen is active. The main route stops before optional references. Include them through Controls → Route.