TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Object detection:
from one label to a set of objects

VISION / OBJECT DETECTION

How can a fixed network produce a variable, unordered object set?

Prof. Nipun Batra
ES 667, IIT Gandhinagar

What does an image-level label omit?

This lecture follows ViT, CLIP and VLMs. Reuse image features and cross-attention; the new problem is the structure of the output.

Teaching note

This lecture follows ViT, CLIP and VLMs. Reuse image features and cross-attention; the new problem is the structure of the output.

1Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Why is classification not enough?

SECTION 00One imageOne label?

What would one image-level label leave out?

Read diagram labels
  • SECTION 00
  • One image
  • One label?

What would one image-level label leave out?

What would one image-level label leave out?

Teaching note

What would one image-level label leave out?

2Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One label?

One cat in the courtyard
catdog / other

Which label would you select?

Assume a mutually exclusive cat/dog/other classification task.

Teaching note

Assume a mutually exclusive cat/dog/other classification task.

3Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One label now?

A cat and a dog
cat or dog?Both are present.

Would a new “both” class solve detection?

The input type has not changed, but the requested description now includes two classes.

Teaching note

The input type has not changed, but the requested description now includes two classes.

4Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The winning class leaves the cat out

Cat and dog

SOFTMAX DISTRIBUTION

cat 0.46
dog 0.51
other 0.03

dog wins

Is this just a training problem?

These illustrative scores sum to one. Selecting the largest score returns one image-level label.

Teaching note

These illustrative scores sum to one. Selecting the largest score returns one image-level label.

5Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which classes are present?

Cat and dog

INDEPENDENT SCORES

cat 0.97
dog 0.95
other 0.02

What happens to dog-present if another dog enters?

Independent class scores can describe co-occurrence. They do not allocate one record per instance.

Teaching note

Independent class scores can describe co-occurrence. They do not allocate one record per instance.

6Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Class presence and object instances

CLASS PRESENCEis not the same as OBJECT INSTANCES

What extra output do we need?

The distinction is about representation, not accuracy.

Teaching note

The distinction is about representation, not accuracy.

7Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

What should one object record contain?

SECTION 01One designated objectClass and geometry

First solve the simpler case: one designated dog.

Read diagram labels
  • SECTION 01
  • One designated object
  • Class and geometry

First solve the simpler case: one designated dog.

First solve the simpler case: one designated dog.

Teaching note

First solve the simpler case: one designated dog.

8Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One dog, one record

The same single dog used as the target identifier

WHAT?

dog

Why is the score absent from the target?

Class and geometry define the target record. Add an inferred score only after establishing that record. A score is not automatically a calibrated probability that the complete detection is correct.

Teaching note

Class and geometry define the target record. Add an inferred score only after establishing that record. A score is not automatically a calibrated probability that the complete detection is correct.

9Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Centre and full box size

centre
b=(xc,yc,w,h)\mathbf b=(x_c,y_c,w,h)

Origin at the top left.
Horizontal coordinates increase rightward.
Vertical coordinates increase downward.

Read diagram labels
  • centre
  • full width
  • h

Is w half the box width?

Use this one format in the main derivation. Width and height span the complete box. Other conventions and letterboxing remain in the reference route.

Teaching note

Use this one format in the main derivation. Width and height span the complete box. Other conventions and letterboxing remain in the reference route.

10Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One backbone, two heads

Shared imagefeaturesClass head: C scoresBox head: 4 values
Read diagram labels
  • Shared image
  • features
  • Class head: C scores
  • Box head: 4 values

What is the limitation?

The feature extractor is shared. One box head with four outputs describes only one designated instance.

Teaching note

The feature extractor is shared. One box head with four outputs describes only one designated instance.

11Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Fixed tensor, variable object set

SECTION 02Fixed network capacityVariable object set

How many records should the model reserve?

Read diagram labels
  • SECTION 02
  • Fixed network capacity
  • Variable object set

How many records should the model reserve?

How many records should the model reserve?

Teaching note

How many records should the model reserve?

12Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Zero, two, or eleven objects

No target objects0211
Read diagram labels
  • No target objects
  • 0
  • 2
  • 11

Can the target have length zero?

Annotation policy defines the target instances. Empty images have an empty target set.

Teaching note

Annotation policy defines the target instances. Empty images have an empty target set.

13Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The annotation order does not change the target

Y={(ci,bi)}i=1N\mathcal Y=\{(c_i,\mathbf b_i)\}_{i=1}^{N}

STORED ROW 1

cat + cat box

STORED ROW 2

dog + dog box

A possible storage order.

What must move together when rows are swapped?

Rows can be permuted without changing the set. The class must stay paired with its own box.

Teaching note

Rows can be permuted without changing the set. The class must stay paired with its own box.

14Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A fixed tensor, a variable unordered set

NETWORK

fθ(x)∈RDf_\theta(x)\in\mathbb R^D
Fixed capacity

TARGET

Y={yi}i=1N\mathcal Y=\{y_i\}_{i=1}^{N}
Variable N

Must the network change its shape for every scene?

For a fixed input shape, the head emits a fixed-capacity tensor. We need a representation and selection rule for a variable set.

Teaching note

For a fixed input shape, the head emits a fixed-capacity tensor. We need a representation and selection rule for a variable set.

15Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

First attempt: K global slots

slot 1slot 2slot 3slot 4slot 5Reserve five records before seeing the image.
Read diagram labels
  • slot 1
  • slot 2
  • slot 3
  • slot 4
  • slot 5
  • Reserve five records before seeing the image.
  • left dog
  • right dog
  • empty
  • The list order is a choice. The objects have not changed.

What do unused slots represent?

K is a capacity, chosen before observing the particular scene. Empty slots need a no-object target.

Teaching note

K is a capacity, chosen before observing the particular scene. Empty slots need a no-object target.

16Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

What if there are more objects than slots?

Eleven people
N=11>K=5N=11>K=5

Five records cannot represent eleven distinct instances.

Would learned queries remove this limit?

Every finite candidate design has a capacity limit.

Teaching note

Every finite candidate design has a capacity limit.

17Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which dog is slot 1?

slot 1left dogslot 2right dogslot 3emptyslot 4emptyslot 5emptyThe list order is a choice. The objects have not changed.

One valid ordering.

Read diagram labels
  • slot 1
  • left dog
  • slot 2
  • right dog
  • slot 3
  • empty
  • slot 4
  • slot 5
  • The list order is a choice. The objects have not changed.

What changed between the two builds?

The target set does not identify a first dog. This is an assignment problem; spatial addresses are one possible response.

Teaching note

The target set does not identify a first dog. This is an assignment problem; spatial addresses are one possible response.

18Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Can the image give candidates an address?

The recurring cat and dog image
Use spatial locations?

Which information survives in a feature map?

Retain the feature map instead of collapsing it immediately to one global vector.

Teaching note

Retain the feature map instead of collapsing it immediately to one global vector.

19Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Give candidates spatial addresses

SECTION 03Spatial feature mapAn address per candidate

Use a feature-map location to index each record.

Read diagram labels
  • SECTION 03
  • Spatial feature map
  • An address per candidate

Use a feature-map location to index each record.

Use a feature-map location to index each record.

Teaching note

Use a feature-map location to index each record.

20Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A location can index a prediction

Spatial feature grid

Keep the spatial feature grid.

Read diagram labels
  • Spatial feature grid
  • One output address per grid location

Does each cell require a different network?

The grid indexes output records. Shared head weights operate across locations.

Teaching note

The grid indexes output records. Shared head weights operate across locations.

21Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

From image features to candidate records

IMAGEBACKBONE / FEATURESspatial featuresDENSE HEADshared weightsCANDIDATE RECORDSfixed capacityexistenceboxclassDECODEimage coordinatesSCORE FILTERrank / thresholdNMSremove duplicatesOBJECT SETvariable length
Read diagram labels
  • IMAGE
  • BACKBONE / FEATURES
  • spatial features
  • DENSE HEAD
  • shared weights
  • CANDIDATE RECORDS
  • fixed capacity
  • existence
  • box
  • class
  • DECODE
  • image coordinates
  • SCORE FILTER
  • rank / threshold
  • NMS
  • remove duplicates
  • OBJECT SET
  • variable length

What does the prediction head add?

This diagram will recur as later operations are introduced. Faded stages are still unresolved.

Teaching note

This diagram will recur as later operations are introduced. Faded stages are still unresolved.

22Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One generic prediction record

OBJECTNESS
pop_o
BOX
(x,y,w,h)(x,y,w,h)
CLASS
(p1,…,pC)(p_1,\ldots,p_C)

1 existence value + 4 box values + C class scores.

Where does the 5 in C+5 come from?

This is a generic teaching abstraction. It is not the exact parameterization of every detector.

Teaching note

This is a generic teaching abstraction. It is not the exact parameterization of every detector.

23Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The dimensions of the dense output

One output address per grid location
S×S×(5+C)S\times S\times(5+C)

S × S addresses
C foreground classes
One record at each address

4×4×84\times4\times8
Read diagram labels
  • One output address per grid location

How many records are reserved?

Here S=4 and C=3: cat, dog and other. The toy head emits 16 records with 8 values each. The class term uses a softmax over the foreground classes for positive candidates.

Teaching note

Here S=4 and C=3: cat, dog and other. The toy head emits 16 records with 8 values each. The class term uses a softmax over the foreground classes for positive candidates.

24Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Grid cell = output address

00112233dogDog 1same targetNormalized teaching grid; zero-based indices
The box can cross cell edges.

A feature at this address can see a much larger receptive field.

Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Must the box fit inside the cell?

A cell is not an independent image crop. The dog box deliberately crosses cell boundaries.

Teaching note

A cell is not an independent image crop. The dog box deliberately crosses cell boundaries.

25Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

One dog, one normalized teaching box

00112233dogDog 1same targetNormalized teaching grid; zero-based indices

DOG 1 / GROUND TRUTH

(xc,yc,w,h)=(x_c,y_c,w,h)=
(0.62, 0.41, 0.28, 0.34)(0.62,\,0.41,\,0.28,\,0.34)
S=4S=4
Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Which cell contains the centre?

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.

Teaching note

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.

26Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which cell owns the centre?

00112233dogDog 1same targetNormalized teaching grid; zero-based indices
j=⌊4×0.62⌋=2j=\lfloor 4\times 0.62\rfloor=2

Column index: 2.

Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Why is the first index not 1?

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. Row and column are zero based.

Teaching note

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. Row and column are zero based.

27Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

How far inside that cell is the centre?

00112233dogDog 1same targetNormalized teaching grid; zero-based indices
tx=Sxc−jt_x=Sx_c-j
=4×0.62−2=0.48=4\times0.62-2=0.48
Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Are all four values measured relative to a cell?

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. The offsets are fractions of one cell. Width and height are fractions of the full image.

Teaching note

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. The offsets are fractions of one cell. Width and height are fractions of the full image.

28Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Construct the target record

TARGET AT ROW 1, COLUMN 2

OBJECTNESS
11
BOX
(0.48, 0.64, 0.28, 0.34)(0.48,\,0.64,\,0.28,\,0.34)
CLASS

dog

Local centre offsets; full-image normalized width and height.

Why is objectness 1?

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.

Teaching note

The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.

29Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The model emits a record at the same address

ILLUSTRATIVE TEACHING PREDICTION

OBJECTNESS
0.800.80
BOX
(0.45, 0.60, 0.30, 0.32)(0.45,\,0.60,\,0.30,\,0.32)
CLASS

cat 0.10
dog 0.85
other 0.05

Is the predicted record identical to the target?

These are fixed illustrative predictions, not measured detector output. Every loss and decoding calculation uses exactly these values from the shared JSON.

Teaching note

These are fixed illustrative predictions, not measured detector output. Every loss and decoding calculation uses exactly these values from the shared JSON.

30Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which candidate should learn which object?

SECTION 04Candidates and annotationsSupervision for each record

A loss needs a target for each prediction.

Read diagram labels
  • SECTION 04
  • Candidates and annotations
  • Supervision for each record

A loss needs a target for each prediction.

A loss needs a target for each prediction.

Teaching note

A loss needs a target for each prediction.

31Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which of these candidates should learn the dog?

00112233dogDog 1same targetNormalized teaching grid; zero-based indices
Several nearby candidates see it.Which one receives its target?
Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Is this the assignment rule of every detector?

The simplified centre-cell policy chooses exactly one positive. Modern detectors use model-specific rules, often with several positives per object.

Teaching note

The simplified centre-cell policy chooses exactly one positive. Modern detectors use model-specific rules, often with several positives per object.

32Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Positive, background, and ignored candidates

00112233dogDog 1same targetNormalized teaching grid; zero-based indices

POSITIVE

The assigned cell learns the dog record.

BACKGROUND

An unassigned cell learns objectness 0.

IGNORE, WHEN USED

An ignored candidate contributes no loss.

Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Does the grey cell receive the dog box?

This toy policy has positives and background, with no ignore band. Other policies can exclude ambiguous candidates. Do not silently treat ignore as background.

Teaching note

This toy policy has positives and background, with no ignore band. Other policies can exclude ambiguous candidates. Do not silently treat ignore as background.

33Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Objectness: is an object assigned here?

TARGET

yo=1y_o=1

PREDICTION

po=0.80p_o=0.80

What happens if the objectness approaches 1?

Binary cross-entropy for this positive target reduces to minus log p. Natural logarithms are used.

Teaching note

Binary cross-entropy for this positive target reduces to minus log p. Natural logarithms are used.

34Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Class: how much probability went to dog?

TARGET CLASS

dog

PREDICTED CLASS SCORES

cat: 0.10
dog: 0.85
other: 0.05

Why does the cat probability not appear as a separate term?

Categorical cross-entropy selects the probability of the annotated class. This term supervises the assigned positive, not the background candidate.

Teaching note

Categorical cross-entropy selects the probability of the annotated class. This term supervises the assigned positive, not the background candidate.

35Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Box: compare one coordinate at a time

SIMPLE TEACHING L1 LOSS / SUM REDUCTION

Target

(0.48, 0.64, 0.28, 0.34)(0.48,\,0.64,\,0.28,\,0.34)

Prediction

(0.45, 0.60, 0.30, 0.32)(0.45,\,0.60,\,0.30,\,0.32)
∣t^x−tx∣=∣0.45−0.48∣=0.03|\hat t_x-t_x|=|0.45-0.48|=0.03

1 of 4 coordinate differences

Would the mean give the same numerical loss?

We sum absolute differences, rather than averaging. Geometry combines two local centre offsets and two image-normalized sizes in this deliberately simple pedagogical loss. It is not attributed to a YOLO version.

Teaching note

We sum absolute differences, rather than averaging. Geometry combines two local centre offsets and two image-normalized sizes in this deliberately simple pedagogical loss. It is not attributed to a YOLO version.

36Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Add the three supervised terms

L=λobjLobj+λboxLbox+λclsLcls\mathcal L=\lambda_{obj}\mathcal L_{obj}+\lambda_{box}\mathcal L_{box}+\lambda_{cls}\mathcal L_{cls}

For this worksheet, set all three weights to 1.

What must happen before computing these terms?

The total is computed from unrounded component values; displayed components are rounded to three decimals. This is the positive-candidate loss, not the complete batch loss.

Teaching note

The total is computed from unrounded component values; displayed components are rounded to three decimals. This is the positive-candidate loss, not the complete batch loss.

37Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A nearby unassigned candidate

00112233dogDog 1same targetNormalized teaching grid; zero-based indices

ROW 1, COLUMN 1

yo=0y_o=0
po=0.20p_o=0.20

No dog class or box target.

Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • dog
  • Dog 1
  • same target
  • Normalized teaching grid; zero-based indices

Does its predicted geometry need to be zero?

Background does not learn the dog box from the adjacent positive cell.

Teaching note

Background does not learn the dog box from the adjacent positive cell.

38Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Many background terms can dominate a sum

Separate 8 × 8 example

64 candidates
2 positive
62 background

Read diagram labels
  • Separate 8 × 8 example

Does low average loss guarantee good object recall?

This separate 8×8 example motivates imbalance; it does not change the 4×4 numerical trace. Easy negative terms can dominate an unweighted sum.

Teaching note

This separate 8×8 example motivates imbalance; it does not change the 4×4 numerical trace. Easy negative terms can dominate an unweighted sum.

39Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A teaching record, not a universal detector format

OBJECTNESSobjectness
BOXbox
CLASSclass

The three fields make the derivation concrete. Modern detectors can factor confidence, classes, and geometry differently.

Must every detector predict an explicit objectness scalar?

The record and losses are a generic dense-detector teaching abstraction. Original YOLOv1 confidence, anchor-free heads, distributional regression and DETR are not identical parameterizations.

Teaching note

The record and losses are a generic dense-detector teaching abstraction. Original YOLOv1 confidence, anchor-free heads, distributional regression and DETR are not identical parameterizations.

40Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

From candidates to final detections

SECTION 05Candidate predictionsReturned object set

At inference there are no annotations to consult.

Read diagram labels
  • SECTION 05
  • Candidate predictions
  • Returned object set

At inference there are no annotations to consult.

At inference there are no annotations to consult.

Teaching note

At inference there are no annotations to consult.

41Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Decode the same predicted box

00112233predictionDog 1same candidateNormalized teaching grid; zero-based indices
x^c=j+t^xS=2+0.454=0.6125\hat x_c=\frac{j+\hat t_x}{S}=\frac{2+0.45}{4}=0.6125
Read diagram labels
  • 0
  • 1
  • 2
  • 3
  • prediction
  • Dog 1
  • same candidate
  • Normalized teaching grid; zero-based indices

Why does the decoded centre differ from the ground truth?

Use the predicted offsets, not the target offsets. The model has not changed its record; decoding only interprets its parameterization.

Teaching note

Use the predicted offsets, not the target offsets. The model has not changed its record; decoding only interprets its parameterization.

42Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A score for this predicted dog

ONE DECLARED SCORING CONVENTION

sdog=po p(dog∣object)s_{dog}=p_o\,p(dog\mid object)
0.80×0.85=0.680.80\times0.85=0.68

The score ranks this candidate.

Does a score of 0.68 prove the box is correct?

This product is the declared score for the toy factorization. Other heads use other definitions. It does not add another learned field to the prediction.

Teaching note

This product is the declared score for the toy factorization. Other heads use other definitions. It does not add another learned field to the prediction.

43Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Decoding interprets the fixed candidate tensor

IMAGEBACKBONE / FEATURESspatial featuresDENSE HEADshared weightsCANDIDATE RECORDSfixed capacityexistenceboxclassDECODEimage coordinatesSCORE FILTERrank / thresholdNMSremove duplicatesOBJECT SETvariable length
Read diagram labels
  • IMAGE
  • BACKBONE / FEATURES
  • spatial features
  • DENSE HEAD
  • shared weights
  • CANDIDATE RECORDS
  • fixed capacity
  • existence
  • box
  • class
  • DECODE
  • image coordinates
  • SCORE FILTER
  • rank / threshold
  • NMS
  • remove duplicates
  • OBJECT SET
  • variable length

Which steps remain to choose what to return?

We have followed one dog candidate through target creation, supervision and decoding. The parcel scene now isolates selection among several candidates.

Teaching note

We have followed one dog candidate through target creation, supervision and decoding. The parcel scene now isolates selection among several candidates.

44Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Now several candidates describe one parcel

A 0.96B 0.93C 0.86D 0.76E 0.28CANDIDATESCORESTATUSA0.96candidateB0.93candidateC0.86candidateD0.76candidateE0.28candidate
Read diagram labels
  • A 0.96
  • B 0.93
  • C 0.86
  • D 0.76
  • E 0.28
  • CANDIDATE
  • SCORE
  • STATUS
  • A
  • 0.96
  • candidate
  • B
  • 0.93
  • C
  • 0.86
  • D
  • 0.76
  • E
  • 0.28

How many object instances do five candidates imply?

The parcel photograph and A-E values are constructed teaching data. Candidate IDs, coordinates, scores and colours stay fixed in every slide and the lab. Colours identify candidates, not correctness.

Teaching note

The parcel photograph and A-E values are constructed teaching data. Candidate IDs, coordinates, scores and colours stay fixed in every slide and the lab. Colours identify candidates, not correctness.

45Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Filter by score

A 0.96B 0.93C 0.86D 0.76E 0.28CANDIDATESCORESTATUSA0.96candidateB0.93candidateC0.86candidateD0.76candidateE0.28candidate

Retain scores at least 0.70.

Read diagram labels
  • A 0.96
  • B 0.93
  • C 0.86
  • D 0.76
  • E 0.28
  • CANDIDATE
  • SCORE
  • STATUS
  • A
  • 0.96
  • candidate
  • B
  • 0.93
  • C
  • 0.86
  • D
  • 0.76
  • E
  • 0.28
  • filtered

Why do A, B and C all remain?

Filtering tests the score, not spatial overlap. E remains faintly visible so the identity and geometry do not disappear from the explanation.

Teaching note

Filtering tests the score, not spatial overlap. E remains faintly visible so the identity and geometry do not disappear from the explanation.

46Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

How much do A and B overlap?

A 0.96B 0.93
IoU⁡(A,B)=∣A∩B∣∣A∪B∣\operatorname{IoU}(A,B)=\frac{|A\cap B|}{|A\cup B|}

Shared area divided by total covered area.

Read diagram labels
  • A 0.96
  • B 0.93
  • Shaded intersection: A and B
  • Shaded union: B lies inside A

What is the union when B lies wholly inside A?

All areas are in normalized image units and come from the displayed parcel boxes. B lies inside A in this example. IoU near 1 means strong overlap; near 0 means little overlap.

Teaching note

All areas are in normalized image units and come from the displayed parcel boxes. B lies inside A in this example. IoU near 1 means strong overlap; near 0 means little overlap.

47Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Non-maximum suppression

A 0.96KEEPB 0.93C 0.86D 0.76E 0.28CANDIDATESCORESTATUSA0.96KEEPB0.93pendingC0.86pendingD0.76pendingE0.28filtered

1. Keep A, the highest-scoring pending candidate.

Read diagram labels
  • A 0.96
  • KEEP
  • B 0.93
  • C 0.86
  • D 0.76
  • E 0.28
  • CANDIDATE
  • SCORE
  • STATUS
  • A
  • 0.96
  • B
  • 0.93
  • pending
  • C
  • 0.86
  • D
  • 0.76
  • E
  • 0.28
  • filtered
  • IoU 0.85
  • IoU 0.71
  • IoU 0.00
  • suppress

Why does D survive?

Scores, geometry and ID colours do not change. Green KEEP and red suppression text encode state separately. NMS is greedy prediction-to-prediction comparison, not truth matching.

Teaching note

Scores, geometry and ID colours do not change. Green KEEP and red suppression text encode state separately. NMS is greedy prediction-to-prediction comparison, not truth matching.

48Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The returned set has variable length

A 0.96KEEPB 0.93C 0.86D 0.76KEEPE 0.28CANDIDATESCORESTATUSA0.96KEEPB0.93suppressC0.86suppressD0.76KEEPE0.28filtered

5 candidates → 4 after filtering → 2 after NMS.

Read diagram labels
  • A 0.96
  • KEEP
  • B 0.93
  • C 0.86
  • D 0.76
  • E 0.28
  • CANDIDATE
  • SCORE
  • STATUS
  • A
  • 0.96
  • B
  • 0.93
  • suppress
  • C
  • 0.86
  • D
  • 0.76
  • E
  • 0.28
  • filtered

Which quantity was fixed?

The reported count is the number of retained predictions, not a guarantee of the true count.

Teaching note

The reported count is the number of retained predictions, not a guarantee of the true count.

49Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Class-aware NMS separates class groups

dogcat

Same geometry.
Different predicted classes.

Class-aware NMS does not make these boxes compete.

Some systems use class-agnostic NMS.

Read diagram labels
  • dog
  • cat

Does class-aware NMS solve crowded same-class instances?

The example deliberately aligns the two boxes so only class grouping changes. Same-class crowded objects can still suppress each other.

Teaching note

The example deliberately aligns the two boxes so only class grouping changes. Same-class crowded objects can still suppress each other.

50Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Fixed capacity, scene-dependent returned set

IMAGEBACKBONE / FEATURESspatial featuresDENSE HEADshared weightsCANDIDATE RECORDSfixed capacityexistenceboxclassDECODEimage coordinatesSCORE FILTERrank / thresholdNMSremove duplicatesOBJECT SETvariable length

FIXED CANDIDATE CAPACITY → VARIABLE RETURNED SET.

Read diagram labels
  • IMAGE
  • BACKBONE / FEATURES
  • spatial features
  • DENSE HEAD
  • shared weights
  • CANDIDATE RECORDS
  • fixed capacity
  • existence
  • box
  • class
  • DECODE
  • image coordinates
  • SCORE FILTER
  • rank / threshold
  • NMS
  • remove duplicates
  • OBJECT SET
  • variable length

Can the returned set be empty?

This is the dense-detector answer to the opening question.

Teaching note

This is the dense-detector answer to the opening question.

51Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Three different matching problems

TRAININGcandidate ↔ ground truthassign supervision
INFERENCEprediction ↔ predictionremove duplicates
EVALUATIONprediction ↔ ground truthmeasure correctness

Which one requires annotations while measuring a test set?

The operands and purpose differ, even when all three use boxes and overlap. Training assignment is not NMS or evaluation matching.

Teaching note

The operands and purpose differ, even when all three use boxes and overlap. Training assignment is not NMS or evaluation matching.

52Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

How do we measure a detector?

SECTION 06Predictions and ground truthMeasured performance

Compare the returned predictions with annotated instances.

Read diagram labels
  • SECTION 06
  • Predictions and ground truth
  • Measured performance

Compare the returned predictions with annotated instances.

Compare the returned predictions with annotated instances.

Teaching note

Compare the returned predictions with annotated instances.

53Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

When does a prediction count as a true positive?

  • Same image
  • Same class
  • IoU at or above the evaluation threshold
  • Ground-truth instance not already claimed

Why is a second accurate box on one object a false positive?

In this simplified rule, process predictions in descending score order and choose the greatest-IoU available same-class truth. Official evaluators also handle crowd, ignore, area and detection limits.

Teaching note and source

In this simplified rule, process predictions in descending score order and choose the greatest-IoU available same-class truth. Official evaluators also handle crowd, ignore, area and detection limits.

Primary source
54Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Precision and recall use different denominators

P=TPTP+FPP=\frac{TP}{TP+FP}

Among the predictions, how many are correct?

P=22+1=23P=\frac{2}{2+1}=\frac{2}{3}

Can one wrong box cause both FP and FN?

The example has two matched predictions, one false positive, and no missed truth.

Teaching note

The example has two matched predictions, one false positive, and no missed truth.

55Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A deployment cutoff and AP serve different purposes

DEPLOYMENT SCORE THRESHOLD

Choose which predictions the system returns at one operating point.

AP EVALUATION

Use the ranked prediction list across operating points.

Should we discard every prediction below a chosen deployment score before comparing AP?

AP is not precision at one arbitrary deployment threshold. Apply the benchmark protocol, including its detection limits. An aggressive prefilter can truncate the ranking and reduce measured recall.

Teaching note

AP is not precision at one arbitrary deployment threshold. Apply the benchmark protocol, including its detection limits. An aggressive prefilter can truncate the ranking and reduce measured recall.

56Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

A ranked prefix traces the PR curve

000.50.511PrecisionRecallTWO GROUND-TRUTH OBJECTSP1TP0.95

After 1: precision 1.000, recall 0.500.

Read diagram labels
  • 0
  • 0.5
  • 1
  • Precision
  • Recall
  • TWO GROUND-TRUTH OBJECTS
  • P1
  • TP
  • 0.95
  • P2
  • FP
  • 0.89
  • P3
  • 0.88

Which step adds a prediction without finding another object?

Each prediction appears with its point; future labels are hidden. The two true positives refer to distinct targets. Axes remain fixed in all builds.

Teaching note

Each prediction appears with its point; future labels are hidden. The two true positives refer to distinct targets. Axes remain fixed in all builds.

57Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

AP summarizes precision across recall

000.50.511PrecisionRecallTWO GROUND-TRUTH OBJECTSP1TP0.95P2FP0.89P3TP0.88

For this toy example, AP is the shaded area under the precision envelope.

Read diagram labels
  • 0
  • 0.5
  • 1
  • Precision
  • Recall
  • TWO GROUND-TRUTH OBJECTS
  • P1
  • TP
  • 0.95
  • P2
  • FP
  • 0.89
  • P3
  • 0.88

Can identical final counts produce different AP?

The green envelope uses the best precision at or beyond each recall. This toy continuous area differs from protocol-specific sampled AP such as COCO 101-point averaging.

Teaching note and source

The green envelope uses the best precision at or beyond each recall. This toy continuous area differs from protocol-specific sampled AP such as COCO 101-point averaging.

Primary source
58Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Average AP across classes

  • dog: AP 0.85
  • cat: AP 0.43
  • bird: AP 0.76
mAP⁡=1C∑c=1CAP⁡c\operatorname{mAP}=\frac1C\sum_{c=1}^{C}\operatorname{AP}_c
=0.85+0.43+0.763=0.68=\frac{0.85+0.43+0.76}{3}=0.68

What can the mean hide?

Constructed per-class AP values use the same protocol. Their mean is not pooled class-agnostic AP.

Teaching note

Constructed per-class AP values use the same protocol. Their mean is not pooled class-agnostic AP.

59Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The IoU protocol changes what counts as correct

mAP@.50

Match at IoU ≥ .50.
Compute AP and average across classes.

COCO-STYLE AP@[.50:.95]

Evaluate at .50, .55, …, .95.
Average across the ten thresholds and categories.

Which protocol demands tighter localization?

COCO additionally specifies detection caps, area ranges, recall sampling and crowd/ignore handling. Evaluation thresholds do not rerun NMS.

Teaching note and source

COCO additionally specifies detection caps, area ranges, recall sampling and crowd/ignore handling. Evaluation thresholds do not rerun NMS.

Primary source
60Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Another answer: set prediction with queries

SECTION 07Learned queriesObject records

Return to the global slots that had no fixed identity.

Read diagram labels
  • SECTION 07
  • Learned queries
  • Object records

Return to the global slots that had no fixed identity.

Return to the global slots that had no fixed identity.

Teaching note

Return to the global slots that had no fixed identity.

61Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

What if the slots could read the whole image?

slot 1right dogslot 2left dogslot 3emptyslot 4emptyslot 5emptyThe list order is a choice. The objects have not changed.

Earlier: which dog belongs to slot 1?

Read diagram labels
  • slot 1
  • right dog
  • slot 2
  • left dog
  • slot 3
  • empty
  • slot 4
  • slot 5
  • The list order is a choice. The objects have not changed.

Is query 1 permanently a dog query?

Queries are learned vectors, not text prompts or fixed class labels. Their count still limits output capacity.

Teaching note

Queries are learned vectors, not text prompts or fixed class labels. Their count still limits output capacity.

62Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Reuse cross-attention from Beyond Attention

LEARNED QUERIESQ: query statesFixed query capacity

Query states supply Q.

Read diagram labels
  • LEARNED QUERIES
  • Q: query states
  • Fixed query capacity
  • Encoded image features
  • K, V: keys and values
  • Cross-attention
  • One class distribution + one box per query

What is read, and what asks the question?

Distinguish K, the number of queries, from K in the attention notation for keys. In original DETR, the decoder combines self-attention among queries with cross-attention to encoded image features.

Teaching note and source

Distinguish K, the number of queries, from K in the attention notation for keys. In original DETR, the decoder combines self-attention among queries with cross-attention to encoded image features.

Primary source
63Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Which prediction should match which target?

THREE QUERY PREDICTIONSTWO TARGETSquery 1: cat + boxquery 2: dog + boxquery 3: second dog boxcatdog

Choose a one-to-one assignment using class and box cost.

Read diagram labels
  • THREE QUERY PREDICTIONS
  • TWO TARGETS
  • query 1: cat + box
  • query 2: dog + box
  • query 3: second dog box
  • cat
  • dog
  • no-object

Can two queries both match the one dog target?

Do not assume the shown pairing follows from row order. This is an illustrative lowest-cost assignment; a model computes costs from class scores and box geometry.

Teaching note

Do not assume the shown pairing follows from row order. This is an illustrative lowest-cost assignment; a model computes costs from class scores and box geometry.

64Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Now name the mechanism: bipartite matching

THREE QUERY PREDICTIONSTWO TARGETSquery 1: cat + boxquery 2: dog + boxquery 3: second dog boxcatdogno-object

Hungarian matching finds a minimum-total-cost one-to-one assignment.

Read diagram labels
  • THREE QUERY PREDICTIONS
  • TWO TARGETS
  • query 1: cat + box
  • query 2: dog + box
  • query 3: second dog box
  • cat
  • dog
  • no-object

Why does this avoid choosing an annotation order?

Matched queries learn class and geometry. Unmatched queries learn no-object. The mechanism and purpose precede the name. Original DETR trains a set predictor with this matching objective.

Teaching note and source

Matched queries learn class and geometry. Unmatched queries learn no-object. The mechanism and purpose precede the name. Original DETR trains a set predictor with this matching objective.

Primary source
65Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Two solutions to the same output problem

DENSE CANDIDATES

  • Spatial locations index records
  • Duplicates can remain
  • NMS selects among overlapping predictions

DETR-STYLE SET PREDICTION

  • Learned queries index records
  • One-to-one training discourages duplicates
  • Original DETR inference does not require NMS

What problem do both solve?

This compares the teaching dense model with original DETR. Neither is universally superior, and both retain finite capacity and can make errors.

Teaching note and source

This compares the teaching dense model with original DETR. Neither is universally superior, and both retain finite capacity and can make errors.

Primary source
66Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Summary: fixed capacity to a variable set

SECTION 08Fixed candidate capacityScene-dependent output

Connect the representation, training and inference decisions.

Read diagram labels
  • SECTION 08
  • Fixed candidate capacity
  • Scene-dependent output

Connect the representation, training and inference decisions.

Connect the representation, training and inference decisions.

Teaching note

Connect the representation, training and inference decisions.

67Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The complete dense-detector path

IMAGEBACKBONE / FEATURESspatial featuresDENSE HEADshared weightsCANDIDATE RECORDSfixed capacityexistenceboxclassDECODEimage coordinatesSCORE FILTERrank / thresholdNMSremove duplicatesOBJECT SETvariable lengthTRAINING ONLY: assign targets → losses → update the head
Read diagram labels
  • IMAGE
  • BACKBONE / FEATURES
  • spatial features
  • DENSE HEAD
  • shared weights
  • CANDIDATE RECORDS
  • fixed capacity
  • existence
  • box
  • class
  • DECODE
  • image coordinates
  • SCORE FILTER
  • rank / threshold
  • NMS
  • remove duplicates
  • OBJECT SET
  • variable length
  • TRAINING ONLY: assign targets → losses → update the head

Where do the annotations enter?

Annotations assign supervision while training the fixed-capacity head. At inference, decoding and selection return a scene-dependent subset.

Teaching note

Annotations assign supervision while training the fixed-capacity head. At inference, decoding and selection return a scene-dependent subset.

68Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

Try the same parcel candidates

The same parcel scene used throughout inference

Open the offline detector lab

  • Move score and IoU thresholds separately
  • Step through the greedy NMS choices
  • Compare class-aware and class-agnostic grouping

Download the PyTorch companion notebook

At IoU threshold 1, which candidates survive NMS?

The offline lab embeds the same canonical JSON and pure numerical functions as this lecture. No detector or network download is needed. The optional class-ambiguity scenario changes a label only, not candidate geometry or score.

Teaching note

The offline lab embeds the same canonical JSON and pure numerical functions as this lecture. No detector or network download is needed. The optional class-ambiguity scenario changes a label only, not candidate geometry or score.

69Frame overflows
TASKTARGETCANDIDATESASSIGNTRAININFEREVALUATE

The network can stay fixed-shape. The scene does not have to.

DENSE ANSWER

Image features → spatial candidate records → decode → filter / NMS → returned set

SET-PREDICTION ANSWER

Image features + learned queries → one-to-one-trained candidate set → object records

What changes from image to image?

Both routes use fixed capacity to describe a variable scene. The next representation question is what to predict when a box is insufficient.

Teaching note

Both routes use fixed capacity to describe a variable scene. The next representation question is what to predict when a box is insufficient.

70Frame overflows

OPTIONAL REFERENCE

Three coordinate formats

(xmin⁡,ymin⁡,xmax⁡,ymax⁡)(x_{\min},y_{\min},x_{\max},y_{\max})

Corners

(x,y,w,h)(x,y,w,h)

Top-left and size

(xc,yc,w,h)(x_c,y_c,w,h)

Centre and size

What would you check when implementing this?

The format must accompany the data. Centre coordinates are the means of corresponding corners. Width and height are corner differences. Do not confuse full sizes with half sizes.

Teaching note

The format must accompany the data. Centre coordinates are the means of corresponding corners. Width and height are corner differences. Do not confuse full sizes with half sizes.

R1Frame overflows

OPTIONAL REFERENCE

Worked coordinate conversion

(130, 330, 240, 310)(130,\,330,\,240,\,310)

Top-left x, y, width, height in pixels.

Image: 1200 × 675 pixels.

CENTRE IN PIXELS

(250, 485)(250,\,485)

NORMALIZED CENTRE AND SIZE

(0.208, 0.719, 0.200, 0.459)(0.208,\,0.719,\,0.200,\,0.459)

What would you check when implementing this?

Continuous coordinates are used. Centre = top-left + half-size. Divide horizontal values by image width and vertical values by image height.

Teaching note

Continuous coordinates are used. Centre = top-left + half-size. Divide horizontal values by image width and vertical values by image height.

R2Frame overflows

OPTIONAL REFERENCE

Letterboxing transforms the boxes too

Image before letterboxing
1200×675→640×3601200\times675\rightarrow640\times360
r=640/1200=0.5333r=640/1200=0.5333
x′=rx,y′=ry+140x'=rx,\quad y'=ry+140

Pad above and below to make a square.

What would you check when implementing this?

Scale the coordinates, then add the padding offsets. The example uses continuous geometry and symmetric vertical padding. Crops and flips also transform targets.

Teaching note

Scale the coordinates, then add the padding offsets. The example uses continuous geometry and symmetric vertical padding. Crops and flips also transform targets.

R3Frame overflows

OPTIONAL REFERENCE

A minimal PyTorch target

boxes = torch.tensor([[130., 330., 370., 640.],
                      [760., 180., 1050., 620.]])
labels = torch.tensor([2, 1], dtype=torch.int64)
target = {"boxes": boxes, "labels": labels}

Torchvision detection boxes use pixel corners.

What would you check when implementing this?

Check image dimensions, class IDs, dtypes and coordinate format. Transform the image and target together. Many torchvision detection models reserve class 0 for background.

Teaching note and source

Check image dimensions, class IDs, dtypes and coordinate format. Transform the image and target together. Many torchvision detection models reserve class 0 for background.

Primary source
R4Frame overflows

OPTIONAL REFERENCE

Target encoding and decoding are inverses

i=1,j=2,S=4i=1,\quad j=2,\quad S=4
tx=0.48,ty=0.64t_x=0.48,\quad t_y=0.64
xc=(2+0.48)/4=0.62x_c=(2+0.48)/4=0.62
yc=(1+0.64)/4=0.41y_c=(1+0.64)/4=0.41

What would you check when implementing this?

This round trip uses the target offsets. The main inference example uses different predicted offsets. Width and height remain normalized to the image. Other architectures use different box parameterizations.

Teaching note

This round trip uses the target offsets. The main inference example uses different predicted offsets. Width and height remain normalized to the image. Other architectures use different box parameterizations.

R5Frame overflows

OPTIONAL REFERENCE

Why YOLOv1 used 7 × 7 × 30

7×7 cells7\times7\text{ cells}

20 class probabilities per cell

20+2×5=3020+2\times5=30

Each box has x, y, w, h and confidence.

What would you check when implementing this?

Original YOLOv1 shares class probabilities between the two box predictors in a cell. Its confidence combines object presence and localization quality: Pr(object) × IoU. This is not the independent objectness term in our toy record.

Teaching note and source

Original YOLOv1 shares class probabilities between the two box predictors in a cell. Its confidence combines object presence and localization quality: Pr(object) × IoU. This is not the independent objectness term in our toy record.

Primary source
R6Frame overflows

OPTIONAL REFERENCE

Two predictors do not guarantee two objects

One location; several geometries.The assignment rule still matters.

Original YOLOv1 shares one class vector per cell.

Read diagram labels
  • One location; several geometries.
  • The assignment rule still matters.

What would you check when implementing this?

The predictors specialize in geometry through training. Two independent objects with centres in the same cell are still difficult to represent. These original predictors are not predefined anchor templates.

Teaching note and source

The predictors specialize in geometry through training. Two independent objects with centres in the same cell are still difficult to represent. These original predictors are not predefined anchor templates.

Primary source
R7Frame overflows

OPTIONAL REFERENCE

Anchors provide reference geometries

One location; several geometries.The assignment rule still matters.

Choose a reference shape, then predict offsets from it.

Read diagram labels
  • One location; several geometries.
  • The assignment rule still matters.

What would you check when implementing this?

Anchor-based heads can emit several records at one location. Assignment is model-specific and can use IoU and heuristics. Anchor-free detectors do not require these reference boxes.

Teaching note

Anchor-based heads can emit several records at one location. Assignment is model-specific and can use IoU and heuristics. Anchor-free detectors do not require these reference boxes.

R8Frame overflows

OPTIONAL REFERENCE

A compact dense prediction head

head = nn.Conv2d(channels, anchors * (5 + classes), 1)
raw = head(features)     # [batch, B*(5+C), Hf, Wf]
raw = raw.permute(0, 2, 3, 1)
raw = raw.reshape(batch, Hf, Wf, anchors, 5 + classes)

Assignment, target encoding and decoding are separate parts.

What would you check when implementing this?

Illustrative code, not a complete detector. Define all dimensions consistently. Separate branches and extra regression channels are common.

Teaching note

Illustrative code, not a complete detector. Define all dimensions consistently. Separate branches and extra regression channels are common.

R9Frame overflows

OPTIONAL REFERENCE

Feature pyramids support different object scales

Street with near and distant objects
  • Fine maps retain spatial detail
  • Coarse maps carry broader context
  • Several levels contribute candidates

What would you check when implementing this?

Feature Pyramid Networks combine coarse semantic features with higher-resolution maps through a top-down pathway and lateral connections. Image resizing alone cannot guarantee small-object performance.

Teaching note and source

Feature Pyramid Networks combine coarse semantic features with higher-resolution maps through a top-down pathway and lateral connections. Image resizing alone cannot guarantee small-object performance.

Primary source
R10Frame overflows

OPTIONAL REFERENCE

The full binary cross-entropy expression

L=−[ylog⁡p+(1−y)log⁡(1−p)]\mathcal L=-[y\log p+(1-y)\log(1-p)]

Positive target with p = 0.80

L=0.223\mathcal L=0.223

Background target with the same p

L=1.609\mathcal L=1.609

What would you check when implementing this?

Natural logarithms. A high probability is good for a positive and poor for background. Use numerically stable losses accepting logits, rather than taking logs manually.

Teaching note

Natural logarithms. A high probability is good for a positive and poor for background. Use numerically stable losses accepting logits, rather than taking logs manually.

R11Frame overflows

OPTIONAL REFERENCE

Box losses measure geometric error

COORDINATE LOSSES

L1 or Smooth L1 compares coordinates.

OVERLAP LOSSES

IoU-based losses compare regions. GIoU also considers the enclosing box.

What would you check when implementing this?

The choice depends on parameterization and failure modes. Ordinary IoU offers limited guidance for disjoint boxes. GIoU adds enclosing-region geometry. There is no universal ranking of losses.

Teaching note and source

The choice depends on parameterization and failure modes. Ordinary IoU offers limited guidance for disjoint boxes. GIoU adds enclosing-region geometry. There is no universal ranking of losses.

Primary source
R12Frame overflows

OPTIONAL REFERENCE

An additional IoU calculation

A=(1, 1, 5, 5)A=(1,\,1,\,5,\,5)
B=(3, 2, 7, 6)B=(3,\,2,\,7,\,6)

Continuous corner coordinates.

∣A∣=16,∣B∣=16|A|=16,\quad |B|=16
∣A∩B∣=6,∣A∪B∣=26|A\cap B|=6,\quad |A\cup B|=26
IoU⁡=6/26≈0.231\operatorname{IoU}=6/26\approx0.231

What would you check when implementing this?

Clamp intersection width and height at zero for non-overlap. Validate corner ordering before computing areas.

Teaching note

Clamp intersection width and height at zero for non-overlap. Validate corner ordering before computing areas.

R13Frame overflows

OPTIONAL REFERENCE

Class-aware NMS in Torchvision

from torchvision.ops import batched_nms

keep = batched_nms(boxes, scores, labels, iou_threshold=0.5)
boxes, scores = boxes[keep], scores[keep]
labels = labels[keep]

Run separately for each image. Labels define suppression groups.

What would you check when implementing this?

batched_nms compares only matching group indices. Do not mix images under the same class IDs without image-specific groups. Decoding and score filtering are still separate operations.

Teaching note and source

batched_nms compares only matching group indices. Do not mix images under the same class IDs without image-specific groups. Decoding and score filtering are still separate operations.

Primary source
R14Frame overflows

OPTIONAL REFERENCE

Duplicate, wrong-class and low-IoU predictions

  • A duplicate of an already claimed object is FP
  • A wrong-class prediction is FP for that class
  • A poorly located box can be FP while its target is FN
  • An object with no matched prediction is FN

What would you check when implementing this?

Apply the declared benchmark protocol. One erroneous prediction and one missed target can occur at the same time.

Teaching note

Apply the declared benchmark protocol. One erroneous prediction and one missed target can occur at the same time.

R15Frame overflows

OPTIONAL REFERENCE

Toy AP from two recall intervals

AP⁡area=0.50×1.00+0.50×23\operatorname{AP}_{area}=0.50\times1.00+0.50\times\frac{2}{3}
≈0.833\approx0.833

Continuous area under the interpolated precision envelope.

What would you check when implementing this?

The envelope is 1 up to recall 0.5 and 2/3 from 0.5 to 1. COCO samples 101 recall levels; do not label this exact continuous toy area as COCO AP.

Teaching note

The envelope is 1 up to recall 0.5 and 2/3 from 0.5 to 1. COCO samples 101 recall levels; do not label this exact continuous toy area as COCO AP.

R16Frame overflows

OPTIONAL REFERENCE

Final counts can hide a different ranking

TP → FP → TP

AP⁡area≈0.833\operatorname{AP}_{area}\approx0.833

TP → TP → FP

AP⁡area=1\operatorname{AP}_{area}=1

What would you check when implementing this?

Both complete lists have the same precision and recall. Putting both correct predictions first improves AP. This controlled example assumes the true positives match distinct targets.

Teaching note

Both complete lists have the same precision and recall. Putting both correct predictions first improves AP. This controlled example assumes the true positives match distinct targets.

R17Frame overflows

OPTIONAL REFERENCE

A high-level API may already run NMS

with torch.inference_mode():
    pred = model([image_tensor])[0]

keep = pred["scores"] >= 0.50
boxes = pred["boxes"][keep]
labels = pred["labels"][keep]

Torchvision Faster R-CNN returns post-processed detections.

What would you check when implementing this?

The model includes proposals, decoding, internal score selection and NMS. This keep line adds a user cutoff. Inspect the API before applying another NMS stage.

Teaching note and source

The model includes proposals, decoding, internal score selection and NMS. This keep line adds a user cutoff. Inspect the API before applying another NMS stage.

Primary source
R18Frame overflows

OPTIONAL REFERENCE

One-stage and two-stage detector families

ONE STAGE

Predict classes and boxes directly from dense features.

TWO STAGE

Generate proposals, then classify and refine regions.

What would you check when implementing this?

This is a structural distinction, not a universal speed or accuracy ranking. Faster R-CNN is a two-stage example. Architecture, resolution, hardware and implementation determine practical performance.

Teaching note and source

This is a structural distinction, not a universal speed or accuracy ranking. Faster R-CNN is a two-stage example. Architecture, resolution, hardware and implementation determine practical performance.

Primary source
R19Frame overflows

OPTIONAL REFERENCE

Sources and further practice

What would you check when implementing this?

Photographs are generated teaching illustrations reused from the existing lecture. Their overlays, predictions and scores are constructed examples, not measured detector outputs. Original provenance remains in lecture9/figures/PROVENANCE.md.

Teaching note

Photographs are generated teaching illustrations reused from the existing lecture. Their overlays, predictions and scores are constructed examples, not measured detector outputs. Original provenance remains in lecture9/figures/PROVENANCE.md.

R20Frame overflows

Lecture overview

Read or present

P: present/read. Arrow keys or Space: next reveal, then next slide. O: overview. Q: question. A: answer. S: notes. M: detector map. N: next. Home/End: first/last slide in the chosen route. Escape closes a dialog, then exits presentation.

Worked calculations and mechanisms unfold one operation at a time. Their buttons also work in reading mode. Controls hide after two seconds; move the pointer, tap, or press a key to show them.

Use the route selector for main lecture only or main plus optional reference.