Object detection:
from one label to a set of objects
VISION / OBJECT DETECTION
How can a fixed network produce a variable, unordered object set?
Prof. Nipun Batra
ES 667, IIT Gandhinagar
What does an image-level label omit?
This lecture follows ViT, CLIP and VLMs. Reuse image features and cross-attention; the new problem is the structure of the output.
Teaching note
This lecture follows ViT, CLIP and VLMs. Reuse image features and cross-attention; the new problem is the structure of the output.
Why is classification not enough?
What would one image-level label leave out?
Read diagram labels
- SECTION 00
- One image
- One label?
What would one image-level label leave out?
What would one image-level label leave out?
Teaching note
What would one image-level label leave out?
One label?
Which label would you select?
Assume a mutually exclusive cat/dog/other classification task.
Teaching note
Assume a mutually exclusive cat/dog/other classification task.
One label now?
Would a new “both” class solve detection?
The input type has not changed, but the requested description now includes two classes.
Teaching note
The input type has not changed, but the requested description now includes two classes.
The winning class leaves the cat out
SOFTMAX DISTRIBUTION
cat 0.46
dog 0.51
other 0.03
Is this just a training problem?
These illustrative scores sum to one. Selecting the largest score returns one image-level label.
Teaching note
These illustrative scores sum to one. Selecting the largest score returns one image-level label.
Which classes are present?
INDEPENDENT SCORES
cat 0.97
dog 0.95
other 0.02
What happens to dog-present if another dog enters?
Independent class scores can describe co-occurrence. They do not allocate one record per instance.
Teaching note
Independent class scores can describe co-occurrence. They do not allocate one record per instance.
Class presence and object instances
What extra output do we need?
The distinction is about representation, not accuracy.
Teaching note
The distinction is about representation, not accuracy.
What should one object record contain?
First solve the simpler case: one designated dog.
Read diagram labels
- SECTION 01
- One designated object
- Class and geometry
First solve the simpler case: one designated dog.
First solve the simpler case: one designated dog.
Teaching note
First solve the simpler case: one designated dog.
One dog, one record
WHAT?
WHERE?
A box enclosing this instance.
TRAINING TARGET
INFERENCE RECORD
The score ranks predicted records.
Why is the score absent from the target?
Class and geometry define the target record. Add an inferred score only after establishing that record. A score is not automatically a calibrated probability that the complete detection is correct.
Teaching note
Class and geometry define the target record. Add an inferred score only after establishing that record. A score is not automatically a calibrated probability that the complete detection is correct.
Centre and full box size
Origin at the top left.
Horizontal coordinates increase rightward.
Vertical coordinates increase downward.
Origin at the top left.
Horizontal coordinates increase rightward.
Vertical coordinates increase downward.
Origin at the top left.
Horizontal coordinates increase rightward.
Vertical coordinates increase downward.
Normalize horizontal values by image width and vertical values by height.
Read diagram labels
- centre
- full width
- h
Is w half the box width?
Use this one format in the main derivation. Width and height span the complete box. Other conventions and letterboxing remain in the reference route.
Teaching note
Use this one format in the main derivation. Width and height span the complete box. Other conventions and letterboxing remain in the reference route.
One backbone, two heads
This localization model has one box slot.
Read diagram labels
- Shared image
- features
- Class head: C scores
- Box head: 4 values
What is the limitation?
The feature extractor is shared. One box head with four outputs describes only one designated instance.
Teaching note
The feature extractor is shared. One box head with four outputs describes only one designated instance.
Fixed tensor, variable object set
How many records should the model reserve?
Read diagram labels
- SECTION 02
- Fixed network capacity
- Variable object set
How many records should the model reserve?
How many records should the model reserve?
Teaching note
How many records should the model reserve?
Zero, two, or eleven objects
Read diagram labels
- No target objects
- 0
- 2
- 11
Can the target have length zero?
Annotation policy defines the target instances. Empty images have an empty target set.
Teaching note
Annotation policy defines the target instances. Empty images have an empty target set.
The annotation order does not change the target
STORED ROW 1
cat + cat box
STORED ROW 2
dog + dog box
A possible storage order.
STORED ROW 1
dog + dog box
STORED ROW 2
cat + cat box
SAME TARGET SET.
What must move together when rows are swapped?
Rows can be permuted without changing the set. The class must stay paired with its own box.
Teaching note
Rows can be permuted without changing the set. The class must stay paired with its own box.
A fixed tensor, a variable unordered set
NETWORK
TARGET
Must the network change its shape for every scene?
For a fixed input shape, the head emits a fixed-capacity tensor. We need a representation and selection rule for a variable set.
Teaching note
For a fixed input shape, the head emits a fixed-capacity tensor. We need a representation and selection rule for a variable set.
First attempt: K global slots
Read diagram labels
- slot 1
- slot 2
- slot 3
- slot 4
- slot 5
- Reserve five records before seeing the image.
- left dog
- right dog
- empty
- The list order is a choice. The objects have not changed.
What do unused slots represent?
K is a capacity, chosen before observing the particular scene. Empty slots need a no-object target.
Teaching note
K is a capacity, chosen before observing the particular scene. Empty slots need a no-object target.
What if there are more objects than slots?
Five records cannot represent eleven distinct instances.
Would learned queries remove this limit?
Every finite candidate design has a capacity limit.
Teaching note
Every finite candidate design has a capacity limit.
Which dog is slot 1?
One valid ordering.
A SLOT NUMBER HAS NO INHERENT VISUAL MEANING.
Read diagram labels
- slot 1
- left dog
- slot 2
- right dog
- slot 3
- empty
- slot 4
- slot 5
- The list order is a choice. The objects have not changed.
What changed between the two builds?
The target set does not identify a first dog. This is an assignment problem; spatial addresses are one possible response.
Teaching note
The target set does not identify a first dog. This is an assignment problem; spatial addresses are one possible response.
Can the image give candidates an address?
Which information survives in a feature map?
Retain the feature map instead of collapsing it immediately to one global vector.
Teaching note
Retain the feature map instead of collapsing it immediately to one global vector.
Give candidates spatial addresses
Use a feature-map location to index each record.
Read diagram labels
- SECTION 03
- Spatial feature map
- An address per candidate
Use a feature-map location to index each record.
Use a feature-map location to index each record.
Teaching note
Use a feature-map location to index each record.
A location can index a prediction
Keep the spatial feature grid.
Select one row and column.
AT THIS ADDRESS
Apply the same prediction head at every location.
Read diagram labels
- Spatial feature grid
- One output address per grid location
Does each cell require a different network?
The grid indexes output records. Shared head weights operate across locations.
Teaching note
The grid indexes output records. Shared head weights operate across locations.
From image features to candidate records
Read diagram labels
- IMAGE
- BACKBONE / FEATURES
- spatial features
- DENSE HEAD
- shared weights
- CANDIDATE RECORDS
- fixed capacity
- existence
- box
- class
- DECODE
- image coordinates
- SCORE FILTER
- rank / threshold
- NMS
- remove duplicates
- OBJECT SET
- variable length
What does the prediction head add?
This diagram will recur as later operations are introduced. Faded stages are still unresolved.
Teaching note
This diagram will recur as later operations are introduced. Faded stages are still unresolved.
One generic prediction record
1 existence value + 4 box values + C class scores.
Where does the 5 in C+5 come from?
This is a generic teaching abstraction. It is not the exact parameterization of every detector.
Teaching note
This is a generic teaching abstraction. It is not the exact parameterization of every detector.
The dimensions of the dense output
S × S addresses
C foreground classes
One record at each address
Read diagram labels
- One output address per grid location
How many records are reserved?
Here S=4 and C=3: cat, dog and other. The toy head emits 16 records with 8 values each. The class term uses a softmax over the foreground classes for positive candidates.
Teaching note
Here S=4 and C=3: cat, dog and other. The toy head emits 16 records with 8 values each. The class term uses a softmax over the foreground classes for positive candidates.
Grid cell = output address
A feature at this address can see a much larger receptive field.
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Must the box fit inside the cell?
A cell is not an independent image crop. The dog box deliberately crosses cell boundaries.
Teaching note
A cell is not an independent image crop. The dog box deliberately crosses cell boundaries.
One dog, one normalized teaching box
DOG 1 / GROUND TRUTH
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Which cell contains the centre?
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.
Teaching note
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.
Which cell owns the centre?
Column index: 2.
Row 1, column 2.
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Why is the first index not 1?
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. Row and column are zero based.
Teaching note
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. Row and column are zero based.
How far inside that cell is the centre?
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Are all four values measured relative to a cell?
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. The offsets are fractions of one cell. Width and height are fractions of the full image.
Teaching note
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace. The offsets are fractions of one cell. Width and height are fractions of the full image.
Construct the target record
TARGET AT ROW 1, COLUMN 2
dog
Local centre offsets; full-image normalized width and height.
Why is objectness 1?
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.
Teaching note
The normalized box is a constructed teaching target. The generated dog photograph identifies the object; it is not a measured dataset annotation. The same values, target record and prediction recur throughout the numerical trace.
The model emits a record at the same address
ILLUSTRATIVE TEACHING PREDICTION
cat 0.10
dog 0.85
other 0.05
Is the predicted record identical to the target?
These are fixed illustrative predictions, not measured detector output. Every loss and decoding calculation uses exactly these values from the shared JSON.
Teaching note
These are fixed illustrative predictions, not measured detector output. Every loss and decoding calculation uses exactly these values from the shared JSON.
Which candidate should learn which object?
A loss needs a target for each prediction.
Read diagram labels
- SECTION 04
- Candidates and annotations
- Supervision for each record
A loss needs a target for each prediction.
A loss needs a target for each prediction.
Teaching note
A loss needs a target for each prediction.
Which of these candidates should learn the dog?
TEACHING ASSIGNMENT RULE
Assign the object to the cell containing its centre.
Modern detectors use model-specific assignment rules.
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Is this the assignment rule of every detector?
The simplified centre-cell policy chooses exactly one positive. Modern detectors use model-specific rules, often with several positives per object.
Teaching note
The simplified centre-cell policy chooses exactly one positive. Modern detectors use model-specific rules, often with several positives per object.
Positive, background, and ignored candidates
POSITIVE
The assigned cell learns the dog record.
BACKGROUND
An unassigned cell learns objectness 0.
IGNORE, WHEN USED
An ignored candidate contributes no loss.
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Does the grey cell receive the dog box?
This toy policy has positives and background, with no ignore band. Other policies can exclude ambiguous candidates. Do not silently treat ignore as background.
Teaching note
This toy policy has positives and background, with no ignore band. Other policies can exclude ambiguous candidates. Do not silently treat ignore as background.
Objectness: is an object assigned here?
TARGET
PREDICTION
TARGET / PREDICTION
What happens if the objectness approaches 1?
Binary cross-entropy for this positive target reduces to minus log p. Natural logarithms are used.
Teaching note
Binary cross-entropy for this positive target reduces to minus log p. Natural logarithms are used.
Class: how much probability went to dog?
TARGET CLASS
PREDICTED CLASS SCORES
cat: 0.10
dog: 0.85
other: 0.05
CORRECT CLASS
Why does the cat probability not appear as a separate term?
Categorical cross-entropy selects the probability of the annotated class. This term supervises the assigned positive, not the background candidate.
Teaching note
Categorical cross-entropy selects the probability of the annotated class. This term supervises the assigned positive, not the background candidate.
Box: compare one coordinate at a time
SIMPLE TEACHING L1 LOSS / SUM REDUCTION
Target
Prediction
1 of 4 coordinate differences
SIMPLE TEACHING L1 LOSS / SUM REDUCTION
Target
Prediction
2 of 4 coordinate differences
SIMPLE TEACHING L1 LOSS / SUM REDUCTION
Target
Prediction
3 of 4 coordinate differences
SIMPLE TEACHING L1 LOSS / SUM REDUCTION
Target
Prediction
4 of 4 coordinate differences
SIMPLE TEACHING L1 LOSS / SUM REDUCTION
Target
Prediction
Real detectors often use IoU-family losses or other box parameterizations.
Would the mean give the same numerical loss?
We sum absolute differences, rather than averaging. Geometry combines two local centre offsets and two image-normalized sizes in this deliberately simple pedagogical loss. It is not attributed to a YOLO version.
Teaching note
We sum absolute differences, rather than averaging. Geometry combines two local centre offsets and two image-normalized sizes in this deliberately simple pedagogical loss. It is not attributed to a YOLO version.
Add the three supervised terms
For this worksheet, set all three weights to 1.
What must happen before computing these terms?
The total is computed from unrounded component values; displayed components are rounded to three decimals. This is the positive-candidate loss, not the complete batch loss.
Teaching note
The total is computed from unrounded component values; displayed components are rounded to three decimals. This is the positive-candidate loss, not the complete batch loss.
A nearby unassigned candidate
ROW 1, COLUMN 1
No dog class or box target.
The box and class terms are masked out.
Read diagram labels
- 0
- 1
- 2
- 3
- dog
- Dog 1
- same target
- Normalized teaching grid; zero-based indices
Does its predicted geometry need to be zero?
Background does not learn the dog box from the adjacent positive cell.
Teaching note
Background does not learn the dog box from the adjacent positive cell.
Many background terms can dominate a sum
64 candidates
2 positive
62 background
2 positive terms
62 background terms
Weighting, focal loss, or hard-example selection can reduce this imbalance.
Read diagram labels
- Separate 8 × 8 example
Does low average loss guarantee good object recall?
This separate 8×8 example motivates imbalance; it does not change the 4×4 numerical trace. Easy negative terms can dominate an unweighted sum.
Teaching note
This separate 8×8 example motivates imbalance; it does not change the 4×4 numerical trace. Easy negative terms can dominate an unweighted sum.
A teaching record, not a universal detector format
The three fields make the derivation concrete. Modern detectors can factor confidence, classes, and geometry differently.
Must every detector predict an explicit objectness scalar?
The record and losses are a generic dense-detector teaching abstraction. Original YOLOv1 confidence, anchor-free heads, distributional regression and DETR are not identical parameterizations.
Teaching note
The record and losses are a generic dense-detector teaching abstraction. Original YOLOv1 confidence, anchor-free heads, distributional regression and DETR are not identical parameterizations.
From candidates to final detections
At inference there are no annotations to consult.
Read diagram labels
- SECTION 05
- Candidate predictions
- Returned object set
At inference there are no annotations to consult.
At inference there are no annotations to consult.
Teaching note
At inference there are no annotations to consult.
Decode the same predicted box
DECODED PREDICTION
Width and height were already normalized to the full image.
Read diagram labels
- 0
- 1
- 2
- 3
- prediction
- Dog 1
- same candidate
- Normalized teaching grid; zero-based indices
Why does the decoded centre differ from the ground truth?
Use the predicted offsets, not the target offsets. The model has not changed its record; decoding only interprets its parameterization.
Teaching note
Use the predicted offsets, not the target offsets. The model has not changed its record; decoding only interprets its parameterization.
A score for this predicted dog
ONE DECLARED SCORING CONVENTION
The score ranks this candidate.
Does a score of 0.68 prove the box is correct?
This product is the declared score for the toy factorization. Other heads use other definitions. It does not add another learned field to the prediction.
Teaching note
This product is the declared score for the toy factorization. Other heads use other definitions. It does not add another learned field to the prediction.
Decoding interprets the fixed candidate tensor
Read diagram labels
- IMAGE
- BACKBONE / FEATURES
- spatial features
- DENSE HEAD
- shared weights
- CANDIDATE RECORDS
- fixed capacity
- existence
- box
- class
- DECODE
- image coordinates
- SCORE FILTER
- rank / threshold
- NMS
- remove duplicates
- OBJECT SET
- variable length
Which steps remain to choose what to return?
We have followed one dog candidate through target creation, supervision and decoding. The parcel scene now isolates selection among several candidates.
Teaching note
We have followed one dog candidate through target creation, supervision and decoding. The parcel scene now isolates selection among several candidates.
Now several candidates describe one parcel
Read diagram labels
- A 0.96
- B 0.93
- C 0.86
- D 0.76
- E 0.28
- CANDIDATE
- SCORE
- STATUS
- A
- 0.96
- candidate
- B
- 0.93
- C
- 0.86
- D
- 0.76
- E
- 0.28
How many object instances do five candidates imply?
The parcel photograph and A-E values are constructed teaching data. Candidate IDs, coordinates, scores and colours stay fixed in every slide and the lab. Colours identify candidates, not correctness.
Teaching note
The parcel photograph and A-E values are constructed teaching data. Candidate IDs, coordinates, scores and colours stay fixed in every slide and the lab. Colours identify candidates, not correctness.
Filter by score
Retain scores at least 0.70.
E fades. Why are four boxes still left for one parcel?
Read diagram labels
- A 0.96
- B 0.93
- C 0.86
- D 0.76
- E 0.28
- CANDIDATE
- SCORE
- STATUS
- A
- 0.96
- candidate
- B
- 0.93
- C
- 0.86
- D
- 0.76
- E
- 0.28
- filtered
Why do A, B and C all remain?
Filtering tests the score, not spatial overlap. E remains faintly visible so the identity and geometry do not disappear from the explanation.
Teaching note
Filtering tests the score, not spatial overlap. E remains faintly visible so the identity and geometry do not disappear from the explanation.
How much do A and B overlap?
Shared area divided by total covered area.
Union = area A + area B − intersection
Read diagram labels
- A 0.96
- B 0.93
- Shaded intersection: A and B
- Shaded union: B lies inside A
What is the union when B lies wholly inside A?
All areas are in normalized image units and come from the displayed parcel boxes. B lies inside A in this example. IoU near 1 means strong overlap; near 0 means little overlap.
Teaching note
All areas are in normalized image units and come from the displayed parcel boxes. B lies inside A in this example. IoU near 1 means strong overlap; near 0 means little overlap.
Non-maximum suppression
1. Keep A, the highest-scoring pending candidate.
2. Compare A with the other pending boxes.
3. Suppress B and C: IoU exceeds 0.50. D stays pending.
4. Keep D. The pending set is empty.
Read diagram labels
- A 0.96
- KEEP
- B 0.93
- C 0.86
- D 0.76
- E 0.28
- CANDIDATE
- SCORE
- STATUS
- A
- 0.96
- B
- 0.93
- pending
- C
- 0.86
- D
- 0.76
- E
- 0.28
- filtered
- IoU 0.85
- IoU 0.71
- IoU 0.00
- suppress
Why does D survive?
Scores, geometry and ID colours do not change. Green KEEP and red suppression text encode state separately. NMS is greedy prediction-to-prediction comparison, not truth matching.
Teaching note
Scores, geometry and ID colours do not change. Green KEEP and red suppression text encode state separately. NMS is greedy prediction-to-prediction comparison, not truth matching.
The returned set has variable length
5 candidates → 4 after filtering → 2 after NMS.
Read diagram labels
- A 0.96
- KEEP
- B 0.93
- C 0.86
- D 0.76
- E 0.28
- CANDIDATE
- SCORE
- STATUS
- A
- 0.96
- B
- 0.93
- suppress
- C
- 0.86
- D
- 0.76
- E
- 0.28
- filtered
Which quantity was fixed?
The reported count is the number of retained predictions, not a guarantee of the true count.
Teaching note
The reported count is the number of retained predictions, not a guarantee of the true count.
Class-aware NMS separates class groups
Same geometry.
Different predicted classes.
Class-aware NMS does not make these boxes compete.
Some systems use class-agnostic NMS.
Read diagram labels
- dog
- cat
Does class-aware NMS solve crowded same-class instances?
The example deliberately aligns the two boxes so only class grouping changes. Same-class crowded objects can still suppress each other.
Teaching note
The example deliberately aligns the two boxes so only class grouping changes. Same-class crowded objects can still suppress each other.
Fixed capacity, scene-dependent returned set
FIXED CANDIDATE CAPACITY → VARIABLE RETURNED SET.
Read diagram labels
- IMAGE
- BACKBONE / FEATURES
- spatial features
- DENSE HEAD
- shared weights
- CANDIDATE RECORDS
- fixed capacity
- existence
- box
- class
- DECODE
- image coordinates
- SCORE FILTER
- rank / threshold
- NMS
- remove duplicates
- OBJECT SET
- variable length
Can the returned set be empty?
This is the dense-detector answer to the opening question.
Teaching note
This is the dense-detector answer to the opening question.
Three different matching problems
Which one requires annotations while measuring a test set?
The operands and purpose differ, even when all three use boxes and overlap. Training assignment is not NMS or evaluation matching.
Teaching note
The operands and purpose differ, even when all three use boxes and overlap. Training assignment is not NMS or evaluation matching.
How do we measure a detector?
Compare the returned predictions with annotated instances.
Read diagram labels
- SECTION 06
- Predictions and ground truth
- Measured performance
Compare the returned predictions with annotated instances.
Compare the returned predictions with annotated instances.
Teaching note
Compare the returned predictions with annotated instances.
When does a prediction count as a true positive?
- Same image
- Same class
- IoU at or above the evaluation threshold
- Ground-truth instance not already claimed
Why is a second accurate box on one object a false positive?
In this simplified rule, process predictions in descending score order and choose the greatest-IoU available same-class truth. Official evaluators also handle crowd, ignore, area and detection limits.
Teaching note and source
In this simplified rule, process predictions in descending score order and choose the greatest-IoU available same-class truth. Official evaluators also handle crowd, ignore, area and detection limits.
Primary sourcePrecision and recall use different denominators
Among the predictions, how many are correct?
Among the annotated objects, how many did we find?
Can one wrong box cause both FP and FN?
The example has two matched predictions, one false positive, and no missed truth.
Teaching note
The example has two matched predictions, one false positive, and no missed truth.
A deployment cutoff and AP serve different purposes
DEPLOYMENT SCORE THRESHOLD
Choose which predictions the system returns at one operating point.
AP EVALUATION
Use the ranked prediction list across operating points.
Should we discard every prediction below a chosen deployment score before comparing AP?
AP is not precision at one arbitrary deployment threshold. Apply the benchmark protocol, including its detection limits. An aggressive prefilter can truncate the ranking and reduce measured recall.
Teaching note
AP is not precision at one arbitrary deployment threshold. Apply the benchmark protocol, including its detection limits. An aggressive prefilter can truncate the ranking and reduce measured recall.
A ranked prefix traces the PR curve
After 1: precision 1.000, recall 0.500.
After 2: precision 0.500, recall 0.500.
After 3: precision 0.667, recall 1.000.
Read diagram labels
- 0
- 0.5
- 1
- Precision
- Recall
- TWO GROUND-TRUTH OBJECTS
- P1
- TP
- 0.95
- P2
- FP
- 0.89
- P3
- 0.88
Which step adds a prediction without finding another object?
Each prediction appears with its point; future labels are hidden. The two true positives refer to distinct targets. Axes remain fixed in all builds.
Teaching note
Each prediction appears with its point; future labels are hidden. The two true positives refer to distinct targets. Axes remain fixed in all builds.
AP summarizes precision across recall
For this toy example, AP is the shaded area under the precision envelope.
Read diagram labels
- 0
- 0.5
- 1
- Precision
- Recall
- TWO GROUND-TRUTH OBJECTS
- P1
- TP
- 0.95
- P2
- FP
- 0.89
- P3
- 0.88
Can identical final counts produce different AP?
The green envelope uses the best precision at or beyond each recall. This toy continuous area differs from protocol-specific sampled AP such as COCO 101-point averaging.
Teaching note and source
The green envelope uses the best precision at or beyond each recall. This toy continuous area differs from protocol-specific sampled AP such as COCO 101-point averaging.
Primary sourceAverage AP across classes
- dog: AP 0.85
- cat: AP 0.43
- bird: AP 0.76
What can the mean hide?
Constructed per-class AP values use the same protocol. Their mean is not pooled class-agnostic AP.
Teaching note
Constructed per-class AP values use the same protocol. Their mean is not pooled class-agnostic AP.
The IoU protocol changes what counts as correct
mAP@.50
Match at IoU ≥ .50.
Compute AP and average across classes.
COCO-STYLE AP@[.50:.95]
Evaluate at .50, .55, …, .95.
Average across the ten thresholds and categories.
Which protocol demands tighter localization?
COCO additionally specifies detection caps, area ranges, recall sampling and crowd/ignore handling. Evaluation thresholds do not rerun NMS.
Teaching note and source
COCO additionally specifies detection caps, area ranges, recall sampling and crowd/ignore handling. Evaluation thresholds do not rerun NMS.
Primary sourceAnother answer: set prediction with queries
Return to the global slots that had no fixed identity.
Read diagram labels
- SECTION 07
- Learned queries
- Object records
Return to the global slots that had no fixed identity.
Return to the global slots that had no fixed identity.
Teaching note
Return to the global slots that had no fixed identity.
What if the slots could read the whole image?
Earlier: which dog belongs to slot 1?
LEARNED QUERY VECTORS
Each query reads the image features and predicts an object record.
Read diagram labels
- slot 1
- right dog
- slot 2
- left dog
- slot 3
- empty
- slot 4
- slot 5
- The list order is a choice. The objects have not changed.
Is query 1 permanently a dog query?
Queries are learned vectors, not text prompts or fixed class labels. Their count still limits output capacity.
Teaching note
Queries are learned vectors, not text prompts or fixed class labels. Their count still limits output capacity.
Reuse cross-attention from Beyond Attention
Query states supply Q.
Encoded image features supply K and V.
Learned queries are another way to index candidate records.
Read diagram labels
- LEARNED QUERIES
- Q: query states
- Fixed query capacity
- Encoded image features
- K, V: keys and values
- Cross-attention
- One class distribution + one box per query
What is read, and what asks the question?
Distinguish K, the number of queries, from K in the attention notation for keys. In original DETR, the decoder combines self-attention among queries with cross-attention to encoded image features.
Teaching note and source
Distinguish K, the number of queries, from K in the attention notation for keys. In original DETR, the decoder combines self-attention among queries with cross-attention to encoded image features.
Primary sourceWhich prediction should match which target?
Choose a one-to-one assignment using class and box cost.
A low-total-cost assignment pairs each target with one prediction.
The unmatched query learns no-object.
Read diagram labels
- THREE QUERY PREDICTIONS
- TWO TARGETS
- query 1: cat + box
- query 2: dog + box
- query 3: second dog box
- cat
- dog
- no-object
Can two queries both match the one dog target?
Do not assume the shown pairing follows from row order. This is an illustrative lowest-cost assignment; a model computes costs from class scores and box geometry.
Teaching note
Do not assume the shown pairing follows from row order. This is an illustrative lowest-cost assignment; a model computes costs from class scores and box geometry.
Now name the mechanism: bipartite matching
Hungarian matching finds a minimum-total-cost one-to-one assignment.
Read diagram labels
- THREE QUERY PREDICTIONS
- TWO TARGETS
- query 1: cat + box
- query 2: dog + box
- query 3: second dog box
- cat
- dog
- no-object
Why does this avoid choosing an annotation order?
Matched queries learn class and geometry. Unmatched queries learn no-object. The mechanism and purpose precede the name. Original DETR trains a set predictor with this matching objective.
Teaching note and source
Matched queries learn class and geometry. Unmatched queries learn no-object. The mechanism and purpose precede the name. Original DETR trains a set predictor with this matching objective.
Primary sourceTwo solutions to the same output problem
DENSE CANDIDATES
- Spatial locations index records
- Duplicates can remain
- NMS selects among overlapping predictions
DETR-STYLE SET PREDICTION
- Learned queries index records
- One-to-one training discourages duplicates
- Original DETR inference does not require NMS
What problem do both solve?
This compares the teaching dense model with original DETR. Neither is universally superior, and both retain finite capacity and can make errors.
Teaching note and source
This compares the teaching dense model with original DETR. Neither is universally superior, and both retain finite capacity and can make errors.
Primary sourceSummary: fixed capacity to a variable set
Connect the representation, training and inference decisions.
Read diagram labels
- SECTION 08
- Fixed candidate capacity
- Scene-dependent output
Connect the representation, training and inference decisions.
Connect the representation, training and inference decisions.
Teaching note
Connect the representation, training and inference decisions.
The complete dense-detector path
Read diagram labels
- IMAGE
- BACKBONE / FEATURES
- spatial features
- DENSE HEAD
- shared weights
- CANDIDATE RECORDS
- fixed capacity
- existence
- box
- class
- DECODE
- image coordinates
- SCORE FILTER
- rank / threshold
- NMS
- remove duplicates
- OBJECT SET
- variable length
- TRAINING ONLY: assign targets → losses → update the head
Where do the annotations enter?
Annotations assign supervision while training the fixed-capacity head. At inference, decoding and selection return a scene-dependent subset.
Teaching note
Annotations assign supervision while training the fixed-capacity head. At inference, decoding and selection return a scene-dependent subset.
Try the same parcel candidates
- Move score and IoU thresholds separately
- Step through the greedy NMS choices
- Compare class-aware and class-agnostic grouping
At IoU threshold 1, which candidates survive NMS?
The offline lab embeds the same canonical JSON and pure numerical functions as this lecture. No detector or network download is needed. The optional class-ambiguity scenario changes a label only, not candidate geometry or score.
Teaching note
The offline lab embeds the same canonical JSON and pure numerical functions as this lecture. No detector or network download is needed. The optional class-ambiguity scenario changes a label only, not candidate geometry or score.
The network can stay fixed-shape. The scene does not have to.
DENSE ANSWER
Image features → spatial candidate records → decode → filter / NMS → returned set
SET-PREDICTION ANSWER
Image features + learned queries → one-to-one-trained candidate set → object records
What changes from image to image?
Both routes use fixed capacity to describe a variable scene. The next representation question is what to predict when a box is insufficient.
Teaching note
Both routes use fixed capacity to describe a variable scene. The next representation question is what to predict when a box is insufficient.
OPTIONAL REFERENCE
Three coordinate formats
Corners
Top-left and size
Centre and size
What would you check when implementing this?
The format must accompany the data. Centre coordinates are the means of corresponding corners. Width and height are corner differences. Do not confuse full sizes with half sizes.
Teaching note
The format must accompany the data. Centre coordinates are the means of corresponding corners. Width and height are corner differences. Do not confuse full sizes with half sizes.
OPTIONAL REFERENCE
Worked coordinate conversion
Top-left x, y, width, height in pixels.
Image: 1200 × 675 pixels.
CENTRE IN PIXELS
NORMALIZED CENTRE AND SIZE
What would you check when implementing this?
Continuous coordinates are used. Centre = top-left + half-size. Divide horizontal values by image width and vertical values by image height.
Teaching note
Continuous coordinates are used. Centre = top-left + half-size. Divide horizontal values by image width and vertical values by image height.
OPTIONAL REFERENCE
Letterboxing transforms the boxes too
Pad above and below to make a square.
What would you check when implementing this?
Scale the coordinates, then add the padding offsets. The example uses continuous geometry and symmetric vertical padding. Crops and flips also transform targets.
Teaching note
Scale the coordinates, then add the padding offsets. The example uses continuous geometry and symmetric vertical padding. Crops and flips also transform targets.
OPTIONAL REFERENCE
A minimal PyTorch target
boxes = torch.tensor([[130., 330., 370., 640.],
[760., 180., 1050., 620.]])
labels = torch.tensor([2, 1], dtype=torch.int64)
target = {"boxes": boxes, "labels": labels}Torchvision detection boxes use pixel corners.
What would you check when implementing this?
Check image dimensions, class IDs, dtypes and coordinate format. Transform the image and target together. Many torchvision detection models reserve class 0 for background.
Teaching note and source
Check image dimensions, class IDs, dtypes and coordinate format. Transform the image and target together. Many torchvision detection models reserve class 0 for background.
Primary sourceOPTIONAL REFERENCE
Target encoding and decoding are inverses
What would you check when implementing this?
This round trip uses the target offsets. The main inference example uses different predicted offsets. Width and height remain normalized to the image. Other architectures use different box parameterizations.
Teaching note
This round trip uses the target offsets. The main inference example uses different predicted offsets. Width and height remain normalized to the image. Other architectures use different box parameterizations.
OPTIONAL REFERENCE
Why YOLOv1 used 7 × 7 × 30
20 class probabilities per cell
Each box has x, y, w, h and confidence.
What would you check when implementing this?
Original YOLOv1 shares class probabilities between the two box predictors in a cell. Its confidence combines object presence and localization quality: Pr(object) × IoU. This is not the independent objectness term in our toy record.
Teaching note and source
Original YOLOv1 shares class probabilities between the two box predictors in a cell. Its confidence combines object presence and localization quality: Pr(object) × IoU. This is not the independent objectness term in our toy record.
Primary sourceOPTIONAL REFERENCE
Two predictors do not guarantee two objects
Original YOLOv1 shares one class vector per cell.
Read diagram labels
- One location; several geometries.
- The assignment rule still matters.
What would you check when implementing this?
The predictors specialize in geometry through training. Two independent objects with centres in the same cell are still difficult to represent. These original predictors are not predefined anchor templates.
Teaching note and source
The predictors specialize in geometry through training. Two independent objects with centres in the same cell are still difficult to represent. These original predictors are not predefined anchor templates.
Primary sourceOPTIONAL REFERENCE
Anchors provide reference geometries
Choose a reference shape, then predict offsets from it.
Read diagram labels
- One location; several geometries.
- The assignment rule still matters.
What would you check when implementing this?
Anchor-based heads can emit several records at one location. Assignment is model-specific and can use IoU and heuristics. Anchor-free detectors do not require these reference boxes.
Teaching note
Anchor-based heads can emit several records at one location. Assignment is model-specific and can use IoU and heuristics. Anchor-free detectors do not require these reference boxes.
OPTIONAL REFERENCE
A compact dense prediction head
head = nn.Conv2d(channels, anchors * (5 + classes), 1)
raw = head(features) # [batch, B*(5+C), Hf, Wf]
raw = raw.permute(0, 2, 3, 1)
raw = raw.reshape(batch, Hf, Wf, anchors, 5 + classes)Assignment, target encoding and decoding are separate parts.
What would you check when implementing this?
Illustrative code, not a complete detector. Define all dimensions consistently. Separate branches and extra regression channels are common.
Teaching note
Illustrative code, not a complete detector. Define all dimensions consistently. Separate branches and extra regression channels are common.
OPTIONAL REFERENCE
Feature pyramids support different object scales
- Fine maps retain spatial detail
- Coarse maps carry broader context
- Several levels contribute candidates
What would you check when implementing this?
Feature Pyramid Networks combine coarse semantic features with higher-resolution maps through a top-down pathway and lateral connections. Image resizing alone cannot guarantee small-object performance.
Teaching note and source
Feature Pyramid Networks combine coarse semantic features with higher-resolution maps through a top-down pathway and lateral connections. Image resizing alone cannot guarantee small-object performance.
Primary sourceOPTIONAL REFERENCE
The full binary cross-entropy expression
Positive target with p = 0.80
Background target with the same p
What would you check when implementing this?
Natural logarithms. A high probability is good for a positive and poor for background. Use numerically stable losses accepting logits, rather than taking logs manually.
Teaching note
Natural logarithms. A high probability is good for a positive and poor for background. Use numerically stable losses accepting logits, rather than taking logs manually.
OPTIONAL REFERENCE
Box losses measure geometric error
COORDINATE LOSSES
L1 or Smooth L1 compares coordinates.
OVERLAP LOSSES
IoU-based losses compare regions. GIoU also considers the enclosing box.
What would you check when implementing this?
The choice depends on parameterization and failure modes. Ordinary IoU offers limited guidance for disjoint boxes. GIoU adds enclosing-region geometry. There is no universal ranking of losses.
Teaching note and source
The choice depends on parameterization and failure modes. Ordinary IoU offers limited guidance for disjoint boxes. GIoU adds enclosing-region geometry. There is no universal ranking of losses.
Primary sourceOPTIONAL REFERENCE
An additional IoU calculation
Continuous corner coordinates.
What would you check when implementing this?
Clamp intersection width and height at zero for non-overlap. Validate corner ordering before computing areas.
Teaching note
Clamp intersection width and height at zero for non-overlap. Validate corner ordering before computing areas.
OPTIONAL REFERENCE
Class-aware NMS in Torchvision
from torchvision.ops import batched_nms
keep = batched_nms(boxes, scores, labels, iou_threshold=0.5)
boxes, scores = boxes[keep], scores[keep]
labels = labels[keep]Run separately for each image. Labels define suppression groups.
What would you check when implementing this?
batched_nms compares only matching group indices. Do not mix images under the same class IDs without image-specific groups. Decoding and score filtering are still separate operations.
Teaching note and source
batched_nms compares only matching group indices. Do not mix images under the same class IDs without image-specific groups. Decoding and score filtering are still separate operations.
Primary sourceOPTIONAL REFERENCE
Duplicate, wrong-class and low-IoU predictions
- A duplicate of an already claimed object is FP
- A wrong-class prediction is FP for that class
- A poorly located box can be FP while its target is FN
- An object with no matched prediction is FN
What would you check when implementing this?
Apply the declared benchmark protocol. One erroneous prediction and one missed target can occur at the same time.
Teaching note
Apply the declared benchmark protocol. One erroneous prediction and one missed target can occur at the same time.
OPTIONAL REFERENCE
Toy AP from two recall intervals
Continuous area under the interpolated precision envelope.
What would you check when implementing this?
The envelope is 1 up to recall 0.5 and 2/3 from 0.5 to 1. COCO samples 101 recall levels; do not label this exact continuous toy area as COCO AP.
Teaching note
The envelope is 1 up to recall 0.5 and 2/3 from 0.5 to 1. COCO samples 101 recall levels; do not label this exact continuous toy area as COCO AP.
OPTIONAL REFERENCE
Final counts can hide a different ranking
TP → FP → TP
TP → TP → FP
What would you check when implementing this?
Both complete lists have the same precision and recall. Putting both correct predictions first improves AP. This controlled example assumes the true positives match distinct targets.
Teaching note
Both complete lists have the same precision and recall. Putting both correct predictions first improves AP. This controlled example assumes the true positives match distinct targets.
OPTIONAL REFERENCE
A high-level API may already run NMS
with torch.inference_mode():
pred = model([image_tensor])[0]
keep = pred["scores"] >= 0.50
boxes = pred["boxes"][keep]
labels = pred["labels"][keep]Torchvision Faster R-CNN returns post-processed detections.
What would you check when implementing this?
The model includes proposals, decoding, internal score selection and NMS. This keep line adds a user cutoff. Inspect the API before applying another NMS stage.
Teaching note and source
The model includes proposals, decoding, internal score selection and NMS. This keep line adds a user cutoff. Inspect the API before applying another NMS stage.
Primary sourceOPTIONAL REFERENCE
One-stage and two-stage detector families
ONE STAGE
Predict classes and boxes directly from dense features.
TWO STAGE
Generate proposals, then classify and refine regions.
What would you check when implementing this?
This is a structural distinction, not a universal speed or accuracy ranking. Faster R-CNN is a two-stage example. Architecture, resolution, hardware and implementation determine practical performance.
Teaching note and source
This is a structural distinction, not a universal speed or accuracy ranking. Faster R-CNN is a two-stage example. Architecture, resolution, hardware and implementation determine practical performance.
Primary sourceOPTIONAL REFERENCE
Sources and further practice
What would you check when implementing this?
Photographs are generated teaching illustrations reused from the existing lecture. Their overlays, predictions and scores are constructed examples, not measured detector outputs. Original provenance remains in lecture9/figures/PROVENANCE.md.
Teaching note
Photographs are generated teaching illustrations reused from the existing lecture. Their overlays, predictions and scores are constructed examples, not measured detector outputs. Original provenance remains in lecture9/figures/PROVENANCE.md.