00 · WHAT CAN A VLM DO?

FROM CLIP TO
VISION-LANGUAGE MODELS

How Images Become Context
for Language Generation

How does visual information become context for next-token prediction?

Newfoundland dog from the ViT and CLIP lectures

ES 667 · Nipun Batra · IIT Gandhinagar

Oxford-IIIT Pet · CC BY-SA 4.001 / 155

Teaching notes

80-minute teaching route; progressive builds are physical slides. The model lab and measured records are separate from illustrative arithmetic.

Ask: ES 667 · Nipun Batra · IIT Gandhinagar

Answer: 80-minute teaching route; progressive builds are physical slides. The model lab and measured records are separate from illustrative arithmetic.

00 · WHAT CAN A VLM DO?

WHAT CAN A VLM DO?

00

Start with the evidence an answer needs.

Original teaching example / diagram02 / 155

Teaching notes

Start with the evidence an answer needs.

Ask:

Answer: Start with the evidence an answer needs.

00 · WHAT CAN A VLM DO?

Can the answer use visible evidence?

Is the laptop plugged in?

QUESTION

Is the laptop plugged in?

Original teaching example / diagram03 / 155

Teaching notes

A visible connection does not prove power is flowing.

Ask: Is the laptop plugged in?

Answer: Yes. A cable connects the laptop to the socket.

00 · WHAT CAN A VLM DO?

Can the answer use visible evidence?

Is the laptop plugged in?

REFERENCE ANSWER

Is the laptop plugged in?

Yes. A cable connects the laptop to the socket.

Original teaching example / diagram04 / 155

Teaching notes

A visible connection does not prove power is flowing.

Ask: Is the laptop plugged in?

Answer: Yes. A cable connects the laptop to the socket.

00 · WHAT CAN A VLM DO?

Can the answer use visible evidence?

Is the laptop plugged in?

ARCHITECTURAL REQUIREMENT

Is the laptop plugged in?

Yes. A cable connects the laptop to the socket.

Trace a connection between objects.

Original teaching example / diagram05 / 155

Teaching notes

A visible connection does not prove power is flowing.

Ask: Is the laptop plugged in?

Answer: Yes. A cable connects the laptop to the socket.

00 · WHAT CAN A VLM DO?

Identity is only part of the question

What is immediately left of the mug?

QUESTION

What is immediately left of the mug?

Original teaching example / diagram06 / 155

Teaching notes

The cube is also above the book; use this as a second oral question.

Ask: What is immediately left of the mug?

Answer: The red ball.

00 · WHAT CAN A VLM DO?

Identity is only part of the question

What is immediately left of the mug?

REFERENCE ANSWER

What is immediately left of the mug?

The red ball.

Original teaching example / diagram07 / 155

Teaching notes

The cube is also above the book; use this as a second oral question.

Ask: What is immediately left of the mug?

Answer: The red ball.

00 · WHAT CAN A VLM DO?

Identity is only part of the question

What is immediately left of the mug?

ARCHITECTURAL REQUIREMENT

What is immediately left of the mug?

The red ball.

Keep object identity and spatial relations.

Original teaching example / diagram08 / 155

Teaching notes

The cube is also above the book; use this as a second oral question.

Ask: What is immediately left of the mug?

Answer: The red ball.

00 · WHAT CAN A VLM DO?

The question selects a subset

How many apples are outside the bowl?

QUESTION

How many apples are outside the bowl?

Original teaching example / diagram09 / 155

Teaching notes

Five apples total; two in the bowl and three outside.

Ask: How many apples are outside the bowl?

Answer: Three.

00 · WHAT CAN A VLM DO?

The question selects a subset

How many apples are outside the bowl?

REFERENCE ANSWER

How many apples are outside the bowl?

Three.

Original teaching example / diagram10 / 155

Teaching notes

Five apples total; two in the bowl and three outside.

Ask: How many apples are outside the bowl?

Answer: Three.

00 · WHAT CAN A VLM DO?

The question selects a subset

How many apples are outside the bowl?

ARCHITECTURAL REQUIREMENT

How many apples are outside the bowl?

Three.

Represent multiple objects and the region that selects them.

Original teaching example / diagram11 / 155

Teaching notes

Five apples total; two in the bowl and three outside.

Ask: How many apples are outside the bowl?

Answer: Three.

00 · WHAT CAN A VLM DO?

Read first, then calculate

What is the invoice ID?
What percentage of the subtotal is GST?

QUESTION

What is the invoice ID?
What percentage of the subtotal is GST?

Original teaching example / diagram12 / 155

Teaching notes

Synthetic document. Separate text extraction errors from arithmetic errors.

Ask: What is the invoice ID? What percentage of the subtotal is GST?

Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.

00 · WHAT CAN A VLM DO?

Read first, then calculate

What is the invoice ID?
What percentage of the subtotal is GST?

REFERENCE ANSWER

What is the invoice ID?
What percentage of the subtotal is GST?

ES667-042. GST is 180 / 1000 × 100 = 18%.

Original teaching example / diagram13 / 155

Teaching notes

Synthetic document. Separate text extraction errors from arithmetic errors.

Ask: What is the invoice ID? What percentage of the subtotal is GST?

Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.

00 · WHAT CAN A VLM DO?

Read first, then calculate

What is the invoice ID?
What percentage of the subtotal is GST?

ARCHITECTURAL REQUIREMENT

What is the invoice ID?
What percentage of the subtotal is GST?

ES667-042. GST is 180 / 1000 × 100 = 18%.

Preserve fine text, numbers and their roles.

Original teaching example / diagram14 / 155

Teaching notes

Synthetic document. Separate text extraction errors from arithmetic errors.

Ask: What is the invoice ID? What percentage of the subtotal is GST?

Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.

00 · WHAT CAN A VLM DO?

The labels change the answer

Which month is highest?
How much higher is March than January?

QUESTION

Which month is highest?
How much higher is March than January?

Original teaching example / diagram15 / 155

Teaching notes

Synthetic values: January 30, February 45, March 75, April 40.

Ask: Which month is highest? How much higher is March than January?

Answer: March. The difference is 45 µg/m³.

00 · WHAT CAN A VLM DO?

The labels change the answer

Which month is highest?
How much higher is March than January?

REFERENCE ANSWER

Which month is highest?
How much higher is March than January?

March. The difference is 45 µg/m³.

Original teaching example / diagram16 / 155

Teaching notes

Synthetic values: January 30, February 45, March 75, April 40.

Ask: Which month is highest? How much higher is March than January?

Answer: March. The difference is 45 µg/m³.

00 · WHAT CAN A VLM DO?

The labels change the answer

Which month is highest?
How much higher is March than January?

ARCHITECTURAL REQUIREMENT

Which month is highest?
How much higher is March than January?

March. The difference is 45 µg/m³.

Bind chart labels to values and units.

Original teaching example / diagram17 / 155

Teaching notes

Synthetic values: January 30, February 45, March 75, April 40.

Ask: Which month is highest? How much higher is March than January?

Answer: March. The difference is 45 µg/m³.

00 · WHAT CAN A VLM DO?

A diagram supplies conditions

What is x?

QUESTION

What is x?

Original teaching example / diagram18 / 155

Teaching notes

The right-angle marker, rather than appearance alone, licenses Pythagoras.

Ask: What is x?

Answer: x = 5, using the marked right angle and sides 3 and 4.

00 · WHAT CAN A VLM DO?

A diagram supplies conditions

What is x?

REFERENCE ANSWER

What is x?

x = 5, using the marked right angle and sides 3 and 4.

Original teaching example / diagram19 / 155

Teaching notes

The right-angle marker, rather than appearance alone, licenses Pythagoras.

Ask: What is x?

Answer: x = 5, using the marked right angle and sides 3 and 4.

00 · WHAT CAN A VLM DO?

A diagram supplies conditions

What is x?

ARCHITECTURAL REQUIREMENT

What is x?

x = 5, using the marked right angle and sides 3 and 4.

Combine diagram evidence with a mathematical rule.

Original teaching example / diagram20 / 155

Teaching notes

The right-angle marker, rather than appearance alone, licenses Pythagoras.

Ask: What is x?

Answer: x = 5, using the marked right angle and sides 3 and 4.

00 · WHAT CAN A VLM DO?

Compare presence and position

What changed from A to B?

QUESTION

What changed from A to B?

Original teaching example / diagram21 / 155

Teaching notes

Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.

Ask: What changed from A to B?

Answer: The chair moved to the right; the cup disappeared.

00 · WHAT CAN A VLM DO?

Compare presence and position

What changed from A to B?

REFERENCE ANSWER

What changed from A to B?

The chair moved to the right; the cup disappeared.

Original teaching example / diagram22 / 155

Teaching notes

Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.

Ask: What changed from A to B?

Answer: The chair moved to the right; the cup disappeared.

00 · WHAT CAN A VLM DO?

Compare presence and position

What changed from A to B?

ARCHITECTURAL REQUIREMENT

What changed from A to B?

The chair moved to the right; the cup disappeared.

Compare evidence across images.

Original teaching example / diagram23 / 155

Teaching notes

Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.

Ask: What changed from A to B?

Answer: The chair moved to the right; the cup disappeared.

00 · WHAT CAN A VLM DO?

A screenshot is a spatial interface

Where should I click to disable notifications?

QUESTION

Where should I click to disable notifications?

Original teaching example / diagram24 / 155

Teaching notes

Describing a control is distinct from executing a click.

Ask: Where should I click to disable notifications?

Answer: The blue toggle on the Notifications row.

00 · WHAT CAN A VLM DO?

A screenshot is a spatial interface

Where should I click to disable notifications?

REFERENCE ANSWER

Where should I click to disable notifications?

The blue toggle on the Notifications row.

Original teaching example / diagram25 / 155

Teaching notes

Describing a control is distinct from executing a click.

Ask: Where should I click to disable notifications?

Answer: The blue toggle on the Notifications row.

00 · WHAT CAN A VLM DO?

A screenshot is a spatial interface

Where should I click to disable notifications?

ARCHITECTURAL REQUIREMENT

Where should I click to disable notifications?

The blue toggle on the Notifications row.

Locate a control relative to its label.

Original teaching example / diagram26 / 155

Teaching notes

Describing a control is distinct from executing a click.

Ask: Where should I click to disable notifications?

Answer: The blue toggle on the Notifications row.

00 · WHAT CAN A VLM DO?

The image can be a scientific figure

What is the hottest measured region?

QUESTION

What is the hottest measured region?

Original teaching example / diagram27 / 155

Teaching notes

Synthetic test-plate measurements; count rows and columns from the top left.

Ask: What is the hottest measured region?

Answer: Row 3, column 4: 62 °C.

00 · WHAT CAN A VLM DO?

The image can be a scientific figure

What is the hottest measured region?

REFERENCE ANSWER

What is the hottest measured region?

Row 3, column 4: 62 °C.

Original teaching example / diagram28 / 155

Teaching notes

Synthetic test-plate measurements; count rows and columns from the top left.

Ask: What is the hottest measured region?

Answer: Row 3, column 4: 62 °C.

00 · WHAT CAN A VLM DO?

The image can be a scientific figure

What is the hottest measured region?

ARCHITECTURAL REQUIREMENT

What is the hottest measured region?

Row 3, column 4: 62 °C.

Read a scale and local measurements.

Original teaching example / diagram29 / 155

Teaching notes

Synthetic test-plate measurements; count rows and columns from the top left.

Ask: What is the hottest measured region?

Answer: Row 3, column 4: 62 °C.

00 · WHAT CAN A VLM DO?

The question can be wrong

What color is the bicycle?

QUESTION

What color is the bicycle?

Original teaching example / diagram30 / 155

Teaching notes

Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.

Ask: What color is the bicycle?

Answer: No bicycle is visible.

00 · WHAT CAN A VLM DO?

The question can be wrong

What color is the bicycle?

REFERENCE ANSWER

What color is the bicycle?

No bicycle is visible.

Original teaching example / diagram31 / 155

Teaching notes

Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.

Ask: What color is the bicycle?

Answer: No bicycle is visible.

00 · WHAT CAN A VLM DO?

The question can be wrong

What color is the bicycle?

ARCHITECTURAL REQUIREMENT

What color is the bicycle?

No bicycle is visible.

Let visual evidence override the wording of the question.

Original teaching example / diagram32 / 155

Teaching notes

Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.

Ask: What color is the bicycle?

Answer: No bicycle is visible.

00 · WHAT CAN A VLM DO?

What must survive the visual representation?

Question familyInformation needed
Counting / spatialMultiple objects and their arrangement
OCR / chartsFine detail, labels, values and units
UI / multiple imagesLocations and correspondence across views
False premiseEvidence that can contradict the prompt

Would one global CLIP vector always be enough?

Original teaching example / diagram33 / 155

Teaching notes

A global vector can be useful, but fine-grained tasks motivate exposing richer visual states. This is a design motivation, not an impossibility theorem.

Ask: Would one global CLIP vector always be enough?

Answer: A global vector can be useful, but fine-grained tasks motivate exposing richer visual states. This is a design motivation, not an impossibility theorem.

00 · WHAT CAN A VLM DO?

Try the same questions on a small model

Newfoundland dog

Predict → run → inspect the exact answer.

Open the VLM lab ↗

11 image presets, recorded outputs, and an image-swap experiment.

A reference answer is not a model prediction.

SmolVLM-256M · model card · Official Transformers.js browser example34 / 155

Teaching notes

Preload the model before class. First download needs network; recorded outputs are available offline. Run the dog, spatial and absent-object cases.

Ask: A reference answer is not a model prediction.

Answer: Preload the model before class. First download needs network; recorded outputs are available offline. Run the dog, spatial and absent-object cases.

01 · FROM MATCHING TO GENERATION

FROM MATCHING TO GENERATION

01

Matching supplied text leaves a generation problem.

Original teaching example / diagram35 / 155

Teaching notes

Matching supplied text leaves a generation problem.

Ask:

Answer: Matching supplied text leaves a generation problem.

01 · FROM MATCHING TO GENERATION

Recall: CLIP compares an image with supplied text

CLIP matching; remove supplied text to expose the missing language generatorImage encoderuSupplied textText encodervMatch score

Image vector u · text vector v · similarity uᵀv

CLIP · Radford et al., 202136 / 155

Teaching notes

Vectors are normalized for the cosine-style comparison. CLIP is a VLM broadly; here we distinguish contrastive matching from generative VLMs.

Ask: Image vector u · text vector v · similarity uᵀv

Answer: Vectors are normalized for the cosine-style comparison. CLIP is a VLM broadly; here we distinguish contrastive matching from generative VLMs.

01 · FROM MATCHING TO GENERATION

Now remove the candidate description

CLIP matching; remove supplied text to expose the missing language generatorImage encoderuNo candidate description supplied.Where do new words come from?

Desired output: ? ? ?

Original teaching example / diagram37 / 155

Teaching notes

The image encoder remains useful. But a matching score alone cannot produce a next-token vocabulary distribution.

Ask: Desired output: ? ? ?

Answer: The image encoder remains useful. But a matching score alone cannot produce a next-token vocabulary distribution.

01 · FROM MATCHING TO GENERATION

We already know how to generate words

Recall the causal language decoder before introducing a visual connectorQuestion / prefixCausal language modelVocabulary → next token

A causal language model maps a prefix to a next-token distribution.

Original teaching example / diagram38 / 155

Teaching notes

Recall the decoder-only model from Beyond Attention before attaching vision.

Ask: A causal language model maps a prefix to a next-token distribution.

Answer: Recall the decoder-only model from Beyond Attention before attaching vision.

01 · FROM MATCHING TO GENERATION

We have two useful pieces

Recall the causal language decoder before introducing a visual connectorVisual representationQuestion / prefixCausal language modelVocabulary → next tokenHow do we connect them?

How can visual states enter the language computation?

Original teaching example / diagram39 / 155

Teaching notes

Keep the two branches disconnected. Do not draw a connector until we derive what it must do.

Ask: How can visual states enter the language computation?

Answer: Keep the two branches disconnected. Do not draw a connector until we derive what it must do.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

HOW CAN VISION ENTER A LANGUAGE MODEL?

02

First, see the three ideas. Then build each connection.

Original teaching example / diagram40 / 155

Teaching notes

First, see the three ideas. Then build each connection.

Ask:

Answer: First, see the three ideas. Then build each connection.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Idea 1 · prepend the image as visual tokens

All N projected patch vectors are concatenated with embeddings of question tokens and answer tokens generated so far, then passed through causal self-attention, decoder updates and a vocabulary headIMAGE PATHVision encoder (ViT)Full self-attention over patchesProject every rowV1V2VN…All N patch rowsTEXT PATHWhat animal is shown?+ answer so far (empty at first)Tokenizer → IDsEmbeddingsONE INPUT SEQUENCE OF VECTORSQuestion embeddingsAnswer-so-far embeddingsCausal self-attention: image + textMLP + residual updatesDECODERText reads earlierimage and text rows.Vocabulary headNext token

Keep all N patch vectors in this design; N depends on the image grid.

LLaVA · Liu et al., 202341 / 155

Teaching notes

Trace both paths before naming the join. The vision encoder produces one contextual feature row per patch; a row-wise projector adapts every row to the decoder width. V1, V2, …, VN represents the full N-row sequence, not a selection of four patches. The question and answer generated so far are token IDs with embedding vectors; concatenate their embeddings after the image vectors. At the first step, no answer tokens exist. Generation appends one token at a time; cached computation can reuse earlier states. Chat/system markers and positional handling are omitted here. No separate text encoder is added: tokenization and embedding lookup feed the causal language decoder. Show all three ideas before deriving patch grids, widths or weights.

Ask: Do we keep only four patches, and does the decoder receive raw question text?

Answer: No. This basic prefix design keeps all N patch rows and projects each one. The question and already-generated answer tokens are embedded; their vectors follow V1, V2, …, VN in one decoder input sequence. Before the first prediction, the answer portion is empty.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Idea 2 · let the decoder look at a visual memory

A vision encoder supplies N feature vectors as separate memory. Question and answer-so-far tokens become text embeddings. The language decoder uses causal self-attention for text and cross-attention to read image memory before decoder updates and the vocabulary head predict the next token.IMAGE PATHVision encoder (ViT)Full self-attention over patchesVISUAL MEMORYh1h2h3hN…The encoder outputs stay available.224 × 224 image, 16 × 16 patches:196 patch rows; 197 if CLS is kept.N depends on the encoder setup.TEXT PATHWhat animal is shown?+ answer so far (empty at first)Tokenizer → IDs → embeddingse1e2eT…TEXT DECODERCausal self-attention: read textCross-attention: read imageSequence-to-sequence:the image is the source.MLP + residual updatesVocabulary head → next token

As in translation: encode the source, then let the text decoder read its states.

Transformer · Vaswani et al., 2017 · ViT · Dosovitskiy et al., 202042 / 155

Teaching notes

Visual memory is the sequence produced by the vision encoder, not a second encoder or a separate learned memory bank. h1, h2, …, hN abbreviates all N retained image feature rows. The question and answer-so-far tokens become e1, e2, …, eT through embedding lookup; T is the current text length. The answer is empty for the first prediction and grows only from generated token IDs. A causal decoder needs cross-attention sublayers to use this route: text states supply queries, visual features supply keys and values, and the attention result updates text states before the vocabulary head scores the next token. This is the sequence-to-sequence connection from English-to-Hindi translation, with the image as source. The overview shows the two attention operations and the MLP/residual updates; normalization, projection details, positional handling, chat markers and repeated blocks are omitted. The example uses a 224×224 image and 16×16 patches: 14×14=196 patch rows; retaining the global CLS row gives 197. Neither is universal. Compression can shorten this memory later; Flamingo also resamples features and uses gated cross-attention, so this is not its exact architecture.

Ask: Is visual memory another model? Does cross-attention itself produce the answer word?

Answer: Visual memory is the encoder’s sequence of image features. The language decoder reads it through cross-attention, like a translation decoder reading source-language states. That updates text states; the vocabulary head then scores the next token. N is not fixed: a 224×224 image with 16×16 patches gives 196 patch rows, or 197 rows if CLS is retained.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Idea 3 · summarize the image, then use either connection

The image becomes all N ViT patch features. M learned queries cross-attend to those features to produce M image-dependent summaries. Alternative A projects the summaries and prepends them to text embeddings for causal self-attention. Alternative B keeps summaries as memory that the text decoder reads through cross-attention before its vocabulary head.ViT encoderFull self-attentionh1h2hN…All N patch featuresLearned summarizerCross-attentionM learned query vectorsQueries read all patch rows.S1…SMM < N summariesImage-dependentA · Prefix: concatenate vectorsProject summariesS1 … SMQuestion + answer so fartokenize → embedv1 … vMe1 … eTDecoder: causal self-attentionVocabulary head → next tokenB · Decoder reads separate memoryQuestion + answer so fartokenize → embed → e1 … eTSummary memoryS1 … SMCausal self-attentionK, VQCross-attentionVocabulary head → next token

Compression reduces N image rows to M summaries; either decoder connection remains possible.

BLIP-2 · Li et al., 2023 · Flamingo · Alayrac et al., 202243 / 155

Teaching notes

Read the common image path first: the ViT encoder uses full self-attention and exposes all N contextual patch rows. A learned-query summarizer uses M shared trainable query vectors to cross-attend to those image features, producing M image-dependent summaries. This is a representative resampler mechanism, not a subset selection and not a complete implementation of a named model. N > M; omitted boxes stand for intermediate rows. Now choose one of two alternative decoder connections. A projects each summary to the decoder width, concatenates those vectors before text embeddings, and uses causal self-attention over the joined sequence. Projection changes width, not the already-reduced M row count. B retains S1…SM as separate memory: text embeddings enter causal self-attention, then text states supply Q while summaries supply K,V for decoder cross-attention. The vocabulary head produces next-token scores in both routes. The compression cross-attention uses learned queries; decoder cross-attention uses text-derived queries. Question tokens and previously generated answer tokens are embedded; the answer starts empty. v labels are projected summary vectors; e labels are text embeddings. A and B are alternatives, not consecutive stages. Decoder MLPs, residuals, normalization, position handling and repeated blocks are abstracted in this overview and restored in the deep dives. Source notes discuss BLIP-2 and Flamingo as examples with additional model-specific components.

Ask: After producing M summaries, must language read them with cross-attention?

Answer: No. In route A, project the M summaries, prepend them to the text embeddings, and use decoder causal self-attention. In route B, keep the M summaries separate and let decoder cross-attention read them. The summarizer itself can use learned-query cross-attention in either route; that is a different operation from the decoder’s text-query cross-attention.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The map: two access routes, one compression choice

IdeaHigh-level decisionThen we will explain…
1 · Visual prefixPut visual vectors before text.How to make them fit the decoder input.
2 · Separate memoryLet language read the image.How a text state retrieves visual information.
3 · Shorter summaryUse fewer visual vectors.How learned readers make the summaries.

Ideas 1 and 2 choose the access route. Idea 3 can be used with either.

Original teaching example / diagram44 / 155

Teaching notes

Spend roughly five minutes on the three preceding idea slides, with no matrix shapes or attention arithmetic. Ask students to explain each idea in one sentence. Now return to Idea 1 and derive it. Ideas 2 and 3 get their own later deep dives.

Ask: Must a model choose only one of these three ideas?

Answer: No. Prefix and separate memory are two ways to access vision. Compression can be combined with either; for example, a short visual summary can be prepended to the text.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Idea 1 in detail · build the visual prefix

Complete prefix architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderProject each visual vectorV1V2…VNSame rows, new width.All N patch rows; middle rows omitted.LANGUAGE: GROWING TEXTQuestion + answer so farV1V2…VNText embeddingsCausal self-attention + residualMLP + residualVocabulary head → next token

One sequence; the same causal self-attention can use both kinds of context.

LLaVA · Liu et al., 202345 / 155

Teaching notes

Trace all N patch rows through the amber projector into the input row. Trace question and answer-so-far tokenization and embedding into the same row. The answer is initially empty, then grows by generated tokens. V1, V2, …, VN abbreviates the full visual sequence. Point to causal self-attention, MLP and vocabulary head from the previous lecture. Positions, normalization and exact chat markers are omitted. We will now unpack only the new visual path.

Ask: One sequence; the same causal self-attention can use both kinds of context.

Answer: Trace all N patch rows through the amber projector into the input row. Trace question and answer-so-far tokenization and embedding into the same row. The answer is initially empty, then grows by generated tokens. V1, V2, …, VN abbreviates the full visual sequence. Point to causal self-attention, MLP and vocabulary head from the previous lecture. Positions, normalization and exact chat markers are omitted. We will now unpack only the new visual path.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

What must we build to make this work?

JobWhy it is needed
1. Keep image feature vectorsThe decoder needs access to visual evidence.
2. Match the decoder input widthAll input rows must have the same number of coordinates.
3. Place them before the questionCausal attention can read earlier context.

We already have the decoder. Now build its visual input.

Original teaching example / diagram46 / 155

Teaching notes

Return to this three-job plan after the projector calculation and after the mask. Students should be able to point to each job in the preceding complete architecture.

Ask: We already have the decoder. Now build its visual input.

Answer: Return to this three-job plan after the projector calculation and after the mask. Students should be able to point to each job in the preceding complete architecture.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Recall ViT: every patch already has an output row

A numbered twelve-patch image enters a ViT encoder. All twelve contextual feature rows h1 through h12 are shown.12345678910111212 patches in this exampleIllustrative 3 × 4 gridViT encoderFull self-attentionALL 12 PATCH ROWSh1h2h3h4h5h6h7h8h9h10h11h12Each row is a feature vector.Each patch row can use the whole image.

Patch 1 has row h1, patch 2 has row h2, and so on through patch 12.

ViT · Dosovitskiy et al., 202047 / 155

Teaching notes

The numbered 3 × 4 grid is an illustrative twelve-patch example, not a model configuration. All twelve patch rows are drawn, in row-major order matching the image labels. Each box denotes one whole feature vector, not one scalar coordinate. ViT embeds the patches, adds positions, and updates their states with full self-attention. The row at patch position i can therefore contain information from the whole image. The CLS row is introduced on the next slide.

Ask: Patch 1 has row h1, patch 2 has row h2, and so on through patch 12.

Answer: The numbered 3 × 4 grid is an illustrative twelve-patch example, not a model configuration. All twelve patch rows are drawn, in row-major order matching the image labels. Each box denotes one whole feature vector, not one scalar coordinate. ViT embeds the patches, adds positions, and updates their states with full self-attention. The row at patch position i can therefore contain information from the whole image. The CLS row is introduced on the next slide.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

ViT outputs a CLS state and one row per patch

A numbered twelve-patch image enters a ViT encoder. All twelve contextual feature rows h1 through h12 are shown, along with the separate CLS summary state.12345678910111212 patches in this exampleIllustrative 3 × 4 gridViT encoderFull self-attentionALL 12 PATCH ROWSh1h2h3h4h5h6h7h8h9h10h11h12Each row is a feature vector.CLS · summary stateOne whole-image summary

Here: 1 CLS row + 12 patch rows. We keep all 12 patch rows.

ViT · Dosovitskiy et al., 202048 / 155

Teaching notes

Use the same twelve numbered image patches and twelve output rows as the previous slide. This standard ViT also has a CLS token, introduced before the encoder; its output state provides a whole-image summary. CLS is an extra sequence position, not a thirteenth image patch or a separate encoder. The diagram abstracts patch embedding, positions, and the input CLS token inside the ViT path. Its twelve patch states are contextual vectors, not isolated patch descriptors. For the visual-prefix design being developed, retain all twelve patch rows and then project each one to the decoder width; the CLS state is not included in that patch sequence. Twelve is the complete illustrative example, not a universal model count. Keep this explanation about ViT outputs; the earlier CLIP matching discussion need not be repeated.

Ask: Here: 1 CLS row + 12 patch rows. We keep all 12 patch rows.

Answer: Use the same twelve numbered image patches and twelve output rows as the previous slide. This standard ViT also has a CLS token, introduced before the encoder; its output state provides a whole-image summary. CLS is an extra sequence position, not a thirteenth image patch or a separate encoder. The diagram abstracts patch embedding, positions, and the input CLS token inside the ViT path. Its twelve patch states are contextual vectors, not isolated patch descriptors. For the visual-prefix design being developed, retain all twelve patch rows and then project each one to the decoder width; the CLS state is not included in that patch sequence. Twelve is the complete illustrative example, not a universal model count. Keep this explanation about ViT outputs; the earlier CLIP matching discussion need not be repeated.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Why keep more than the global summary?

Invoice from the opening examples

Recall our opening question

What is the invoice ID?
What percentage is GST?

The answer needs small text, several numbers, and their roles.

Give language access to more visual evidence before asking it to answer.

Original teaching example / diagram49 / 155

Teaching notes

Return to the Newfoundland image on the next frame. This invoice explains the design motivation. Keeping patch states does not guarantee OCR success; resolution and training still matter.

Ask: Give language access to more visual evidence before asking it to answer.

Answer: Return to the Newfoundland image on the next frame. This invoice explains the design motivation. Keeping patch states does not guarantee OCR success; resolution and training still matter.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

A “visual state” is just one row of features

One image → many rows

HV=[h1h2⋮hN]H_V=\begin{bmatrix}h_1\\h_2\\\vdots\\h_N\end{bmatrix}

Read the notation in plain language

N = number of visual rows.

dV = numbers in each row.

hᵢ = the feature row at patch position i.

Rows are vectors, not object labels or named semantic attributes.

Original teaching example / diagram50 / 155

Teaching notes

Use “row” and “vector” interchangeably here. Attention has mixed image context into each row. Do not label one row “dog” or claim that one coordinate is fur colour.

Ask: Rows are vectors, not object labels or named semantic attributes.

Answer: Use “row” and “vector” interchangeably here. Attention has mixed image context into each row. Do not label one row “dog” or claim that one coordinate is fur colour.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Recall: the decoder already consumes vectors

Word IDs become embedding vectors before the decoder; image vectors enter at the same input interfaceWhat animal is shown?Tokenizer → IDsEmbedding lookupe1e2e3e4e5e6Inside the decoder, text is already a sequence of vectors.Visual vectors can enter at this vector interface.

Text IDs become embeddings before reaching the Transformer blocks.

Original teaching example / diagram51 / 155

Teaching notes

This is the connection to the earlier text architecture: tokenization selects embedding rows. Image vectors bypass a word-ID lookup and enter at the vector interface. They also need the implementation’s positional treatment.

Ask: Text IDs become embeddings before reaching the Transformer blocks.

Answer: This is the connection to the earlier text architecture: tokenization selects embedding rows. Image vectors bypass a word-ID lookup and enter at the vector interface. They also need the implementation’s positional treatment.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The image rows have 2 coordinates; text rows have 3

The dog image supplies N toy two-coordinate image rows. The question and answer so far supply T toy three-coordinate text embeddings. Their widths differ, so they cannot yet be stacked as decoder input.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Can these rows enter the same decoder?h121e11022 coordinates per image row3 coordinates per text row≠ widthWe need to turn each image row into a 3-coordinate vector.

To form one input sequence, every vector must have the decoder’s input width.

Original teaching example / diagram52 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Ask students to point to one row, then count its coordinates. The image and text may have different numbers of rows; the problem is their different row widths.

Ask: Do N and T need to match, or must the number of coordinates per row match?

Answer: The row widths must match. N image rows and T text rows can have different counts. After projecting the image rows to width 3, concatenate N + T rows for the decoder.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Add a projector to the image path

The same image and text remain visible. N by 2 image features multiply one shared 2 by 3 projector to give N by 3 visual vectors. T by 3 text embeddings remain unchanged.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Insert a shared projector on the image pathImage features HᵥN × 22 × 3Projector WₚN × 3Visual vectors ZText stays ET × 3Z = Hᵥ Wₚ · Keep N rows. Change 2 coordinates to 3.

The projector changes N × 2 image features into N × 3 visual vectors.

Original teaching example / diagram53 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Read each shape beneath the corresponding visible module. The text embeddings are already T × 3 and remain unchanged. In general Hᵥ is N × dV, Wₚ is dV × dL, and Z = HᵥWₚ is N × dL. Bias is omitted. Real connectors can use an MLP. Matching shapes only makes the interface valid; learning must make the projected contents useful.

Ask: The projector changes N × 2 image features into N × 3 visual vectors.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Read each shape beneath the corresponding visible module. The text embeddings are already T × 3 and remain unchanged. In general Hᵥ is N × dV, Wₚ is dV × dL, and Z = HᵥWₚ is N × dL. Bias is omitted. Real connectors can use an MLP. Matching shapes only makes the interface valid; learning must make the projected contents useful.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Zoom into one row from the image

Projector worked example, step 1. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. h1 is one of the image rows above. Its two coordinates are 2 and 1.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×Shared projectorWₚ2 × 3=???z11 × 3: same width as e1h1 is one of the image rows above. Its two coordinates are 2 and 1.

Take one 2-coordinate row from the image feature sequence.

Original teaching example / diagram54 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: Take one 2-coordinate row from the image feature sequence.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Use a 2 × 3 matrix to change the width

Projector worked example, step 2. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. Each of the three columns of Wₚ will make one output coordinate.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=???z11 × 3: same width as e1Each of the three columns of Wₚ will make one output coordinate.

Each column makes one of the 3 decoder-input coordinates.

Original teaching example / diagram55 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: Each column makes one of the 3 decoder-input coordinates.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The first column makes the first coordinate

Projector worked example, step 3. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. First column: how much of 2, and how much of 1, should we use?IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=???z11 × 3: same width as e1First column: how much of 2, and how much of 1, should we use?

Which input coordinates does this column combine?

Original teaching example / diagram56 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: Which input coordinates does this column combine?

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The first output coordinate is 2

Projector worked example, step 4. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. First coordinate: 2 × 1 + 1 × 0 = 2IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=2??z11 × 3: same width as e1First coordinate: 2 × 1 + 1 × 0 = 2

The first output coordinate copies the first input in this toy map.

Original teaching example / diagram57 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: The first output coordinate copies the first input in this toy map.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The second column reuses the same image row

Projector worked example, step 5. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. Second column: use the same input [2, 1] again.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=2??z11 × 3: same width as e1Second column: use the same input [2, 1] again.

Now use the same input row with the second column.

Original teaching example / diagram58 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: Now use the same input row with the second column.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The second output coordinate is 1

Projector worked example, step 6. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. Second coordinate: 2 × 0 + 1 × 1 = 1IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=21?z11 × 3: same width as e1Second coordinate: 2 × 0 + 1 × 1 = 1

The second output coordinate copies the second input in this toy map.

Original teaching example / diagram59 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: The second output coordinate copies the second input in this toy map.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

The third column combines both input coordinates

Projector worked example, step 7. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. Third coordinate: 2 × 1 + 1 × 1 = 3IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=213z11 × 3: same width as e1Third coordinate: 2 × 1 + 1 × 1 = 3

The third output coordinate adds the two inputs in this toy map.

Original teaching example / diagram60 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: The third output coordinate adds the two inputs in this toy map.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Now this image vector has the required width

Projector worked example, step 8. The same image, encoder, image feature table, question, answer prefix and text embedding table remain visible. The image row [2, 1] has become [2, 1, 3]. The text rows stay as they are.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Zoom into h1211 × 2×101011Wₚ2 × 3=213z11 × 3: same width as e1The image row [2, 1] has become [2, 1, 3]. The text rows stay as they are.

The row [2, 1, 3] fits beside the 3-coordinate text embeddings.

Original teaching example / diagram61 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

Ask: The row [2, 1, 3] fits beside the 3-coordinate text embeddings.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Use the same matrix for all N image rows

Apply the same 2 by 3 matrix to every image row. [2,1] becomes [2,1,3], [1,3] becomes [1,3,4], and [0,2] becomes [0,2,2]. All N rows remain. All projected rows now have the same width as the text embeddings.IMAGE → FEATURE VECTORSViT encoderAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Every image rowh121h213…⋮⋮hN02N × 22 × 3×Same Wₚ101011N × 3=All N projected rowsz1213z2134…⋮⋮⋮zN022Width now matches EN is unchanged.

N rows stay N rows; every projected row now has 3 coordinates.

Original teaching example / diagram62 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The same Wₚ gives [2,1,3] from [2,1], [1,3,4] from [1,3], and [0,2,2] from [0,2]. No patch is dropped and no independent projector is created for each patch. Text E is unchanged.

Ask: N rows stay N rows; every projected row now has 3 coordinates.

Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The same Wₚ gives [2,1,3] from [2,1], [1,3,4] from [1,3], and [0,2,2] from [0,2]. No patch is dropped and no independent projector is created for each patch. Text E is unchanged.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Check: did the projector create more visual tokens?

Before

576 visual rows.
Each row has 1,024 coordinates.

After a row-wise projector

Each row has 4,096 coordinates.

How many rows now?

Predict the row count before advancing.

Original teaching example / diagram63 / 155

Teaching notes

These are illustrative model widths. The operation acts on the last dimension, not the row dimension.

Ask: 576 rows × 1,024 coordinates become how many rows × 4,096 coordinates?

Answer: 576 rows. The projector changes each vector’s width; it does not compress or select patch positions.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Still 576 rows: width and count are different

HV⏟576×1024  WP⏟1024×4096=Z⏟576×4096\underbrace{H_V}_{576\times1024}\;\underbrace{W_P}_{1024\times4096}=\underbrace{Z}_{576\times4096}

Same positions. Different number of features at each position.

We will learn how to reduce the row count in the compression section.

Original teaching example / diagram64 / 155

Teaching notes

If a student answered 32, explain that they have proposed an extra summarization operation. The simple LLaVA-style row-wise projection does not do that.

Ask: We will learn how to reduce the row count in the compression section.

Answer: If a student answered 32, explain that they have proposed an extra summarization operation. The simple LLaVA-style row-wise projection does not do that.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Now the two sequences fit one decoder input

The image rows are projected to width three and stacked before width-three text embeddings, including the question and the answer generated so far. The ViT uses full attention. The joined N plus T rows feed a language decoder whose causal mask applies from the first projected image row z1, including both image and text positions, followed by a vocabulary head. Projected visual rows are continuous features, not word IDs.IMAGE → FEATURE VECTORSViT encoderFull attentionAll N patch rowsh121h213…⋮⋮hN02Hᵥ: N × 2TEXT → EMBEDDING VECTORSWhat animal is shown?Answer so far: “A black”Tokenize → IDs → embede1102e2021…⋮⋮⋮eT210E: T × 3Toy widths and values; not measured from this image or text. Middle rows omitted.Image → shared Wₚ → ZN projected rowsz1213…⋮⋮⋮zN022+ T text rows(N + T) × 3e1102…⋮⋮⋮eT210Language decoderCausal from z1 onwardThe decoder mask covers image AND text rows.Head → next token

The mask applies inside the decoder, including its visual positions.

Original teaching example / diagram65 / 155

Teaching notes

The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Z has N rows and E has T rows; their joined input is (N + T) × 3 in this toy example. Distinguish the two attention stages: the ViT uses full image self-attention before projection; the standard causal decoder applies a lower-triangular mask across every joined position, starting at z1 rather than switching it on only at the first text position. At a decoder layer, z1 reads itself, z2 reads z1 and itself, and the first text row reads all N preceding visual rows and itself. The visual vectors already contain whole-image context from the ViT, so a causal decoder does not undo that encoder context. With the image and prompt supplied, all their positions can be computed in parallel within each decoder layer during prefill; this does not generate image patches one at a time. Only new answer tokens are generated autoregressively. This is the basic LLaVA-style causal-prefix design, not a universal rule: some other VLMs allow bidirectional attention within an image block. System/role markers and positions are abstracted. Visual rows remain continuous vectors, not vocabulary IDs.

Ask: Does the causal mask start only when the decoder reaches the first text row?

Answer: In this design it covers the whole decoder sequence, including visual positions from z1. The ViT earlier used full image attention. Known image and prompt rows are processed together in a masked prefill pass; later answer tokens are generated one by one. The upcoming mask shows exactly which visual and text positions can read each other.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Build the input row: one projected visual vector

Progressive visual prefix: visual vectors before user prompt and current assistant positionV1Visual vectors

Visual vectors and text embeddings share the decoder’s input width.

Original teaching example / diagram66 / 155

Teaching notes

Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

Ask: Visual vectors and text embeddings share the decoder’s input width.

Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Keep the rest of the projected visual sequence

Progressive visual prefix: visual vectors before user prompt and current assistant positionV1V2V3Visual vectors

Visual vectors and text embeddings share the decoder’s input width.

Original teaching example / diagram67 / 155

Teaching notes

Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

Ask: Visual vectors and text embeddings share the decoder’s input width.

Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Then place the user’s text

Progressive visual prefix: visual vectors before user prompt and current assistant positionV1V2V3USERVisual vectorsText embeddings

Visual vectors and text embeddings share the decoder’s input width.

Original teaching example / diagram68 / 155

Teaching notes

Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

Ask: Visual vectors and text embeddings share the decoder’s input width.

Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Finish the question; start the assistant’s answer

Progressive visual prefix: visual vectors before user prompt and current assistant positionV1V2V3USERWhatanimalisshown?ASSTVisual vectorsText embeddings

Visual vectors and text embeddings share the decoder’s input width.

Original teaching example / diagram69 / 155

Teaching notes

Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

Ask: Visual vectors and text embeddings share the decoder’s input width.

Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

What may the current text position read?

Course mask grammar: queries in rows, keys in columns, allowed coloured cells and blocked grey cellsCurrent text queryASSISTANT positionV1✓V2✓V3✓What✓animal✓?✓All six are earlier context.No future text position is available.

The assistant’s state can use both the picture and the question.

Original teaching example / diagram70 / 155

Teaching notes

Read one query row first. The current position may also attend to itself; the simplified row here emphasizes earlier context.

Ask: The assistant’s state can use both the picture and the question.

Answer: Read one query row first. The current position may also attend to itself; the simplified row here emphasizes earlier context.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

In this decoder, the causal mask starts at V1

Queries are rows and keys are columns. In this standard causal-prefix decoder, the mask includes visual positions: V1 reads itself; V2 reads V1 and itself. Text reads all preceding image rows and causal text context. ViT has already contextualized each image row. Colored cells are allowed; grey cells are blocked.Key / value columns →Query ↓V1V2V3T1T2T3AV1V2V3T1T2T3AVISUAL POSITIONSV1 reads V1.V2 reads V1, V2.V3 reads V1, V2, V3.TEXT POSITIONST1 reads V1–V3 + T1.Later text reads images+ earlier text + itself.T1: What · T2: animalT3: ? · A: ASSTViT has already mixedthe whole imageinto each visual row.Known image + prompt rows run together in one masked pass.

Colored: allowed · grey: blocked. The triangle includes the visual positions.

ViT · Dosovitskiy et al., 2020 · LLaVA · Liu et al., 2023 · Transformer · Vaswani et al., 201771 / 155

Teaching notes

V1–V3 are the projected image vectors called z1…zN in the earlier numeric example; three visual positions make the mask readable. Rows are queries and columns are keys. Read the first three rows before the text: V1 can read V1, V2 can read V1 and V2, and V3 can read V1 through V3. No visual decoder query reads later visual positions or later text, but every visual input already incorporates image-wide context from the ViT encoder. Each text row can read all preceding visual rows, preceding text rows and itself. All supplied image/prompt rows are processed together with this mask during prefill, not generated one by one. New answer tokens are generated autoregressively afterwards. This is the ordinary LLaVA/LLaMA causal-mask example; some architectures instead use full attention within the visual prefix. The mask controls attention edges inside the decoder, not the ViT encoder.

Ask: Colored: allowed · grey: blocked. The triangle includes the visual positions.

Answer: V1–V3 are the projected image vectors called z1…zN in the earlier numeric example; three visual positions make the mask readable. Rows are queries and columns are keys. Read the first three rows before the text: V1 can read V1, V2 can read V1 and V2, and V3 can read V1 through V3. No visual decoder query reads later visual positions or later text, but every visual input already incorporates image-wide context from the ViT encoder. Each text row can read all preceding visual rows, preceding text rows and itself. All supplied image/prompt rows are processed together with this mask during prefill, not generated one by one. New answer tokens are generated autoregressively afterwards. This is the ordinary LLaVA/LLaMA causal-mask example; some architectures instead use full attention within the visual prefix. The mask controls attention edges inside the decoder, not the ViT encoder.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Back to the big picture: the image is earlier context

Complete prefix architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderProject each visual vectorV1V2…VNSame rows, new width.All N patch rows; middle rows omitted.LANGUAGE: GROWING TEXTQuestion + answer so farV1V2…VNText embeddingsCausal self-attention + residualMLP + residualVocabulary head → next token

We have built all three jobs: visual rows, matching width, and a prefix.

Original teaching example / diagram72 / 155

Teaching notes

No extra cross-attention layer is required by this access pattern.

Ask: We have built all three jobs: visual rows, matching width, and a prefix.

Answer: No extra cross-attention layer is required by this access pattern.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

This is the basic LLaVA-style connection

Complete prefix architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderProject each visual vectorV1V2…VNSame rows, new width.All N patch rows; middle rows omitted.LANGUAGE: GROWING TEXTQuestion + answer so farV1V2…VNText embeddingsCausal self-attention + residualMLP + residualVocabulary head → next token

Projected patch vectors occupy earlier positions; they do not become word IDs.

LLaVA · Liu et al., 202373 / 155

Teaching notes

A simple row-wise linear or MLP projector keeps the patch-row count. The few boxes are a drawn subset, not token selection. LLaVA variants differ in details. Later compression methods deliberately reduce the visual sequence.

Ask: Projected patch vectors occupy earlier positions; they do not become word IDs.

Answer: A simple row-wise linear or MLP projector keeps the patch-row count. The few boxes are a drawn subset, not token selection. LLaVA variants differ in details. Later compression methods deliberately reduce the visual sequence.

PAPER · LLaVA · NeurIPS 2023

Visual Instruction Tuning

Haotian Liu · Chunyuan Li · Qingyang Wu · Yong Jae Lee

Read the paper ↗
Original LLaVA architecture: image Xv enters a vision encoder, features Zv pass through projection W to visual embeddings Hv, and these join instruction embeddings Hq in a language model producing response Xa.
Original Figure 1 · Image features and instruction embeddings enter the same language model.

1 · Vision encoder

CLIP ViT-L/14 supplies image features Zᵥ.

2 · Projection W

The original paper uses a linear map to match the LM embedding width.

3 · Language model

Vicuna reads the visual vectors and text, then predicts the response.

LLaVA · Liu et al., 202374 / 155

Teaching notes

Figure 1 is cropped from page 4 of arXiv 2304.08485v2; the architecture itself is unaltered. Read the diagram bottom to top. The paper calls unprojected features Zv and projected visual embeddings Hv; our lecture uses Hv and Z respectively, so map by role rather than letter. The original connector is linear; later variants can use an MLP. The few drawn tokens are schematic. The figure abstracts tokenization, positions and repeated decoder blocks. Visual instruction tuning also contributes a data/training recipe; this slide focuses on the architecture.

Ask: Point to the modules that implement the connection we just derived.

Answer: Vision encoder: CLIP ViT-L/14 supplies image features Zᵥ. Projection W: The original paper uses a linear map to match the LM embedding width. Language model: Vicuna reads the visual vectors and text, then predicts the response.

02 · HOW CAN VISION ENTER A LANGUAGE MODEL?

Idea 1, end to end in one function

The same Newfoundland image, supplied to the ViT

image
One image

question = “What animal is shown?”

answer_so_far = “It is a” (empty at the start)

One next-token step · high-level pseudocodeShapes · batch omitted
def next_token(image, question, answer_so_far):
patches = vit(image).patch_statesN × dV
visual = projector(patches)N × dLM
ids = tokenize(question, answer_so_far)T token IDs
text = embed(ids)T × dLM
inputs = concat([visual, text], axis="tokens")(N + T) × dLM
hidden = causal_decoder(inputs)(N + T) × dLM
scores = lm_head(hidden[-1])Vocabulary scores
return argmax(scores)Next text token ID

causal_decoder stacks causal self-attention + MLP blocks over all input positions.

Display the returned token with the tokenizer; append its ID to the answer and predict again.

LLaVA · Liu et al., 2023 · Oxford-IIIT Pet · CC BY-SA 4.075 / 155

Teaching notes

High-level pseudocode for one next-token prediction, not executable library calls. One image is shown and the batch dimension is omitted. Image preprocessing is folded into vit; patch_states selects all N contextual patch rows and excludes CLS in this design. ViT uses full image attention. The shared row-wise projector changes dV to dLM while preserving N. tokenize(question, answer_so_far) is a conceptual helper for constructing and tokenizing the text context in that order, not the tokenizer API for paired sequences. In a real model, use its chat template and processor; keep generated answer IDs rather than repeatedly retokenizing them. T counts all supplied text positions, including already-generated answer tokens. The answer portion is empty at the first prediction. concat stacks visual rows before text rows along the token axis, not the feature axis. The decoder consumes embeddings directly, with positional handling and a causal mask from the first visual position. causal_decoder means the full repeated Transformer block stack, including MLPs, normalization and residuals; it is not only one attention operation. hidden[-1] is the last supplied position and predicts the next, as-yet-unknown answer token. The language head produces a score for every vocabulary entry. argmax shows greedy selection; softmax is unnecessary for choosing the largest score. The result is a text token ID, not a visual token. Decode it with the tokenizer to display text. This deliberately unoptimized function summarizes the dependency graph. During actual generation, encode the fixed image once and reuse the decoder KV cache rather than recomputing the whole function on every token. Question and answer text are illustrative; no measured answer or embedding is claimed.

Ask: Which line combines image and text, and which state predicts the next token?

Answer: concat joins N projected visual rows and T text embeddings along the token axis. The causal decoder processes their combined sequence; the language head reads hidden[-1], the last supplied position, to score the next text token.

03 · COULD VISION REMAIN SEPARATE?

COULD VISION REMAIN SEPARATE?

03

Return to Idea 2: how can language read separate visual memory?

Original teaching example / diagram76 / 155

Teaching notes

Return to Idea 2: how can language read separate visual memory?

Ask:

Answer: Return to Idea 2: how can language read separate visual memory?

03 · COULD VISION REMAIN SEPARATE?

Idea 2 in detail · read a separate visual memory

Complete cross architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderh1h2h3hN…Memory = N encoder feature rowsBatch shape: B × N × dVB: batch size · N: rows · dV: row widthLANGUAGE: GROWING TEXTQuestion + answer so farText embeddings + positionsCausal self-attention + residualCross-attention + residualK, VMLP + residualVocabulary head → next token

Causal self-attention reads text; cross-attention reads image memory.

Transformer · Vaswani et al., 201777 / 155

Teaching notes

Trace the two inputs separately. The purple path embeds the question and tokens generated so far; the answer portion is empty initially. The vision encoder exposes h1, h2, …, hN, a sequence of feature rows. B × N × dV means batch size × retained visual rows per example × features per row; the drawing shows one image. Text states supply Q, image states are projected to K and V. The attention result updates the language states before the vocabulary head. N varies with image resolution, patch size, CLS handling and any compression; 197 is only a conditional example. This schematic shows the standard sequence-to-sequence mechanism, not Flamingo’s exact layer ordering or gates.

Ask: Causal self-attention reads text; cross-attention reads image memory.

Answer: Trace the two inputs separately. The purple path embeds the question and tokens generated so far; the answer portion is empty initially. The vision encoder exposes h1, h2, …, hN, a sequence of feature rows. B × N × dV means batch size × retained visual rows per example × features per row; the drawing shows one image. Text states supply Q, image states are projected to K and V. The attention result updates the language states before the vocabulary head. N varies with image resolution, patch size, CLS handling and any compression; 197 is only a conditional example. This schematic shows the standard sequence-to-sequence mechanism, not Flamingo’s exact layer ordering or gates.

03 · COULD VISION REMAIN SEPARATE?

We already did this with English-to-Hindi translation

Complete translation architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: ENGLISHRaghav goes to school in DelhiEnglish encoderh1h2h3h4h5Keep the source states.Schematic; vector boxes show a subset.LANGUAGE: GROWING TEXTHindi prefix generated so farText embeddings + positionsCausal self-attention + residualCross-attention + residualK, VMLP + residualVocabulary head → next token

English source states stayed available while the Hindi prefix grew.

Original teaching example / diagram78 / 155

Teaching notes

Use the complete source-left, decoder-right architecture from the final diagrams of the earlier lecture. The current target state first uses causal target self-attention and then reads source states. Replace only the source modality on the next slide.

Ask: English source states stayed available while the Hindi prefix grew.

Answer: Use the complete source-left, decoder-right architecture from the final diagrams of the earlier lecture. The current target state first uses causal target self-attention and then reads source states. Replace only the source modality on the next slide.

03 · COULD VISION REMAIN SEPARATE?

Keep the decoder idea; change English to an image

Complete cross architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderh1h2h3hN…Memory = N encoder feature rowsBatch shape: B × N × dVB: batch size · N: rows · dV: row widthLANGUAGE: GROWING TEXTQuestion + answer so farText embeddings + positionsCausal self-attention + residualCross-attention + residualK, VMLP + residualVocabulary head → next token

The source changes. The decoder still reads a separate memory.

Original teaching example / diagram79 / 155

Teaching notes

The text sequence remains separate from visual memory. We have not yet computed a visual message.

Ask: The source changes. The decoder still reads a separate memory.

Answer: The text sequence remains separate from visual memory. We have not yet computed a visual message.

03 · COULD VISION REMAIN SEPARATE?

What does the visual lookup need?

PartRole in the familiar attention operation
Query: from the current language stateWhat information would help this position?
Keys: from the visual statesCompute a match score for each memory row.
Values: from the visual statesSupply the features that the weights will mix.

Q decides how to read; K is compared; V supplies the returned content.

Original teaching example / diagram80 / 155

Teaching notes

These phrases explain data flow, not interpretable concepts encoded by every learned coordinate. A query is a vector derived from the current text computation, not the entire question string. We now derive these three pieces on a zoomed-in view.

Ask: If we swap the image but keep this incoming language state fixed, what changes?

Answer: The visual keys and values change. The incoming query is fixed for this isolated lookup; in the full model later language states can also change.

03 · COULD VISION REMAIN SEPARATE?

The language state asks the question

Reuse source encoder states above and updated target states below; cross-attention reads a different source modalityVISUAL SOURCE STATESh1h2h3h4h5h6LANGUAGE STATE · AFTER CAUSAL SELF-ATTENTIONCurrent decoder stateNew query Q
Q=HLWQQ=H_LW_Q

Queries come from the language state.

Original teaching example / diagram81 / 155

Teaching notes

H_L is the state entering this cross-attention sublayer, after causal self-attention and any normalization. Cross-attention has its own projection weights.

Ask: Queries come from the language state.

Answer: H_L is the state entering this cross-attention sublayer, after causal self-attention and any normalization. Cross-attention has its own projection weights.

03 · COULD VISION REMAIN SEPARATE?

The image supplies the keys

Reuse source encoder states above and updated target states below; cross-attention reads a different source modalityVISUAL SOURCE STATESh1h2h3h4h5h6Visual keys KLANGUAGE STATE · AFTER CAUSAL SELF-ATTENTIONCurrent decoder stateNew query Q
K=HVWKK=H_VW_K

Keys describe what the visual memory can match.

Original teaching example / diagram82 / 155

Teaching notes

Q and K share the per-head width d_k; their sequence lengths need not match.

Ask: Keys describe what the visual memory can match.

Answer: Q and K share the per-head width d_k; their sequence lengths need not match.

03 · COULD VISION REMAIN SEPARATE?

The image also supplies the values

Reuse source encoder states above and updated target states below; cross-attention reads a different source modalityVISUAL SOURCE STATESh1h2h3h4h5h6Visual keys KVisual values VLANGUAGE STATE · AFTER CAUSAL SELF-ATTENTIONCurrent decoder stateNew query Q
V=HVWVV=H_VW_V

Values provide the content that attention will mix.

Original teaching example / diagram83 / 155

Teaching notes

W_K and W_V are separate learned matrices. Per-input K and V are computed activations.

Ask: Values provide the content that attention will mix.

Answer: W_K and W_V are separate learned matrices. Per-input K and V are computed activations.

03 · COULD VISION REMAIN SEPARATE?

Compare a query with every visual key

s=qKTdks=\frac{qK^{\mathsf T}}{\sqrt{d_k}}
Language query×Visual keys→Match scores

The result is one score per visual state.

Original teaching example / diagram84 / 155

Teaching notes

Scaling controls dot-product magnitude before normalization. One attention head shown.

Ask: The result is one score per visual state.

Answer: Scaling controls dot-product magnitude before normalization. One attention head shown.

03 · COULD VISION REMAIN SEPARATE?

Turn the scores into mixing weights

α=softmax⁡(s)\alpha=\operatorname{softmax}(s)
Scores→Nonnegative weights summing to one

Softmax normalizes across the visual memory positions.

Original teaching example / diagram85 / 155

Teaching notes

These are attention weights, not probabilities of answer words.

Ask: Softmax normalizes across the visual memory positions.

Answer: These are attention weights, not probabilities of answer words.

03 · COULD VISION REMAIN SEPARATE?

Read a visual message

c=αV=∑iαivic=\alpha V=\sum_i\alpha_i v_i
Visual values→Weighted mixture

One query produces one context vector per head.

Original teaching example / diagram86 / 155

Teaching notes

Heads are concatenated and projected in multi-head attention; the result returns to the language computation.

Ask: One query produces one context vector per head.

Answer: Heads are concatenated and projected in multi-head attention; the result returns to the language computation.

03 · COULD VISION REMAIN SEPARATE?

One tiny lookup: from a language query to a visual message

q=[1,0],dk=2q=[1,0],\quad d_k=2
Teaching labelKeyValue
face[0.9, 0.1][1, 0]
fur[0.6, 0.4][0, 2]
background[0.1, 0.9][1, 1]

Toy numbers and labels; no measured attention map is being shown.

Original teaching example / diagram87 / 155

Teaching notes

Return to the visual-memory architecture. We are inside its teal cross-attention box, computing one message for one language position. The three row labels and all numbers are hand-set, not measured regions from the dog image. A large attention weight does not by itself prove grounding.

Ask: Toy numbers and labels; no measured attention map is being shown.

Answer: Return to the visual-memory architecture. We are inside its teal cross-attention box, computing one message for one language position. The three row labels and all numbers are hand-set, not measured regions from the dog image. A large attention weight does not by itself prove grounding.

03 · COULD VISION REMAIN SEPARATE?

Query × face key

q⋅k1=1(0.9)+0(0.1)=0.9q\cdot k_1=1(0.9)+0(0.1)=0.9

The same language query is compared with another visual key.

Original teaching example / diagram88 / 155

Teaching notes

Unscaled toy dot product. Earlier results remain on the board.

Ask: The same language query is compared with another visual key.

Answer: Unscaled toy dot product. Earlier results remain on the board.

03 · COULD VISION REMAIN SEPARATE?

Query × fur key

q⋅k2=1(0.6)+0(0.4)=0.6q\cdot k_2=1(0.6)+0(0.4)=0.6

The same language query is compared with another visual key.

Original teaching example / diagram89 / 155

Teaching notes

Unscaled toy dot product. Earlier results remain on the board.

Ask: The same language query is compared with another visual key.

Answer: Unscaled toy dot product. Earlier results remain on the board.

03 · COULD VISION REMAIN SEPARATE?

Query × background key

q⋅k3=1(0.1)+0(0.9)=0.1q\cdot k_3=1(0.1)+0(0.9)=0.1

The same language query is compared with another visual key.

Original teaching example / diagram90 / 155

Teaching notes

Unscaled toy dot product. Earlier results remain on the board.

Ask: The same language query is compared with another visual key.

Answer: Unscaled toy dot product. Earlier results remain on the board.

03 · COULD VISION REMAIN SEPARATE?

Scale, then normalize this row

s=[0.9,0.6,0.1]/2≈[0.636,0.424,0.071]s=[0.9,0.6,0.1]/\sqrt2\approx[0.636,0.424,0.071]
α=softmax⁡(s)≈[0.421,0.340,0.239]\alpha=\operatorname{softmax}(s)\approx[0.421,0.340,0.239]

The weights were calculated from Q and K.

Original teaching example / diagram91 / 155

Teaching notes

Full precision: 0.4207287, 0.34030973, 0.23896158. They sum to one.

Ask: The weights were calculated from Q and K.

Answer: Full precision: 0.4207287, 0.34030973, 0.23896158. They sum to one.

03 · COULD VISION REMAIN SEPARATE?

Use those weights to mix the values

c=0.421[1,0]+0.340[0,2]+0.239[1,1]c=0.421[1,0]+0.340[0,2]+0.239[1,1]
c≈[0.660,0.920]c\approx[0.660,0.920]

The result is a vector, not an output word.

Original teaching example / diagram92 / 155

Teaching notes

Rounded display; the executable arithmetic check uses full precision. Values need not equal keys.

Ask: The result is a vector, not an output word.

Answer: Rounded display; the executable arithmetic check uses full precision. Values need not equal keys.

03 · COULD VISION REMAIN SEPARATE?

The message updates the language state; it is not a word

The projected cross-attention message returns to the current language stateCurrent language stateVisual message cResidual +Updated language state
HL′=HL+CrossAttn⁡(HL,HV)H_L^{\prime}=H_L+\operatorname{CrossAttn}(H_L,H_V)

The projected attention output updates the language state.

Original teaching example / diagram93 / 155

Teaching notes

Schematic residual update. Normalization, multi-head output projection and model-specific gates are omitted from the drawing; they do not turn the context vector directly into a word.

Ask: The projected attention output updates the language state.

Answer: Schematic residual update. Normalization, multi-head output projection and model-specific gates are omitted from the drawing; they do not turn the context vector directly into a word.

03 · COULD VISION REMAIN SEPARATE?

Check: has cross-attention already answered “dog”?

Lookup result

c≈[0.660,0.920]c\approx[0.660,0.920]

What happens next?

Does c name a word?

Or does it update the decoder’s current state?

Predict before advancing: where does a vocabulary token actually appear?

Original teaching example / diagram94 / 155

Teaching notes

The numerical vector is not a word ID. Ask students to point to the vocabulary head in the whole architecture.

Ask: Does the cross-attention output directly name the answer token?

Answer: No. It updates a language hidden state. The vocabulary head later scores output token IDs.

03 · COULD VISION REMAIN SEPARATE?

The vocabulary head still produces the next-token scores

Complete cross architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderh1h2h3hN…Memory = N encoder feature rowsBatch shape: B × N × dVB: batch size · N: rows · dV: row widthLANGUAGE: GROWING TEXTQuestion + answer so farText embeddings + positionsCausal self-attention + residualCross-attention + residualK, VMLP + residualVocabulary head → next token

Visual lookup changes the language state; the language head chooses words.

Original teaching example / diagram95 / 155

Teaching notes

Trace memory → cross-attention → updated decoder computation → vocabulary head. Cross-attention returns features, not an English answer. Model-specific blocks can repeat this lookup in depth.

Ask: Visual lookup changes the language state; the language head chooses words.

Answer: Trace memory → cross-attention → updated decoder computation → vocabulary head. Cross-attention returns features, not an English answer. Model-specific blocks can repeat this lookup in depth.

03 · COULD VISION REMAIN SEPARATE?

Two ways for language to access vision

Same image and prompt: visual context in sequence versus visual context in separate memorySame image. Same question: What animal is shown?A · VISUAL PREFIXV1V2V3questionASSTCausal self-attentionB · SEPARATE MEMORYVisual K, VText query QCross-attention

In-sequence visual context or a separate visual memory.

Original teaching example / diagram96 / 155

Teaching notes

Same image and prompt. Neither access pattern is universally better; training, model size, compute and task matter.

Ask: In-sequence visual context or a separate visual memory.

Answer: Same image and prompt. Neither access pattern is universally better; training, model size, compute and task matter.

03 · COULD VISION REMAIN SEPARATE?

Idea 2, end to end in one function

The same Newfoundland image, supplied to the vision encoder

image
One image

question = “What animal is shown?”

answer_so_far = “It is a” (empty at the start)

One next-token step · high-level pseudocodeShapes · batch omitted
def next_token(image, question, answer_so_far):
memory = vision_encoder(image).featuresN × dV
ids = tokenize(question, answer_so_far)T token IDs
text = embed(ids)T × dLM
for block in decoder_blocks:
text = block.causal_self_attn(text)T × dLM
text = block.cross_attn(text, memory)T × dLM
text = block.mlp(text)T × dLM
return argmax(lm_head(text[-1]))Next text token ID

cross_attn: Q from text; K and V from image memory. The two sequences stay separate.

Sublayer helpers include residuals and norms. Append the returned token ID to the answer; repeat.

Transformer · Vaswani et al., 2017 · Oxford-IIIT Pet · CC BY-SA 4.097 / 155

Teaching notes

A conceptual encoder-decoder computation for one image and one next-token prediction, batch omitted. vision_encoder includes image preprocessing and produces all N feature rows. For the ViT teaching example these are contextual patch states, with full image attention. The tokenizer helper formats question and answer prefix, as on the Idea 1 slide; T includes all supplied text positions. Each causal_self_attn helper returns updated text states using causal self-attention, residual addition and normalization; cross_attn and mlp also include their residual and normalization steps. Positional handling and the final stack normalization are abstracted. The cross-attention helper has separate learned projections: Q comes from text and K/V from image memory. Per-head Q and K widths must match, but dV and dLM need not match; the internal projections and output projection handle this. No visual/text concatenation occurs. The decoder text stream keeps T rows, and the final text position feeds the vocabulary head. Image memory is constant within this prediction. This is the familiar standard sequence-to-sequence ordering, not an exact implementation of Flamingo. Flamingo adds gated visual cross-attention at selected points among pretrained language blocks and uses a resampled visual memory. The later Flamingo paper slide maps those differences explicitly. For efficient generation, encode the image once and cache reusable attention keys and values. argmax is greedy decoding; display via the tokenizer and append the ID, stopping at EOS or the generation limit. All APIs and the answer prefix are illustrative.

Ask: Where do Q, K and V come from, and how many text states leave each block?

Answer: Cross-attention gets Q from text states and K/V from image memory. It updates T text states; the N image rows remain separate. The last text state supplies the next-token scores.

04 · DO WE NEED EVERY VISUAL TOKEN?

DO WE NEED EVERY VISUAL TOKEN?

04

Return to Idea 3: how can we make a shorter visual summary?

Original teaching example / diagram98 / 155

Teaching notes

Return to Idea 3: how can we make a shorter visual summary?

Ask:

Answer: Return to Idea 3: how can we make a shorter visual summary?

04 · DO WE NEED EVERY VISUAL TOKEN?

Idea 3 in detail · make the visual context shorter

Complete compressed architecture using the source-left decoder-right grammar of the earlier attention lectureSOURCE: IMAGEVision encoderSummarize + projectv1v2v3v4Fewer rows, then match width.Schematic; vector boxes show a subset.LANGUAGE: GROWING TEXTWhat animal is shown? + prefixv1v2v3v4Text embeddingsCausal self-attention + residualMLP + residualVocabulary head → next token

Replace the row-wise projector with a learned summary followed by projection.

Original teaching example / diagram99 / 155

Teaching notes

Compare with the first complete prefix architecture. Only the left-hand connector changes. The decoder still consumes visual vectors followed by text. The same compression idea can instead provide the separate memory used by cross-attention.

Ask: Replace the row-wise projector with a learned summary followed by projection.

Answer: Compare with the first complete prefix architecture. Only the left-hand connector changes. The decoder still consumes visual vectors followed by text. The same compression idea can instead provide the separate memory used by cross-attention.

04 · DO WE NEED EVERY VISUAL TOKEN?

How long is the combined context?

To-scale sequence lengths: 640 positions compared with 96576 visual + 64 text = 640 positions576 visual states64 text

576 visual states + 64 text positions = 640 positions.

Original teaching example / diagram100 / 155

Teaching notes

Illustrative 24 by 24 patch grid, with no CLS state included. Token counts depend on image resolution and preprocessing.

Ask: 576 visual states + 64 text positions = 640 positions.

Answer: Illustrative 24 by 24 patch grid, with no CLS state included. Token counts depend on image resolution and preprocessing.

04 · DO WE NEED EVERY VISUAL TOKEN?

What if the image became 32 context vectors?

To-scale sequence lengths: 640 positions compared with 96576 visual + 64 text = 640 positions576 visual states64 text32 visual + 64 text = 96 positions3264Same scale in both rows.

32 visual states + 64 text positions = 96 positions.

Original teaching example / diagram101 / 155

Teaching notes

Both bars use the same scale. Dense attention score counts are 640² versus 96², about 44.4 times fewer pairs, not 44.4 times total speedup.

Ask: 32 visual states + 64 text positions = 96 positions.

Answer: Both bars use the same scale. Dense attention score counts are 640² versus 96², about 44.4 times fewer pairs, not 44.4 times total speedup.

04 · DO WE NEED EVERY VISUAL TOKEN?

Now build the summary: start with the image states

32 learned queries attend to 576 image states and produce 32 output vectors; subsets explicitly labelled576 image-dependent visual states · showing 6 of 576h1h2h3h4h5h6

How can we read a fixed-size summary of a variable visual memory?

Original teaching example / diagram102 / 155

Teaching notes

Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

Ask: How can we read a fixed-size summary of a variable visual memory?

Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

04 · DO WE NEED EVERY VISUAL TOKEN?

Use 32 learned query vectors as 32 readers

32 learned queries attend to 576 image states and produce 32 output vectors; subsets explicitly labelled576 image-dependent visual states · showing 6 of 576h1h2h3h4h5h632 learned queriesCross-attentionImage supplies K, V

How can we read a fixed-size summary of a variable visual memory?

Original teaching example / diagram103 / 155

Teaching notes

Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

Ask: How can we read a fixed-size summary of a variable visual memory?

Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

04 · DO WE NEED EVERY VISUAL TOKEN?

Each reader returns one image-dependent summary vector

32 learned queries attend to 576 image states and produce 32 output vectors; subsets explicitly labelled576 image-dependent visual states · showing 6 of 576h1h2h3h4h5h632 learned queriesCross-attentionImage supplies K, Vz1z2z3z4Showing 4 of 32

The query count fixes the number of output vectors.

Original teaching example / diagram104 / 155

Teaching notes

Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

Ask: The query count fixes the number of output vectors.

Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.

04 · DO WE NEED EVERY VISUAL TOKEN?

Which detail survives the bottleneck?

Fine text in the invoice

32 outputs must carry the evidence needed by the question.

A shorter sequence can lose small text or spatial detail.

Compression is a separate choice from prefix versus cross-attention access.

Original teaching example / diagram105 / 155

Teaching notes

A resampler can feed either access pattern. Do not equate fewer visual tokens with universally better performance.

Ask: Compression is a separate choice from prefix versus cross-attention access.

Answer: A resampler can feed either access pattern. Do not equate fewer visual tokens with universally better performance.

04 · DO WE NEED EVERY VISUAL TOKEN?

Attach names after the design choices

ExampleVisual statesConnectorAccess
LLaVA-styleViT patchesLinear / MLPVisual prefix
FlamingoVision feature gridPerceiver ResamplerGated cross-attention
BLIP-2ViT patchesQ-Former + projectionQuery-output prefix*

*Decoder-only LLM variant. Other language backbones use different interfaces.

Flamingo and BLIP-2 use learned-query mechanisms in different systems.

LLaVA · Liu et al., 2023 · Flamingo · Alayrac et al., 2022 · BLIP-2 · Li et al., 2023106 / 155

Teaching notes

This is a design comparison, not a ranking. BLIP-2 supports different language backbones; decoder-only examples use visual prefix conditioning. Training objectives and exact modules differ.

Ask: Flamingo and BLIP-2 use learned-query mechanisms in different systems.

Answer: This is a design comparison, not a ranking. BLIP-2 supports different language backbones; decoder-only examples use visual prefix conditioning. Training objectives and exact modules differ.

PAPER · Flamingo · NeurIPS 2022

Flamingo: a Visual Language Model for Few-Shot Learning

Jean-Baptiste Alayrac · Jeff Donahue · Pauline Luc · et al.

Read the paper ↗
Original Flamingo overview: frozen vision encoders feed Perceiver Resamplers whose outputs supply gated cross-attention blocks interleaved with frozen language-model blocks. Interleaved images and text provide context for generated text.
Original Figure 3 · Image/video features condition a separate language stream.

1 · Idea 3 · compress visual features

Frozen NFNet features → Perceiver Resampler → fixed-size visual memory.

2 · Idea 2 · read separate memory

Gated cross-attention inserts visual information between LM blocks.

3 · Read the freeze symbols

Vision encoder and LM stay frozen; the added bridge modules are trained.

Flamingo · Alayrac et al., 2022107 / 155

Teaching notes

Original Figure 3 from arXiv 2204.14198v1. Trace both visual paths into the language stack. The original vision backbone is an NFNet-F6 convolutional network, not the generic ViT used for our teaching diagram. The Perceiver Resampler compresses the visual features; gated cross-attention keeps them separate from the text stream. The diagram also shows interleaved image/text context for few-shot prediction. Snowflakes indicate frozen pretrained modules. The paper has an image-conditioning mask and other details beyond our generic cross-attention illustration. Its figure colors follow the paper, not the lecture palette.

Ask: Point to the modules that implement the connection we just derived.

Answer: Idea 3 · compress visual features: Frozen NFNet features → Perceiver Resampler → fixed-size visual memory. Idea 2 · read separate memory: Gated cross-attention inserts visual information between LM blocks. Read the freeze symbols: Vision encoder and LM stay frozen; the added bridge modules are trained.

PAPER · BLIP-2 · ICML 2023

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Junnan Li · Dongxu Li · Silvio Savarese · Steven Hoi

Read the paper ↗
Original BLIP-2 Figure 3: frozen image encoder feeds learned queries in a Q-Former; an FC projection connects their outputs to either an OPT-style decoder or a FlanT5-style encoder-decoder.
Original Figure 3 · Top: decoder-only LLM. Bottom: encoder–decoder LLM.

1 · Idea 3 · learned-query summaries

32 learned queries produce 32 image-dependent output vectors.

2 · Idea 1 · prefix for OPT (top)

Project the summaries; prepend them to text embeddings.

3 · Bottom: encoder–decoder

Visual vectors + text enter the encoder; its decoder generates the answer.

BLIP-2 · Li et al., 2023108 / 155

Teaching notes

Original Figure 3 from arXiv 2301.12597v3, showing the second generative-pretraining stage. The top OPT route maps to our compressed visual-prefix connection. The lower FlanT5 route conditions the language encoder, so it is not the same decoder-only interface. The fully connected layer changes vector width after Q-Former compression. Query outputs depend on the image, while the learned query parameters are shared. Image encoder and LLM remain frozen in this stage; Q-Former and the connecting projection are trained. The first representation-learning stage and its objectives are described elsewhere in the paper; the simple learned-reader diagram was not a full Q-Former specification.

Ask: Point to the modules that implement the connection we just derived.

Answer: Idea 3 · learned-query summaries: 32 learned queries produce 32 image-dependent output vectors. Idea 1 · prefix for OPT (top): Project the summaries; prepend them to text embeddings. Bottom: encoder–decoder: Visual vectors + text enter the encoder; its decoder generates the answer.

04 · DO WE NEED EVERY VISUAL TOKEN?

The learned readers stay; their image summaries change

Stored model parameters versus activations computed for this image and token contextLEARNED / STOREDCOMPUTED FOR THIS INPUTVision weightsImage: visual states H and projected ZConnector weightsPrompt: language hidden statesLanguage weights + token embeddingsPositions: queries, keys and valuesCross-attention weights, if presentAttention weights + vocabulary logitsLearned queries, if presentSelected tokens appended at inference

Parameters persist across examples; activations depend on the input.

Original teaching example / diagram109 / 155

Teaching notes

A pretrained but frozen weight is still a learned parameter. The frozen/trainable decision will be a separate overlay during training. Learned query embeddings are parameters; actual projected attention queries are activations.

Ask: Parameters persist across examples; activations depend on the input.

Answer: A pretrained but frozen weight is still a learned parameter. The frozen/trainable decision will be a separate overlay during training. Learned query embeddings are parameters; actual projected attention queries are activations.

04 · DO WE NEED EVERY VISUAL TOKEN?

Idea 3, end to end in one function

The same Newfoundland image, supplied to the vision encoder

image
One image

question = “What animal is shown?”

answer_so_far = “It is a” (empty at the start)

One next-token step · high-level pseudocodeShapes · batch omitted
def next_token(image, question, answer_so_far):
patches = vit(image).patch_statesN × dV
summary = summarizer(queries, patches)M × dS; M < N
visual = projector(summary)M × dLM
ids = tokenize(question, answer_so_far)T token IDs
text = embed(ids)T × dLM
inputs = concat([visual, text], axis="tokens")(M + T) × dLM
hidden = causal_decoder(inputs)(M + T) × dLM
return argmax(lm_head(hidden[-1]))Next text token ID

summarizer: M learned queries read all N patch rows through cross-attention; outputs depend on the image.

Shown: BLIP-2 / OPT-style prefix. For a Flamingo-style connection, use the summaries as Idea 2’s memory.

BLIP-2 · Li et al., 2023 · Flamingo · Alayrac et al., 2022 · Oxford-IIIT Pet · CC BY-SA 4.0110 / 155

Teaching notes

This conceptual function demonstrates the compressed-prefix route, corresponding to the top OPT-style path in the BLIP-2 architecture figure. The image enters a ViT and all N contextual patch features enter the summarizer. queries is an M-row learned parameter bank shared across images, not an input text question. The summarizer forms attention queries from this bank and its evolving states; patch features supply image keys and values. Its M output rows depend on the image and are mixtures, not M selected original patches. dS is the summarizer output width. The summarizer can contain repeated self-attention, cross-attention and MLP blocks, with residuals and norms; this one call is not a full specification of Q-Former or Perceiver Resampler. BLIP-2 uses 32 learned queries. M < N expresses the compression example rather than a universal constraint at every possible resolution. The projector changes width dS to dLM, preserving M. Then the same tokenize, embed, concat, causal decoder and language head as Idea 1 produce the next answer token. For the other connection, keep summary as a separate memory and pass it to the Idea 2 cross-attention decoder with the T text embeddings; do not also concatenate. Flamingo exemplifies compression plus separate memory, with NFNet visual features rather than this ViT and a gated cross-attention architecture. Its memory need not be externally projected to dLM; cross-attention has its own K/V and output projections. BLIP-2 also has an encoder-decoder FlanT5 route, shown in the paper figure but not implemented by this decoder-only sketch. The input formatting, positional handling, masking, full decoder stack, final normalization and generation caching conventions are the same conceptual abstractions as Idea 1. The summarizer queries are learned parameters; its summaries are newly computed activations. Freeze/train choices are taught in the training section.

Ask: Which line changes the number of visual rows, and which line changes their width?

Answer: summarizer changes N patch rows into M image-dependent summaries. projector changes each summary’s width to dLM while retaining M. This sketch uses a prefix; the same summaries could instead become separate memory.

05 · FOLLOW ONE ANSWER TOKEN END TO END

FOLLOW ONE ANSWER TOKEN END TO END

05

Now follow the connection we built all the way to a word.

Original teaching example / diagram111 / 155

Teaching notes

Now follow the connection we built all the way to a word.

Ask:

Answer: Now follow the connection we built all the way to a word.

05 · FOLLOW ONE ANSWER TOKEN END TO END

The same Newfoundland image enters the encoder

Canonical VLM: image, encoder, visual states, connector, visual context, prompt, decoder state, LM head and tokenImageVision encoderVisual statesHVN × dVConnectorVisual contextZM × dLQuestion + prefixWhat animal is shown?Language decoderStateLM headTokenlogits → softmax

The encoder computes the visual state matrix.

Oxford-IIIT Pet · CC BY-SA 4.0112 / 155

Teaching notes

Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

Ask: The encoder computes the visual state matrix.

Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

05 · FOLLOW ONE ANSWER TOKEN END TO END

The connector exposes usable visual context

Canonical VLM: image, encoder, visual states, connector, visual context, prompt, decoder state, LM head and tokenImageVision encoderVisual statesHVN × dVConnectorVisual contextZM × dLQuestion + prefixWhat animal is shown?Language decoderStateLM headTokenlogits → softmax

The visual context is available to the language computation.

Oxford-IIIT Pet · CC BY-SA 4.0113 / 155

Teaching notes

Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

Ask: The visual context is available to the language computation.

Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Combine image, question and answer prefix

Canonical VLM: image, encoder, visual states, connector, visual context, prompt, decoder state, LM head and tokenImageVision encoderVisual statesHVN × dVConnectorVisual contextZM × dLQuestion + prefixWhat animal is shown?Language decoderStateLM headTokenlogits → softmax

The current language state will drive the vocabulary head.

Oxford-IIIT Pet · CC BY-SA 4.0114 / 155

Teaching notes

Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

Ask: The current language state will drive the vocabulary head.

Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.

05 · FOLLOW ONE ANSWER TOKEN END TO END

The LM head scores vocabulary entries

ℓt=Wvocabht+b\ell_t=W_{\mathrm{vocab}}h_t+b
Vocabulary entryLogit
dog
ℓdog\ell_{\mathrm{dog}}
cat
ℓcat\ell_{\mathrm{cat}}
horse / …
ℓhorse  /  …\ell_{\mathrm{horse}}\;/\;\ldots

Visual evidence can change hₜ and therefore every vocabulary score.

Original teaching example / diagram115 / 155

Teaching notes

Scores are symbolic; the browser API does not expose reliable per-token logits here. No numerical probabilities are invented.

Ask: Visual evidence can change hₜ and therefore every vocabulary score.

Answer: Scores are symbolic; the browser API does not expose reliable per-token logits here. No numerical probabilities are invented.

05 · FOLLOW ONE ANSWER TOKEN END TO END

From vocabulary scores to a selected token

p(yt∣y<t,I,P)=softmax⁡(ℓt)p(y_t\mid y_{<t},I,P)=\operatorname{softmax}(\ell_t)
Vocabulary distribution→select “dog”

Selection can be greedy or sampled; the output is a vocabulary token.

Original teaching example / diagram116 / 155

Teaching notes

“dog” is an illustrative selection explaining the computation, not a claim that the tiny model chose it. The measured counterfactual follows. Real tokenizers may split displayed words.

Ask: Selection can be greedy or sampled; the output is a vocabulary token.

Answer: “dog” is an illustrative selection explaining the computation, not a claim that the tiny model chose it. The measured counterfactual follows. Real tokenizers may split displayed words.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Start with image + prompt

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

〈empty〉
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram117 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Append the first token

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

A
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram118 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Use the growing prefix

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

Ablack
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram119 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Append the object word

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

Ablackdog
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram120 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Finish the sentence

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

Ablackdog.
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram121 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Stop at the end token

Fixed Newfoundland image

Fixed prompt

Describe the animal.

Generated prefix

Ablackdog.EOS
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Illustrative trace: the image and prompt remain fixed.

Original teaching example / diagram122 / 155

Teaching notes

The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

Ask: Illustrative trace: the image and prompt remain fixed.

Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.

05 · FOLLOW ONE ANSWER TOKEN END TO END

Keep the prompt fixed. Change only the image.

What animal is shown?

Dog

Dog

There is a black dog in the foreground.

Cat

Cat

A cat is visible.

Blank

Blank

A squirrel.

p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Actual saved answers from the pinned browser model; probabilities are not exposed.

SmolVLM-256M · model card123 / 155

Teaching notes

Blank means a white image passed through the same image processor, not omission of the image input. Responses test sensitivity to image changes; they do not by themselves establish reliable grounding. Full exact records and metadata are in results/vlm_lab_outputs.json.

Ask: Actual saved answers from the pinned browser model; probabilities are not exposed.

Answer: Blank means a white image passed through the same image processor, not omission of the image input. Responses test sensitivity to image changes; they do not by themselves establish reliable grounding. Full exact records and metadata are in results/vlm_lab_outputs.json.

06 · HOW DO WE TRAIN THE CONNECTION?

HOW DO WE TRAIN THE CONNECTION?

06

Make the visual connection useful by training on reference answers.

Original teaching example / diagram124 / 155

Teaching notes

Make the visual connection useful by training on reference answers.

Ask:

Answer: Make the visual connection useful by training on reference answers.

06 · HOW DO WE TRAIN THE CONNECTION?

The wiring fits. How does it learn to use the image?

Reference targets score the model next-token distribution; the loss updates trainable weightsQuestion: What colour is the dog?Connected VLMnext-token scoresReference answer: Black.The reference supplies training targets.Score the targetsNext-token lossUpdate weights

Use the correct answer tokens to train the visual-to-language connection.

Original teaching example / diagram125 / 155

Teaching notes

Shape matching made the forward computation possible. It did not teach what the image features mean to language. Use image/answer pairs, next-token loss and gradients. The next slides identify which parameters receive updates in example training recipes.

Ask: Use the correct answer tokens to train the visual-to-language connection.

Answer: Shape matching made the forward computation possible. It did not teach what the image features mean to language. Use image/answer pairs, next-token loss and gradients. The next slides identify which parameters receive updates in example training recipes.

06 · HOW DO WE TRAIN THE CONNECTION?

Start from two pretrained components

Frozen weights remain fixed while gradients can reach the trainable connector through their computationsVision encoderFROZENConnector+NEW WEIGHTSLanguage modelFROZENNew connector weights; pretrained vision and language.

The connector is new; vision and language already have learned weights.

Original teaching example / diagram126 / 155

Teaching notes

The snowflake denotes frozen parameters; the flame in later slides denotes parameters being optimized.

Ask: The connector is new; vision and language already have learned weights.

Answer: The snowflake denotes frozen parameters; the flame in later slides denotes parameters being optimized.

06 · HOW DO WE TRAIN THE CONNECTION?

Stage 1 · align visual features with language

Frozen weights remain fixed while gradients can reach the trainable connector through their computationsVision encoderFROZENConnectorTRAINABLELanguage modelFROZENCaption alignment

Example recipe: caption alignment; freezing choices depend on the model.

LLaVA · Liu et al., 2023127 / 155

Teaching notes

Simplified LLaVA-style recipe. Frozen vision encoder and frozen LLM; connector trainable. Exact objectives and freeze choices are architecture-dependent.

Ask: Example recipe: caption alignment; freezing choices depend on the model.

Answer: Simplified LLaVA-style recipe. Frozen vision encoder and frozen LLM; connector trainable. Exact objectives and freeze choices are architecture-dependent.

06 · HOW DO WE TRAIN THE CONNECTION?

Supply the correct previous caption tokens

Dog caption training image

Reference caption

Ablackdog.EOS
L=−∑tlog⁡p(yt∗∣y<t∗,I)\mathcal{L}=-\sum_t\log p(y_t^*\mid y_{<t}^*,I)

Teacher forcing supplies the reference prefix during training.

Original teaching example / diagram128 / 155

Teaching notes

Targets and logits are shifted: each state predicts the next target. With a causal mask all known reference positions can be evaluated in parallel.

Ask: Teacher forcing supplies the reference prefix during training.

Answer: Targets and logits are shifted: each state predicts the next target. With a causal mask all known reference positions can be evaluated in parallel.

06 · HOW DO WE TRAIN THE CONNECTION?

Stage 2 · learn to follow visual instructions

Frozen weights remain fixed while gradients can reach the trainable connector through their computationsVision encoderFROZENConnectorTRAINABLELanguage modelTRAINABLEVisual instruction tuning

Example recipe: image + question + answer; train the LLM or its adapters.

LLaVA · Liu et al., 2023129 / 155

Teaching notes

Example recipe: frozen vision, trainable connector and LLM, or LLM adapters. Some models also unfreeze vision. These choices are model-dependent.

Ask: Example recipe: image + question + answer; train the LLM or its adapters.

Answer: Example recipe: frozen vision, trainable connector and LLM, or LLM adapters. Some models also unfreeze vision. These choices are model-dependent.

06 · HOW DO WE TRAIN THE CONNECTION?

The same image can support different questions

Newfoundland instruction image

USER: What colour is the dog?

ASSISTANT: Black.

L=−∑t∈Alog⁡p(yt∗∣y<t∗,I,P)\mathcal{L}=-\sum_{t\in\mathcal{A}}\log p(y_t^*\mid y_{<t}^*,I,P)

The loss supervises the answer targets selected by the training recipe.

Original teaching example / diagram130 / 155

Teaching notes

A denotes selected assistant answer targets. This answer is an authored training example. Dataset coverage matters: captions alone may not teach counting, OCR or justified rejection.

Ask: The loss supervises the answer targets selected by the training recipe.

Answer: A denotes selected assistant answer targets. This answer is an authored training example. Dataset coverage matters: captions alone may not teach counting, OCR or justified rejection.

06 · HOW DO WE TRAIN THE CONNECTION?

Mark the target at every position

Every input position has a loss-target marker: only Black, period and EOS are targetsVIS1× no targetVIS2× no targetVIS3× no targetUSER× no targetWhat× no targetcolour× no targetis× no targetthe× no targetdog× no target?× no targetASST× no targetBlack✓ target.✓ targetEOS✓ target

Only Black, period and EOS are targets; preceding-position logits predict them.

Original teaching example / diagram131 / 155

Teaching notes

Every displayed token has a target/no-target marker. The loss uses logits at the preceding positions to predict these targets; a prompt-position logit can predict the first answer. Masked context still influences answer gradients.

Ask: Only Black, period and EOS are targets; preceding-position logits predict them.

Answer: Every displayed token has a target/no-target marker. The loss uses logits at the preceding positions to predict these targets; a prompt-position logit can predict the first answer. Masked context still influences answer gradients.

06 · HOW DO WE TRAIN THE CONNECTION?

Two masks answer different questions

Attention mask

V1V2text

↑ available to the current query

Which earlier states may this query read?

Loss mask

× USER✓ Black✓ EOS

Which next-token targets contribute to the loss?

Removing a target loss does not hide a context token.

Original teaching example / diagram132 / 155

Teaching notes

The attention mask controls information flow; the loss mask selects supervised target positions. They are not interchangeable.

Ask: Removing a target loss does not hide a context token.

Answer: The attention mask controls information flow; the loss mask selects supervised target positions. They are not interchangeable.

06 · HOW DO WE TRAIN THE CONNECTION?

A frozen LLM still passes the gradient

Frozen weights remain fixed while gradients can reach the trainable connector through their computationsVision encoderFROZENConnectorTRAINABLELanguage modelFROZENAnswer lossBackward through frozen language operationsNO WEIGHT UPDATE

Frozen ≠ stop-gradient.

Original teaching example / diagram133 / 155

Teaching notes

Differentiate the answer loss through the frozen language operations with respect to their input, so connector weights receive a gradient. The LLM weights are not updated. Wrapping that LLM forward in no_grad would sever the required path.

Ask: Frozen ≠ stop-gradient.

Answer: Differentiate the answer loss through the frozen language operations with respect to their input, so connector weights receive a gradient. The LLM weights are not updated. Wrapping that LLM forward in no_grad would sever the required path.

07 · WHAT HAPPENS AT INFERENCE?

WHAT HAPPENS AT INFERENCE?

07

At inference: prepare the image once, then grow the answer.

Original teaching example / diagram134 / 155

Teaching notes

At inference: prepare the image once, then grow the answer.

Ask:

Answer: At inference: prepare the image once, then grow the answer.

07 · WHAT HAPPENS AT INFERENCE?

The whole generation loop before its individual steps

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2Decode next token3Append selected token4Reuse cached keys / values5Repeat until EOS / limit6

Prepare the fixed context once; repeatedly predict and append a token.

Original teaching example / diagram135 / 155

Teaching notes

This is the same autoregressive loop from the earlier language-model lecture, with one new preparation step for the image. The next frames walk through each step in this complete loop. A cache stores computed keys and values; it does not store the future answer.

Ask: Prepare the fixed context once; repeatedly predict and append a token.

Answer: This is the same autoregressive loop from the earlier language-model lecture, with one new preparation step for the image. The next frames walk through each step in this complete loop. A cache stores computed keys and values; it does not store the future answer.

07 · WHAT HAPPENS AT INFERENCE?

Encode the fixed image once

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1

Only the generated prefix changes at each decoding step.

Original teaching example / diagram136 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Prefill the image and question context

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2

Only the generated prefix changes at each decoding step.

Original teaching example / diagram137 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Decode one new token

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2Decode next token3

Only the generated prefix changes at each decoding step.

Original teaching example / diagram138 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Append the selected token

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2Decode next token3Append selected token4

Only the generated prefix changes at each decoding step.

Original teaching example / diagram139 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Reuse the cached keys and values

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2Decode next token3Append selected token4Reuse cached keys / values5

Only the generated prefix changes at each decoding step.

Original teaching example / diagram140 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Repeat until EOS or the generation limit

Encode once, prefill once, then decode, append and reuse cached computationEncode fixed image1Prefill visual + text context2Decode next token3Append selected token4Reuse cached keys / values5Repeat until EOS / limit6

Only the generated prefix changes at each decoding step.

Original teaching example / diagram141 / 155

Teaching notes

This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

Ask: Only the generated prefix changes at each decoding step.

Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.

07 · WHAT HAPPENS AT INFERENCE?

Same conditional model, different available prefix

Training

A*black*dog*

Reference prefix supplied.
Target loss and gradients.

Inference

Ablack?

Generated prefix grows.
No reference answer or backward pass.

Known targets allow parallel teacher-forced positions; generation is sequential.

Original teaching example / diagram142 / 155

Teaching notes

The asterisk marks reference tokens. At inference each selected output becomes part of the next step’s input.

Ask: Known targets allow parallel teacher-forced positions; generation is sequential.

Answer: The asterisk marks reference tokens. At inference each selected output becomes part of the next step’s input.

07 · WHAT HAPPENS AT INFERENCE?

Reuse the image; keep the conversation context

Black Newfoundland dog outdoors

User: What animal is this?
Assistant: A dog.

User: What colour is its fur?
Assistant: Black.

Illustrative conversation, not a captured model transcript.

Visual features can persist. Decoder cache reuse also depends on the exact prefix.

Oxford-IIIT Pet · CC BY-SA 4.0143 / 155

Teaching notes

69–71 min. Reprocessing the full conversation is valid but less efficient. Image feature caching and decoder KV caching are distinct. The included classroom lab evaluates independent questions; it does not claim multi-turn KV reuse.

Ask: Visual features can persist. Decoder cache reuse also depends on the exact prefix.

Answer: 69–71 min. Reprocessing the full conversation is valid but less efficient. Image feature caching and decoder KV caching are distinct. The included classroom lab evaluates independent questions; it does not claim multi-turn KV reuse.

08 · CAN WE TRUST THE ANSWER?

CAN WE TRUST THE ANSWER?

08

Check evidence use, not just fluent language.

Original teaching example / diagram144 / 155

Teaching notes

Check evidence use, not just fluent language.

Ask:

Answer: Check evidence use, not just fluent language.

08 · CAN WE TRUST THE ANSWER?

A fluent sentence can have no visual support

Coffee photograph without a bicycle

What color is the bicycle?

Reference: No bicycle is visible.

Recorded baseline: “The bicycle is red.”

Fluent ≠ grounded.

Original teaching example / diagram145 / 155

Teaching notes

The displayed raw answer was measured in the earlier six-case browser run and is preserved in output/verification/live-model-results.json. The rebuilt lab also records this preset with the fixed decoding settings.

Ask: Fluent ≠ grounded.

Answer: The displayed raw answer was measured in the earlier six-case browser run and is preserved in output/verification/live-model-results.json. The rebuilt lab also records this preset with the fixed decoding settings.

08 · CAN WE TRUST THE ANSWER?

Keep a small diagnostic suite

ImageFixed questionReference
DogWhat animal is this?Dog
DeskIs the laptop plugged in?Yes, visible cable connected
ShapesWhat is immediately left of the mug?Red ball
InvoiceWhat is the invoice number?ES667-042
ChartWhich month has the highest value?March
CoffeeWhat color is the bicycle?No bicycle visible
Run / record the diagnostic suite ↗Offline fallback: recorded model outputs

Preserve the raw answer, reveal the reference, then judge support.

Original teaching example / diagram146 / 155

Teaching notes

Dog, desk, spatial, invoice, chart and false-premise questions span distinct failure modes. Mark supported / partial / unsupported in the lab. Six items are a classroom diagnostic, not benchmark performance.

Ask: Preserve the raw answer, reveal the reference, then judge support.

Answer: Dog, desk, spatial, invoice, chart and false-premise questions span distinct failure modes. Mark supported / partial / unsupported in the lab. Six items are a classroom diagnostic, not benchmark performance.

09 · SUMMARY

SUMMARY

09

Represent the image, expose context, predict the next token.

Original teaching example / diagram147 / 155

Teaching notes

Represent the image, expose context, predict the next token.

Ask:

Answer: Represent the image, expose context, predict the next token.

09 · SUMMARY

ViT → CLIP → generative VLM

ViT represents vision; CLIP compares independently encoded vision and text; generative VLM conditions language on visual contextViTImageVisual statesTask headCLIPImage → uText → vSimilarityGenerative VLMImage → contextQuestion + prefixDecoder → next token

Visual representation → cross-modal matching → conditioned generation.

Original teaching example / diagram148 / 155

Teaching notes

CLIP remains a vision-language model in the broad sense. The progression distinguishes what each system is trained and used to produce.

Ask: Visual representation → cross-modal matching → conditioned generation.

Answer: CLIP remains a vision-language model in the broad sense. The progression distinguishes what each system is trained and used to produce.

09 · SUMMARY

One computation to remember

Canonical VLM: image, encoder, visual states, connector, visual context, prompt, decoder state, LM head and tokenImageVision encoderVisual statesHVN × dVConnectorVisual contextZM × dLQuestion + prefixWhat animal is shown?Language decoderStateLM headTokenlogits → softmax
p(yt∣y<t,I,P)p(y_t\mid y_{<t},I,P)

Visual information becomes context for next-token prediction.

Original teaching example / diagram149 / 155

Teaching notes

Exit ticket: distinguish width matching, token compression, and access pattern. Then explain how an answer target can train a connector through a frozen language model.

Ask: Visual information becomes context for next-token prediction.

Answer: Exit ticket: distinguish width matching, token compression, and access pattern. Then explain how an answer target can train a connector through a frozen language model.

OPTIONAL · NEXT LECTURE / SOURCES

Different benchmarks probe different capabilities

Capability map to representative benchmark families, not a leaderboardPerceptionOCR · documentsOCRBench · DocVQAReasoningCharts · maths · subjectsChartQA · MathVista · MMMUIntegrated skillsBroad + combined tasksMMBench · MM-VetReliabilityIllusions · hallucinationHallusionBench

Choose the evaluation from the capability you need.

OCRBench · DocVQA · ChartQA · MathVista · MMMU · MMBench · MM-Vet · HallusionBench150 / 155

Teaching notes

71–73 min. Representative benchmark families, not current rankings. MMMU covers 30 subjects; MathVista emphasizes visual mathematical reasoning. MMBench and MM-Vet test different broad or integrated capabilities. Multi-turn and multi-image interaction need additional protocols.

Ask: Would high OCR performance establish good spatial reasoning?

Answer: No. The capabilities and failure modes differ.

OPTIONAL · NEXT LECTURE / SOURCES

Inspect the questions before interpreting a score

Invoice

Read: What is the invoice ID?

Chart

Calculate: March minus January?

Triangle

Reason: What is x?

Coffee

Reject: What colour is the bicycle?

Authored classroom examples, not copied benchmark items.

A wrong answer may come from perception, reasoning, or accepting a false premise.

Original teaching example / diagram151 / 155

Teaching notes

73–74 min. Ask students to localize the error: wrong extracted number versus wrong subtraction, for example. The benchmark names motivate task families; these four examples are not official items.

Ask: How could you separate a chart-reading error from an arithmetic error?

Answer: First ask for the two values, then ask for their difference.

OPTIONAL · NEXT LECTURE / SOURCES

Use a metric that matches the answer format

Output / useMeasureWhat still needs inspection
Fixed answerAccuracyAmbiguity and answer normalization
Document textExact / normalized string match; ANLSOCR errors and layout dependence
Open explanationHuman or specified model rubricJudge bias and unsupported details
Grounding coordinatesIoU / pointing accuracyObject identity and spatial accuracy
DeploymentLatency, memory, output tokens/sHardware and decoding settings

Report the evaluation protocol along with the number.

DocVQA152 / 155

Teaching notes

74–75 min. Recall@K is relevant to retrieval systems such as CLIP, not the default metric for generated answers. ANLS is average normalized Levenshtein similarity; DocVQA uses a thresholded normalized edit-similarity convention. Do not treat model grading as ground truth.

Ask: Report the evaluation protocol along with the number.

Answer: 74–75 min. Recall@K is relevant to retrieval systems such as CLIP, not the default metric for generated answers. ANLS is average normalized Levenshtein similarity; DocVQA uses a thresholded normalized edit-similarity convention. Do not treat model grading as ground truth.

OPTIONAL · NEXT LECTURE / SOURCES

Further training can change how the model answers

Behaviour

Instruction following
Preference optimisation
Safer responses

Evidence use

Better grounding data
Hard negative examples
Missing-evidence responses

The basic inference target remains image-conditioned next-token prediction.

Original teaching example / diagram153 / 155

Teaching notes

65–66 min. Optional post-training is a single-slide preview. Do not claim every model uses every method. Return to the original conditional distribution.

Ask: The basic inference target remains image-conditioned next-token prediction.

Answer: 65–66 min. Optional post-training is a single-slide preview. Do not claim every model uses every method. Return to the original conditional distribution.

OPTIONAL · NEXT LECTURE / SOURCES

Sources · architectures and browser implementation

CLIP · Radford et al., 2021
https://arxiv.org/abs/2103.00020

LLaVA · Liu et al., 2023
https://arxiv.org/abs/2304.08485

Flamingo · Alayrac et al., 2022
https://arxiv.org/abs/2204.14198

BLIP-2 · Li et al., 2023
https://arxiv.org/abs/2301.12597

SmolVLM-256M · model card
https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct

Official Transformers.js browser example
https://github.com/huggingface/transformers.js-examples/tree/main/smolvlm-webgpu

Oxford-IIIT Pet · CC BY-SA 4.0
https://www.robots.ox.ac.uk/~vgg/data/pets/

Primary sources; no leaderboard scores are used in this lecture.

Original teaching example / diagram154 / 155

Teaching notes

Source links checked 1 October 2026. Full provenance and limitations are in SOURCES.md. Sources describe named architectures; drawings are original simplifications.

Ask: Primary sources; no leaderboard scores are used in this lecture.

Answer: Source links checked 1 October 2026. Full provenance and limitations are in SOURCES.md. Sources describe named architectures; drawings are original simplifications.

OPTIONAL · NEXT LECTURE / SOURCES

Sources · evaluation

OCRBench
https://arxiv.org/abs/2305.07895

DocVQA
https://arxiv.org/abs/2007.00398

ChartQA
https://arxiv.org/abs/2203.10244

MathVista
https://arxiv.org/abs/2310.02255

MMMU
https://arxiv.org/abs/2311.16502

MMBench
https://arxiv.org/abs/2307.06281

MM-Vet
https://arxiv.org/abs/2308.02490

HallusionBench
https://arxiv.org/abs/2310.14566

The local six-image suite is an instructional diagnostic, not one of these benchmarks.

Original teaching example / diagram155 / 155

Teaching notes

Benchmark examples in the deck are authored analogues and are not official samples. Do not compare the six-case result against published benchmark scores.

Ask: The local six-image suite is an instructional diagnostic, not one of these benchmarks.

Answer: Benchmark examples in the deck are authored analogues and are not official samples. Do not compare the six-case result against published benchmark scores.