00 · WHAT CAN A VLM DO?
FROM CLIP TO
VISION-LANGUAGE MODELS
How Images Become Context
for Language Generation
How does visual information become context for next-token prediction?
ES 667 · Nipun Batra · IIT Gandhinagar
How does visual information become context for next-token prediction?
Oxford-IIIT Pet · CC BY-SA 4.001 / 155
80-minute teaching route; progressive builds are physical slides. The model lab and measured records are separate from illustrative arithmetic.
Ask: ES 667 · Nipun Batra · IIT Gandhinagar
Answer: 80-minute teaching route; progressive builds are physical slides. The model lab and measured records are separate from illustrative arithmetic.
00 · WHAT CAN A VLM DO?
WHAT CAN A VLM DO?
00
Start with the evidence an answer needs.
How does visual information become context for next-token prediction?
Original teaching example / diagram02 / 155
Start with the evidence an answer needs.
Ask:
Answer: Start with the evidence an answer needs.
00 · WHAT CAN A VLM DO?
Can the answer use visible evidence?
QUESTION
Is the laptop plugged in?
How does visual information become context for next-token prediction?
Original teaching example / diagram03 / 155
A visible connection does not prove power is flowing.
Ask: Is the laptop plugged in?
Answer: Yes. A cable connects the laptop to the socket.
00 · WHAT CAN A VLM DO?
Can the answer use visible evidence?
REFERENCE ANSWER
Is the laptop plugged in?
Yes. A cable connects the laptop to the socket.
How does visual information become context for next-token prediction?
Original teaching example / diagram04 / 155
A visible connection does not prove power is flowing.
Ask: Is the laptop plugged in?
Answer: Yes. A cable connects the laptop to the socket.
00 · WHAT CAN A VLM DO?
Can the answer use visible evidence?
ARCHITECTURAL REQUIREMENT
Is the laptop plugged in?
Yes. A cable connects the laptop to the socket.
Trace a connection between objects.
How does visual information become context for next-token prediction?
Original teaching example / diagram05 / 155
A visible connection does not prove power is flowing.
Ask: Is the laptop plugged in?
Answer: Yes. A cable connects the laptop to the socket.
00 · WHAT CAN A VLM DO?
Identity is only part of the question
QUESTION
What is immediately left of the mug?
How does visual information become context for next-token prediction?
Original teaching example / diagram06 / 155
The cube is also above the book; use this as a second oral question.
Ask: What is immediately left of the mug?
Answer: The red ball.
00 · WHAT CAN A VLM DO?
Identity is only part of the question
REFERENCE ANSWER
What is immediately left of the mug?
The red ball.
How does visual information become context for next-token prediction?
Original teaching example / diagram07 / 155
The cube is also above the book; use this as a second oral question.
Ask: What is immediately left of the mug?
Answer: The red ball.
00 · WHAT CAN A VLM DO?
Identity is only part of the question
ARCHITECTURAL REQUIREMENT
What is immediately left of the mug?
The red ball.
Keep object identity and spatial relations.
How does visual information become context for next-token prediction?
Original teaching example / diagram08 / 155
The cube is also above the book; use this as a second oral question.
Ask: What is immediately left of the mug?
Answer: The red ball.
00 · WHAT CAN A VLM DO?
The question selects a subset
QUESTION
How many apples are outside the bowl?
How does visual information become context for next-token prediction?
Original teaching example / diagram09 / 155
Five apples total; two in the bowl and three outside.
Ask: How many apples are outside the bowl?
Answer: Three.
00 · WHAT CAN A VLM DO?
The question selects a subset
REFERENCE ANSWER
How many apples are outside the bowl?
Three.
How does visual information become context for next-token prediction?
Original teaching example / diagram10 / 155
Five apples total; two in the bowl and three outside.
Ask: How many apples are outside the bowl?
Answer: Three.
00 · WHAT CAN A VLM DO?
The question selects a subset
ARCHITECTURAL REQUIREMENT
How many apples are outside the bowl?
Three.
Represent multiple objects and the region that selects them.
How does visual information become context for next-token prediction?
Original teaching example / diagram11 / 155
Five apples total; two in the bowl and three outside.
Ask: How many apples are outside the bowl?
Answer: Three.
00 · WHAT CAN A VLM DO?
Read first, then calculate
QUESTION
What is the invoice ID?
What percentage of the subtotal is GST?
How does visual information become context for next-token prediction?
Original teaching example / diagram12 / 155
Synthetic document. Separate text extraction errors from arithmetic errors.
Ask: What is the invoice ID? What percentage of the subtotal is GST?
Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.
00 · WHAT CAN A VLM DO?
Read first, then calculate
REFERENCE ANSWER
What is the invoice ID?
What percentage of the subtotal is GST?
ES667-042. GST is 180 / 1000 × 100 = 18%.
How does visual information become context for next-token prediction?
Original teaching example / diagram13 / 155
Synthetic document. Separate text extraction errors from arithmetic errors.
Ask: What is the invoice ID? What percentage of the subtotal is GST?
Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.
00 · WHAT CAN A VLM DO?
Read first, then calculate
ARCHITECTURAL REQUIREMENT
What is the invoice ID?
What percentage of the subtotal is GST?
ES667-042. GST is 180 / 1000 × 100 = 18%.
Preserve fine text, numbers and their roles.
How does visual information become context for next-token prediction?
Original teaching example / diagram14 / 155
Synthetic document. Separate text extraction errors from arithmetic errors.
Ask: What is the invoice ID? What percentage of the subtotal is GST?
Answer: ES667-042. GST is 180 / 1000 × 100 = 18%.
00 · WHAT CAN A VLM DO?
The labels change the answer
QUESTION
Which month is highest?
How much higher is March than January?
How does visual information become context for next-token prediction?
Original teaching example / diagram15 / 155
Synthetic values: January 30, February 45, March 75, April 40.
Ask: Which month is highest? How much higher is March than January?
Answer: March. The difference is 45 µg/m³.
00 · WHAT CAN A VLM DO?
The labels change the answer
REFERENCE ANSWER
Which month is highest?
How much higher is March than January?
March. The difference is 45 µg/m³.
How does visual information become context for next-token prediction?
Original teaching example / diagram16 / 155
Synthetic values: January 30, February 45, March 75, April 40.
Ask: Which month is highest? How much higher is March than January?
Answer: March. The difference is 45 µg/m³.
00 · WHAT CAN A VLM DO?
The labels change the answer
ARCHITECTURAL REQUIREMENT
Which month is highest?
How much higher is March than January?
March. The difference is 45 µg/m³.
Bind chart labels to values and units.
How does visual information become context for next-token prediction?
Original teaching example / diagram17 / 155
Synthetic values: January 30, February 45, March 75, April 40.
Ask: Which month is highest? How much higher is March than January?
Answer: March. The difference is 45 µg/m³.
00 · WHAT CAN A VLM DO?
A diagram supplies conditions
QUESTION
What is x?
How does visual information become context for next-token prediction?
Original teaching example / diagram18 / 155
The right-angle marker, rather than appearance alone, licenses Pythagoras.
Ask: What is x?
Answer: x = 5, using the marked right angle and sides 3 and 4.
00 · WHAT CAN A VLM DO?
A diagram supplies conditions
REFERENCE ANSWER
What is x?
x = 5, using the marked right angle and sides 3 and 4.
How does visual information become context for next-token prediction?
Original teaching example / diagram19 / 155
The right-angle marker, rather than appearance alone, licenses Pythagoras.
Ask: What is x?
Answer: x = 5, using the marked right angle and sides 3 and 4.
00 · WHAT CAN A VLM DO?
A diagram supplies conditions
ARCHITECTURAL REQUIREMENT
What is x?
x = 5, using the marked right angle and sides 3 and 4.
Combine diagram evidence with a mathematical rule.
How does visual information become context for next-token prediction?
Original teaching example / diagram20 / 155
The right-angle marker, rather than appearance alone, licenses Pythagoras.
Ask: What is x?
Answer: x = 5, using the marked right angle and sides 3 and 4.
00 · WHAT CAN A VLM DO?
Compare presence and position
QUESTION
What changed from A to B?
How does visual information become context for next-token prediction?
Original teaching example / diagram21 / 155
Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.
Ask: What changed from A to B?
Answer: The chair moved to the right; the cup disappeared.
00 · WHAT CAN A VLM DO?
Compare presence and position
REFERENCE ANSWER
What changed from A to B?
The chair moved to the right; the cup disappeared.
How does visual information become context for next-token prediction?
Original teaching example / diagram22 / 155
Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.
Ask: What changed from A to B?
Answer: The chair moved to the right; the cup disappeared.
00 · WHAT CAN A VLM DO?
Compare presence and position
ARCHITECTURAL REQUIREMENT
What changed from A to B?
The chair moved to the right; the cup disappeared.
Compare evidence across images.
How does visual information become context for next-token prediction?
Original teaching example / diagram23 / 155
Two panels in a single composite for the small browser model. Native multi-image systems can take separate ordered images.
Ask: What changed from A to B?
Answer: The chair moved to the right; the cup disappeared.
00 · WHAT CAN A VLM DO?
A screenshot is a spatial interface
QUESTION
Where should I click to disable notifications?
How does visual information become context for next-token prediction?
Original teaching example / diagram24 / 155
Describing a control is distinct from executing a click.
Ask: Where should I click to disable notifications?
Answer: The blue toggle on the Notifications row.
00 · WHAT CAN A VLM DO?
A screenshot is a spatial interface
REFERENCE ANSWER
Where should I click to disable notifications?
The blue toggle on the Notifications row.
How does visual information become context for next-token prediction?
Original teaching example / diagram25 / 155
Describing a control is distinct from executing a click.
Ask: Where should I click to disable notifications?
Answer: The blue toggle on the Notifications row.
00 · WHAT CAN A VLM DO?
A screenshot is a spatial interface
ARCHITECTURAL REQUIREMENT
Where should I click to disable notifications?
The blue toggle on the Notifications row.
Locate a control relative to its label.
How does visual information become context for next-token prediction?
Original teaching example / diagram26 / 155
Describing a control is distinct from executing a click.
Ask: Where should I click to disable notifications?
Answer: The blue toggle on the Notifications row.
00 · WHAT CAN A VLM DO?
The image can be a scientific figure
QUESTION
What is the hottest measured region?
How does visual information become context for next-token prediction?
Original teaching example / diagram27 / 155
Synthetic test-plate measurements; count rows and columns from the top left.
Ask: What is the hottest measured region?
Answer: Row 3, column 4: 62 °C.
00 · WHAT CAN A VLM DO?
The image can be a scientific figure
REFERENCE ANSWER
What is the hottest measured region?
Row 3, column 4: 62 °C.
How does visual information become context for next-token prediction?
Original teaching example / diagram28 / 155
Synthetic test-plate measurements; count rows and columns from the top left.
Ask: What is the hottest measured region?
Answer: Row 3, column 4: 62 °C.
00 · WHAT CAN A VLM DO?
The image can be a scientific figure
ARCHITECTURAL REQUIREMENT
What is the hottest measured region?
Row 3, column 4: 62 °C.
Read a scale and local measurements.
How does visual information become context for next-token prediction?
Original teaching example / diagram29 / 155
Synthetic test-plate measurements; count rows and columns from the top left.
Ask: What is the hottest measured region?
Answer: Row 3, column 4: 62 °C.
00 · WHAT CAN A VLM DO?
The question can be wrong
QUESTION
What color is the bicycle?
How does visual information become context for next-token prediction?
Original teaching example / diagram30 / 155
Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.
Ask: What color is the bicycle?
Answer: No bicycle is visible.
00 · WHAT CAN A VLM DO?
The question can be wrong
REFERENCE ANSWER
What color is the bicycle?
No bicycle is visible.
How does visual information become context for next-token prediction?
Original teaching example / diagram31 / 155
Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.
Ask: What color is the bicycle?
Answer: No bicycle is visible.
00 · WHAT CAN A VLM DO?
The question can be wrong
ARCHITECTURAL REQUIREMENT
What color is the bicycle?
No bicycle is visible.
Let visual evidence override the wording of the question.
How does visual information become context for next-token prediction?
Original teaching example / diagram32 / 155
Coffee photo by Rachel Michetti, CC0 via scikit-image. A color answer would be unsupported.
Ask: What color is the bicycle?
Answer: No bicycle is visible.
00 · WHAT CAN A VLM DO?
What must survive the visual representation?
| Question family | Information needed |
|---|---|
| Counting / spatial | Multiple objects and their arrangement |
| OCR / charts | Fine detail, labels, values and units |
| UI / multiple images | Locations and correspondence across views |
| False premise | Evidence that can contradict the prompt |
Would one global CLIP vector always be enough?
How does visual information become context for next-token prediction?
Original teaching example / diagram33 / 155
A global vector can be useful, but fine-grained tasks motivate exposing richer visual states. This is a design motivation, not an impossibility theorem.
Ask: Would one global CLIP vector always be enough?
Answer: A global vector can be useful, but fine-grained tasks motivate exposing richer visual states. This is a design motivation, not an impossibility theorem.
00 · WHAT CAN A VLM DO?
Try the same questions on a small model
Predict → run → inspect the exact answer.
Open the VLM lab ↗11 image presets, recorded outputs, and an image-swap experiment.
A reference answer is not a model prediction.
How does visual information become context for next-token prediction?
SmolVLM-256M · model card · Official Transformers.js browser example34 / 155
Preload the model before class. First download needs network; recorded outputs are available offline. Run the dog, spatial and absent-object cases.
Ask: A reference answer is not a model prediction.
Answer: Preload the model before class. First download needs network; recorded outputs are available offline. Run the dog, spatial and absent-object cases.
01 · FROM MATCHING TO GENERATION
FROM MATCHING TO GENERATION
01
Matching supplied text leaves a generation problem.
How does visual information become context for next-token prediction?
Original teaching example / diagram35 / 155
Matching supplied text leaves a generation problem.
Ask:
Answer: Matching supplied text leaves a generation problem.
01 · FROM MATCHING TO GENERATION
Recall: CLIP compares an image with supplied text
Image vector u · text vector v · similarity uᵀv
How does visual information become context for next-token prediction?
CLIP · Radford et al., 202136 / 155
Vectors are normalized for the cosine-style comparison. CLIP is a VLM broadly; here we distinguish contrastive matching from generative VLMs.
Ask: Image vector u · text vector v · similarity uᵀv
Answer: Vectors are normalized for the cosine-style comparison. CLIP is a VLM broadly; here we distinguish contrastive matching from generative VLMs.
01 · FROM MATCHING TO GENERATION
Now remove the candidate description
Desired output: ? ? ?
How does visual information become context for next-token prediction?
Original teaching example / diagram37 / 155
The image encoder remains useful. But a matching score alone cannot produce a next-token vocabulary distribution.
Ask: Desired output: ? ? ?
Answer: The image encoder remains useful. But a matching score alone cannot produce a next-token vocabulary distribution.
01 · FROM MATCHING TO GENERATION
We already know how to generate words
A causal language model maps a prefix to a next-token distribution.
How does visual information become context for next-token prediction?
Original teaching example / diagram38 / 155
Recall the decoder-only model from Beyond Attention before attaching vision.
Ask: A causal language model maps a prefix to a next-token distribution.
Answer: Recall the decoder-only model from Beyond Attention before attaching vision.
01 · FROM MATCHING TO GENERATION
We have two useful pieces
How can visual states enter the language computation?
How does visual information become context for next-token prediction?
Original teaching example / diagram39 / 155
Keep the two branches disconnected. Do not draw a connector until we derive what it must do.
Ask: How can visual states enter the language computation?
Answer: Keep the two branches disconnected. Do not draw a connector until we derive what it must do.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
HOW CAN VISION ENTER A LANGUAGE MODEL?
02
First, see the three ideas. Then build each connection.
How does visual information become context for next-token prediction?
Original teaching example / diagram40 / 155
First, see the three ideas. Then build each connection.
Ask:
Answer: First, see the three ideas. Then build each connection.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Idea 1 · prepend the image as visual tokens
Keep all N patch vectors in this design; N depends on the image grid.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 202341 / 155
Trace both paths before naming the join. The vision encoder produces one contextual feature row per patch; a row-wise projector adapts every row to the decoder width. V1, V2, …, VN represents the full N-row sequence, not a selection of four patches. The question and answer generated so far are token IDs with embedding vectors; concatenate their embeddings after the image vectors. At the first step, no answer tokens exist. Generation appends one token at a time; cached computation can reuse earlier states. Chat/system markers and positional handling are omitted here. No separate text encoder is added: tokenization and embedding lookup feed the causal language decoder. Show all three ideas before deriving patch grids, widths or weights.
Ask: Do we keep only four patches, and does the decoder receive raw question text?
Answer: No. This basic prefix design keeps all N patch rows and projects each one. The question and already-generated answer tokens are embedded; their vectors follow V1, V2, …, VN in one decoder input sequence. Before the first prediction, the answer portion is empty.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Idea 2 · let the decoder look at a visual memory
As in translation: encode the source, then let the text decoder read its states.
How does visual information become context for next-token prediction?
Transformer · Vaswani et al., 2017 · ViT · Dosovitskiy et al., 202042 / 155
Visual memory is the sequence produced by the vision encoder, not a second encoder or a separate learned memory bank. h1, h2, …, hN abbreviates all N retained image feature rows. The question and answer-so-far tokens become e1, e2, …, eT through embedding lookup; T is the current text length. The answer is empty for the first prediction and grows only from generated token IDs. A causal decoder needs cross-attention sublayers to use this route: text states supply queries, visual features supply keys and values, and the attention result updates text states before the vocabulary head scores the next token. This is the sequence-to-sequence connection from English-to-Hindi translation, with the image as source. The overview shows the two attention operations and the MLP/residual updates; normalization, projection details, positional handling, chat markers and repeated blocks are omitted. The example uses a 224×224 image and 16×16 patches: 14×14=196 patch rows; retaining the global CLS row gives 197. Neither is universal. Compression can shorten this memory later; Flamingo also resamples features and uses gated cross-attention, so this is not its exact architecture.
Ask: Is visual memory another model? Does cross-attention itself produce the answer word?
Answer: Visual memory is the encoder’s sequence of image features. The language decoder reads it through cross-attention, like a translation decoder reading source-language states. That updates text states; the vocabulary head then scores the next token. N is not fixed: a 224×224 image with 16×16 patches gives 196 patch rows, or 197 rows if CLS is retained.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Idea 3 · summarize the image, then use either connection
Compression reduces N image rows to M summaries; either decoder connection remains possible.
How does visual information become context for next-token prediction?
BLIP-2 · Li et al., 2023 · Flamingo · Alayrac et al., 202243 / 155
Read the common image path first: the ViT encoder uses full self-attention and exposes all N contextual patch rows. A learned-query summarizer uses M shared trainable query vectors to cross-attend to those image features, producing M image-dependent summaries. This is a representative resampler mechanism, not a subset selection and not a complete implementation of a named model. N > M; omitted boxes stand for intermediate rows. Now choose one of two alternative decoder connections. A projects each summary to the decoder width, concatenates those vectors before text embeddings, and uses causal self-attention over the joined sequence. Projection changes width, not the already-reduced M row count. B retains S1…SM as separate memory: text embeddings enter causal self-attention, then text states supply Q while summaries supply K,V for decoder cross-attention. The vocabulary head produces next-token scores in both routes. The compression cross-attention uses learned queries; decoder cross-attention uses text-derived queries. Question tokens and previously generated answer tokens are embedded; the answer starts empty. v labels are projected summary vectors; e labels are text embeddings. A and B are alternatives, not consecutive stages. Decoder MLPs, residuals, normalization, position handling and repeated blocks are abstracted in this overview and restored in the deep dives. Source notes discuss BLIP-2 and Flamingo as examples with additional model-specific components.
Ask: After producing M summaries, must language read them with cross-attention?
Answer: No. In route A, project the M summaries, prepend them to the text embeddings, and use decoder causal self-attention. In route B, keep the M summaries separate and let decoder cross-attention read them. The summarizer itself can use learned-query cross-attention in either route; that is a different operation from the decoder’s text-query cross-attention.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The map: two access routes, one compression choice
| Idea | High-level decision | Then we will explain… |
|---|---|---|
| 1 · Visual prefix | Put visual vectors before text. | How to make them fit the decoder input. |
| 2 · Separate memory | Let language read the image. | How a text state retrieves visual information. |
| 3 · Shorter summary | Use fewer visual vectors. | How learned readers make the summaries. |
Ideas 1 and 2 choose the access route. Idea 3 can be used with either.
How does visual information become context for next-token prediction?
Original teaching example / diagram44 / 155
Spend roughly five minutes on the three preceding idea slides, with no matrix shapes or attention arithmetic. Ask students to explain each idea in one sentence. Now return to Idea 1 and derive it. Ideas 2 and 3 get their own later deep dives.
Ask: Must a model choose only one of these three ideas?
Answer: No. Prefix and separate memory are two ways to access vision. Compression can be combined with either; for example, a short visual summary can be prepended to the text.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Idea 1 in detail · build the visual prefix
One sequence; the same causal self-attention can use both kinds of context.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 202345 / 155
Trace all N patch rows through the amber projector into the input row. Trace question and answer-so-far tokenization and embedding into the same row. The answer is initially empty, then grows by generated tokens. V1, V2, …, VN abbreviates the full visual sequence. Point to causal self-attention, MLP and vocabulary head from the previous lecture. Positions, normalization and exact chat markers are omitted. We will now unpack only the new visual path.
Ask: One sequence; the same causal self-attention can use both kinds of context.
Answer: Trace all N patch rows through the amber projector into the input row. Trace question and answer-so-far tokenization and embedding into the same row. The answer is initially empty, then grows by generated tokens. V1, V2, …, VN abbreviates the full visual sequence. Point to causal self-attention, MLP and vocabulary head from the previous lecture. Positions, normalization and exact chat markers are omitted. We will now unpack only the new visual path.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
What must we build to make this work?
| Job | Why it is needed |
|---|---|
| 1. Keep image feature vectors | The decoder needs access to visual evidence. |
| 2. Match the decoder input width | All input rows must have the same number of coordinates. |
| 3. Place them before the question | Causal attention can read earlier context. |
We already have the decoder. Now build its visual input.
How does visual information become context for next-token prediction?
Original teaching example / diagram46 / 155
Return to this three-job plan after the projector calculation and after the mask. Students should be able to point to each job in the preceding complete architecture.
Ask: We already have the decoder. Now build its visual input.
Answer: Return to this three-job plan after the projector calculation and after the mask. Students should be able to point to each job in the preceding complete architecture.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Recall ViT: every patch already has an output row
Patch 1 has row h1, patch 2 has row h2, and so on through patch 12.
How does visual information become context for next-token prediction?
ViT · Dosovitskiy et al., 202047 / 155
The numbered 3 × 4 grid is an illustrative twelve-patch example, not a model configuration. All twelve patch rows are drawn, in row-major order matching the image labels. Each box denotes one whole feature vector, not one scalar coordinate. ViT embeds the patches, adds positions, and updates their states with full self-attention. The row at patch position i can therefore contain information from the whole image. The CLS row is introduced on the next slide.
Ask: Patch 1 has row h1, patch 2 has row h2, and so on through patch 12.
Answer: The numbered 3 × 4 grid is an illustrative twelve-patch example, not a model configuration. All twelve patch rows are drawn, in row-major order matching the image labels. Each box denotes one whole feature vector, not one scalar coordinate. ViT embeds the patches, adds positions, and updates their states with full self-attention. The row at patch position i can therefore contain information from the whole image. The CLS row is introduced on the next slide.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
ViT outputs a CLS state and one row per patch
Here: 1 CLS row + 12 patch rows. We keep all 12 patch rows.
How does visual information become context for next-token prediction?
ViT · Dosovitskiy et al., 202048 / 155
Use the same twelve numbered image patches and twelve output rows as the previous slide. This standard ViT also has a CLS token, introduced before the encoder; its output state provides a whole-image summary. CLS is an extra sequence position, not a thirteenth image patch or a separate encoder. The diagram abstracts patch embedding, positions, and the input CLS token inside the ViT path. Its twelve patch states are contextual vectors, not isolated patch descriptors. For the visual-prefix design being developed, retain all twelve patch rows and then project each one to the decoder width; the CLS state is not included in that patch sequence. Twelve is the complete illustrative example, not a universal model count. Keep this explanation about ViT outputs; the earlier CLIP matching discussion need not be repeated.
Ask: Here: 1 CLS row + 12 patch rows. We keep all 12 patch rows.
Answer: Use the same twelve numbered image patches and twelve output rows as the previous slide. This standard ViT also has a CLS token, introduced before the encoder; its output state provides a whole-image summary. CLS is an extra sequence position, not a thirteenth image patch or a separate encoder. The diagram abstracts patch embedding, positions, and the input CLS token inside the ViT path. Its twelve patch states are contextual vectors, not isolated patch descriptors. For the visual-prefix design being developed, retain all twelve patch rows and then project each one to the decoder width; the CLS state is not included in that patch sequence. Twelve is the complete illustrative example, not a universal model count. Keep this explanation about ViT outputs; the earlier CLIP matching discussion need not be repeated.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Why keep more than the global summary?
Recall our opening question
What is the invoice ID?
What percentage is GST?
The answer needs small text, several numbers, and their roles.
Give language access to more visual evidence before asking it to answer.
How does visual information become context for next-token prediction?
Original teaching example / diagram49 / 155
Return to the Newfoundland image on the next frame. This invoice explains the design motivation. Keeping patch states does not guarantee OCR success; resolution and training still matter.
Ask: Give language access to more visual evidence before asking it to answer.
Answer: Return to the Newfoundland image on the next frame. This invoice explains the design motivation. Keeping patch states does not guarantee OCR success; resolution and training still matter.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
A “visual state” is just one row of features
One image → many rows
Read the notation in plain language
N = number of visual rows.
dV = numbers in each row.
hᵢ = the feature row at patch position i.
Rows are vectors, not object labels or named semantic attributes.
How does visual information become context for next-token prediction?
Original teaching example / diagram50 / 155
Use “row” and “vector” interchangeably here. Attention has mixed image context into each row. Do not label one row “dog” or claim that one coordinate is fur colour.
Ask: Rows are vectors, not object labels or named semantic attributes.
Answer: Use “row” and “vector” interchangeably here. Attention has mixed image context into each row. Do not label one row “dog” or claim that one coordinate is fur colour.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Recall: the decoder already consumes vectors
Text IDs become embeddings before reaching the Transformer blocks.
How does visual information become context for next-token prediction?
Original teaching example / diagram51 / 155
This is the connection to the earlier text architecture: tokenization selects embedding rows. Image vectors bypass a word-ID lookup and enter at the vector interface. They also need the implementation’s positional treatment.
Ask: Text IDs become embeddings before reaching the Transformer blocks.
Answer: This is the connection to the earlier text architecture: tokenization selects embedding rows. Image vectors bypass a word-ID lookup and enter at the vector interface. They also need the implementation’s positional treatment.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The image rows have 2 coordinates; text rows have 3
To form one input sequence, every vector must have the decoder’s input width.
How does visual information become context for next-token prediction?
Original teaching example / diagram52 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Ask students to point to one row, then count its coordinates. The image and text may have different numbers of rows; the problem is their different row widths.
Ask: Do N and T need to match, or must the number of coordinates per row match?
Answer: The row widths must match. N image rows and T text rows can have different counts. After projecting the image rows to width 3, concatenate N + T rows for the decoder.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Add a projector to the image path
The projector changes N × 2 image features into N × 3 visual vectors.
How does visual information become context for next-token prediction?
Original teaching example / diagram53 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Read each shape beneath the corresponding visible module. The text embeddings are already T × 3 and remain unchanged. In general Hᵥ is N × dV, Wₚ is dV × dL, and Z = HᵥWₚ is N × dL. Bias is omitted. Real connectors can use an MLP. Matching shapes only makes the interface valid; learning must make the projected contents useful.
Ask: The projector changes N × 2 image features into N × 3 visual vectors.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Read each shape beneath the corresponding visible module. The text embeddings are already T × 3 and remain unchanged. In general Hᵥ is N × dV, Wₚ is dV × dL, and Z = HᵥWₚ is N × dL. Bias is omitted. Real connectors can use an MLP. Matching shapes only makes the interface valid; learning must make the projected contents useful.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Zoom into one row from the image
Take one 2-coordinate row from the image feature sequence.
How does visual information become context for next-token prediction?
Original teaching example / diagram54 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: Take one 2-coordinate row from the image feature sequence.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Use a 2 × 3 matrix to change the width
Each column makes one of the 3 decoder-input coordinates.
How does visual information become context for next-token prediction?
Original teaching example / diagram55 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: Each column makes one of the 3 decoder-input coordinates.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The first column makes the first coordinate
Which input coordinates does this column combine?
How does visual information become context for next-token prediction?
Original teaching example / diagram56 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: Which input coordinates does this column combine?
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The first output coordinate is 2
The first output coordinate copies the first input in this toy map.
How does visual information become context for next-token prediction?
Original teaching example / diagram57 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: The first output coordinate copies the first input in this toy map.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The second column reuses the same image row
Now use the same input row with the second column.
How does visual information become context for next-token prediction?
Original teaching example / diagram58 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: Now use the same input row with the second column.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The second output coordinate is 1
The second output coordinate copies the second input in this toy map.
How does visual information become context for next-token prediction?
Original teaching example / diagram59 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: The second output coordinate copies the second input in this toy map.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
The third column combines both input coordinates
The third output coordinate adds the two inputs in this toy map.
How does visual information become context for next-token prediction?
Original teaching example / diagram60 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: The third output coordinate adds the two inputs in this toy map.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Now this image vector has the required width
The row [2, 1, 3] fits beside the 3-coordinate text embeddings.
How does visual information become context for next-token prediction?
Original teaching example / diagram61 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
Ask: The row [2, 1, 3] fits beside the 3-coordinate text embeddings.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The highlighted h1 is the first image feature row in the persistent top diagram. Wₚ = [[1,0,1],[0,1,1]] is shared across all image rows. The weights are chosen for easy arithmetic; actual connector weights are learned. Each output coordinate is the input row dotted with one matrix column. The lower zoom changes while both input paths remain fixed. Predict the next coordinate before advancing.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Use the same matrix for all N image rows
N rows stay N rows; every projected row now has 3 coordinates.
How does visual information become context for next-token prediction?
Original teaching example / diagram62 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The same Wₚ gives [2,1,3] from [2,1], [1,3,4] from [1,3], and [0,2,2] from [0,2]. No patch is dropped and no independent projector is created for each patch. Text E is unchanged.
Ask: N rows stay N rows; every projected row now has 3 coordinates.
Answer: The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. The same Wₚ gives [2,1,3] from [2,1], [1,3,4] from [1,3], and [0,2,2] from [0,2]. No patch is dropped and no independent projector is created for each patch. Text E is unchanged.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Check: did the projector create more visual tokens?
Before
576 visual rows.
Each row has 1,024 coordinates.
After a row-wise projector
Each row has 4,096 coordinates.
How many rows now?
Predict the row count before advancing.
How does visual information become context for next-token prediction?
Original teaching example / diagram63 / 155
These are illustrative model widths. The operation acts on the last dimension, not the row dimension.
Ask: 576 rows × 1,024 coordinates become how many rows × 4,096 coordinates?
Answer: 576 rows. The projector changes each vector’s width; it does not compress or select patch positions.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Still 576 rows: width and count are different
Same positions. Different number of features at each position.
We will learn how to reduce the row count in the compression section.
How does visual information become context for next-token prediction?
Original teaching example / diagram64 / 155
If a student answered 32, explain that they have proposed an extra summarization operation. The simple LLaVA-style row-wise projection does not do that.
Ask: We will learn how to reduce the row count in the compression section.
Answer: If a student answered 32, explain that they have proposed an extra summarization operation. The simple LLaVA-style row-wise projection does not do that.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Now the two sequences fit one decoder input
The mask applies inside the decoder, including its visual positions.
How does visual information become context for next-token prediction?
Original teaching example / diagram65 / 155
The photograph and actual question provide the same visual context on every frame. All displayed embedding values and widths are hand-set, not measured from either input. Two and three mean coordinates per feature vector, not image dimensions or counts of tokens. N is the full retained patch-row count; T includes question and already-generated answer tokens. Middle rows are omitted only from the drawing. The illustrative answer so far is A black; it is empty before the first prediction. Real token boundaries, chat markers and position handling are omitted. Hᵥ contains contextual vision-encoder features. The text branch tokenizes to IDs and looks up embeddings; no separate text encoder is added. Z has N rows and E has T rows; their joined input is (N + T) × 3 in this toy example. Distinguish the two attention stages: the ViT uses full image self-attention before projection; the standard causal decoder applies a lower-triangular mask across every joined position, starting at z1 rather than switching it on only at the first text position. At a decoder layer, z1 reads itself, z2 reads z1 and itself, and the first text row reads all N preceding visual rows and itself. The visual vectors already contain whole-image context from the ViT, so a causal decoder does not undo that encoder context. With the image and prompt supplied, all their positions can be computed in parallel within each decoder layer during prefill; this does not generate image patches one at a time. Only new answer tokens are generated autoregressively. This is the basic LLaVA-style causal-prefix design, not a universal rule: some other VLMs allow bidirectional attention within an image block. System/role markers and positions are abstracted. Visual rows remain continuous vectors, not vocabulary IDs.
Ask: Does the causal mask start only when the decoder reaches the first text row?
Answer: In this design it covers the whole decoder sequence, including visual positions from z1. The ViT earlier used full image attention. Known image and prompt rows are processed together in a masked prefill pass; later answer tokens are generated one by one. The upcoming mask shows exactly which visual and text positions can read each other.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Build the input row: one projected visual vector
Visual vectors and text embeddings share the decoder’s input width.
How does visual information become context for next-token prediction?
Original teaching example / diagram66 / 155
Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
Ask: Visual vectors and text embeddings share the decoder’s input width.
Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Keep the rest of the projected visual sequence
Visual vectors and text embeddings share the decoder’s input width.
How does visual information become context for next-token prediction?
Original teaching example / diagram67 / 155
Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
Ask: Visual vectors and text embeddings share the decoder’s input width.
Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Then place the user’s text
Visual vectors and text embeddings share the decoder’s input width.
How does visual information become context for next-token prediction?
Original teaching example / diagram68 / 155
Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
Ask: Visual vectors and text embeddings share the decoder’s input width.
Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Finish the question; start the assistant’s answer
Visual vectors and text embeddings share the decoder’s input width.
How does visual information become context for next-token prediction?
Original teaching example / diagram69 / 155
Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
Ask: Visual vectors and text embeddings share the decoder’s input width.
Answer: Simplified chat template. Actual image markers, role markers and positional treatment depend on the implementation.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
What may the current text position read?
The assistant’s state can use both the picture and the question.
How does visual information become context for next-token prediction?
Original teaching example / diagram70 / 155
Read one query row first. The current position may also attend to itself; the simplified row here emphasizes earlier context.
Ask: The assistant’s state can use both the picture and the question.
Answer: Read one query row first. The current position may also attend to itself; the simplified row here emphasizes earlier context.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
In this decoder, the causal mask starts at V1
Colored: allowed · grey: blocked. The triangle includes the visual positions.
How does visual information become context for next-token prediction?
ViT · Dosovitskiy et al., 2020 · LLaVA · Liu et al., 2023 · Transformer · Vaswani et al., 201771 / 155
V1–V3 are the projected image vectors called z1…zN in the earlier numeric example; three visual positions make the mask readable. Rows are queries and columns are keys. Read the first three rows before the text: V1 can read V1, V2 can read V1 and V2, and V3 can read V1 through V3. No visual decoder query reads later visual positions or later text, but every visual input already incorporates image-wide context from the ViT encoder. Each text row can read all preceding visual rows, preceding text rows and itself. All supplied image/prompt rows are processed together with this mask during prefill, not generated one by one. New answer tokens are generated autoregressively afterwards. This is the ordinary LLaVA/LLaMA causal-mask example; some architectures instead use full attention within the visual prefix. The mask controls attention edges inside the decoder, not the ViT encoder.
Ask: Colored: allowed · grey: blocked. The triangle includes the visual positions.
Answer: V1–V3 are the projected image vectors called z1…zN in the earlier numeric example; three visual positions make the mask readable. Rows are queries and columns are keys. Read the first three rows before the text: V1 can read V1, V2 can read V1 and V2, and V3 can read V1 through V3. No visual decoder query reads later visual positions or later text, but every visual input already incorporates image-wide context from the ViT encoder. Each text row can read all preceding visual rows, preceding text rows and itself. All supplied image/prompt rows are processed together with this mask during prefill, not generated one by one. New answer tokens are generated autoregressively afterwards. This is the ordinary LLaVA/LLaMA causal-mask example; some architectures instead use full attention within the visual prefix. The mask controls attention edges inside the decoder, not the ViT encoder.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Back to the big picture: the image is earlier context
We have built all three jobs: visual rows, matching width, and a prefix.
How does visual information become context for next-token prediction?
Original teaching example / diagram72 / 155
No extra cross-attention layer is required by this access pattern.
Ask: We have built all three jobs: visual rows, matching width, and a prefix.
Answer: No extra cross-attention layer is required by this access pattern.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
This is the basic LLaVA-style connection
Projected patch vectors occupy earlier positions; they do not become word IDs.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 202373 / 155
A simple row-wise linear or MLP projector keeps the patch-row count. The few boxes are a drawn subset, not token selection. LLaVA variants differ in details. Later compression methods deliberately reduce the visual sequence.
Ask: Projected patch vectors occupy earlier positions; they do not become word IDs.
Answer: A simple row-wise linear or MLP projector keeps the patch-row count. The few boxes are a drawn subset, not token selection. LLaVA variants differ in details. Later compression methods deliberately reduce the visual sequence.
PAPER · LLaVA · NeurIPS 2023
Visual Instruction Tuning
Haotian Liu · Chunyuan Li · Qingyang Wu · Yong Jae Lee
Read the paper ↗1 · Vision encoder
CLIP ViT-L/14 supplies image features Zᵥ.
2 · Projection W
The original paper uses a linear map to match the LM embedding width.
3 · Language model
Vicuna reads the visual vectors and text, then predicts the response.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 202374 / 155
Figure 1 is cropped from page 4 of arXiv 2304.08485v2; the architecture itself is unaltered. Read the diagram bottom to top. The paper calls unprojected features Zv and projected visual embeddings Hv; our lecture uses Hv and Z respectively, so map by role rather than letter. The original connector is linear; later variants can use an MLP. The few drawn tokens are schematic. The figure abstracts tokenization, positions and repeated decoder blocks. Visual instruction tuning also contributes a data/training recipe; this slide focuses on the architecture.
Ask: Point to the modules that implement the connection we just derived.
Answer: Vision encoder: CLIP ViT-L/14 supplies image features Zᵥ. Projection W: The original paper uses a linear map to match the LM embedding width. Language model: Vicuna reads the visual vectors and text, then predicts the response.
02 · HOW CAN VISION ENTER A LANGUAGE MODEL?
Idea 1, end to end in one function
image
One image
question = “What animal is shown?”
answer_so_far = “It is a” (empty at the start)
def next_token(image, question, answer_so_far): patches = vit(image).patch_statesN × dV visual = projector(patches)N × dLM ids = tokenize(question, answer_so_far)T token IDs text = embed(ids)T × dLM inputs = concat([visual, text], axis="tokens")(N + T) × dLM hidden = causal_decoder(inputs)(N + T) × dLM scores = lm_head(hidden[-1])Vocabulary scores return argmax(scores)Next text token IDcausal_decoder stacks causal self-attention + MLP blocks over all input positions.
Display the returned token with the tokenizer; append its ID to the answer and predict again.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 2023 · Oxford-IIIT Pet · CC BY-SA 4.075 / 155
High-level pseudocode for one next-token prediction, not executable library calls. One image is shown and the batch dimension is omitted. Image preprocessing is folded into vit; patch_states selects all N contextual patch rows and excludes CLS in this design. ViT uses full image attention. The shared row-wise projector changes dV to dLM while preserving N. tokenize(question, answer_so_far) is a conceptual helper for constructing and tokenizing the text context in that order, not the tokenizer API for paired sequences. In a real model, use its chat template and processor; keep generated answer IDs rather than repeatedly retokenizing them. T counts all supplied text positions, including already-generated answer tokens. The answer portion is empty at the first prediction. concat stacks visual rows before text rows along the token axis, not the feature axis. The decoder consumes embeddings directly, with positional handling and a causal mask from the first visual position. causal_decoder means the full repeated Transformer block stack, including MLPs, normalization and residuals; it is not only one attention operation. hidden[-1] is the last supplied position and predicts the next, as-yet-unknown answer token. The language head produces a score for every vocabulary entry. argmax shows greedy selection; softmax is unnecessary for choosing the largest score. The result is a text token ID, not a visual token. Decode it with the tokenizer to display text. This deliberately unoptimized function summarizes the dependency graph. During actual generation, encode the fixed image once and reuse the decoder KV cache rather than recomputing the whole function on every token. Question and answer text are illustrative; no measured answer or embedding is claimed.
Ask: Which line combines image and text, and which state predicts the next token?
Answer: concat joins N projected visual rows and T text embeddings along the token axis. The causal decoder processes their combined sequence; the language head reads hidden[-1], the last supplied position, to score the next text token.
03 · COULD VISION REMAIN SEPARATE?
COULD VISION REMAIN SEPARATE?
03
Return to Idea 2: how can language read separate visual memory?
How does visual information become context for next-token prediction?
Original teaching example / diagram76 / 155
Return to Idea 2: how can language read separate visual memory?
Ask:
Answer: Return to Idea 2: how can language read separate visual memory?
03 · COULD VISION REMAIN SEPARATE?
Idea 2 in detail · read a separate visual memory
Causal self-attention reads text; cross-attention reads image memory.
How does visual information become context for next-token prediction?
Transformer · Vaswani et al., 201777 / 155
Trace the two inputs separately. The purple path embeds the question and tokens generated so far; the answer portion is empty initially. The vision encoder exposes h1, h2, …, hN, a sequence of feature rows. B × N × dV means batch size × retained visual rows per example × features per row; the drawing shows one image. Text states supply Q, image states are projected to K and V. The attention result updates the language states before the vocabulary head. N varies with image resolution, patch size, CLS handling and any compression; 197 is only a conditional example. This schematic shows the standard sequence-to-sequence mechanism, not Flamingo’s exact layer ordering or gates.
Ask: Causal self-attention reads text; cross-attention reads image memory.
Answer: Trace the two inputs separately. The purple path embeds the question and tokens generated so far; the answer portion is empty initially. The vision encoder exposes h1, h2, …, hN, a sequence of feature rows. B × N × dV means batch size × retained visual rows per example × features per row; the drawing shows one image. Text states supply Q, image states are projected to K and V. The attention result updates the language states before the vocabulary head. N varies with image resolution, patch size, CLS handling and any compression; 197 is only a conditional example. This schematic shows the standard sequence-to-sequence mechanism, not Flamingo’s exact layer ordering or gates.
03 · COULD VISION REMAIN SEPARATE?
We already did this with English-to-Hindi translation
English source states stayed available while the Hindi prefix grew.
How does visual information become context for next-token prediction?
Original teaching example / diagram78 / 155
Use the complete source-left, decoder-right architecture from the final diagrams of the earlier lecture. The current target state first uses causal target self-attention and then reads source states. Replace only the source modality on the next slide.
Ask: English source states stayed available while the Hindi prefix grew.
Answer: Use the complete source-left, decoder-right architecture from the final diagrams of the earlier lecture. The current target state first uses causal target self-attention and then reads source states. Replace only the source modality on the next slide.
03 · COULD VISION REMAIN SEPARATE?
Keep the decoder idea; change English to an image
The source changes. The decoder still reads a separate memory.
How does visual information become context for next-token prediction?
Original teaching example / diagram79 / 155
The text sequence remains separate from visual memory. We have not yet computed a visual message.
Ask: The source changes. The decoder still reads a separate memory.
Answer: The text sequence remains separate from visual memory. We have not yet computed a visual message.
03 · COULD VISION REMAIN SEPARATE?
What does the visual lookup need?
| Part | Role in the familiar attention operation |
|---|---|
| Query: from the current language state | What information would help this position? |
| Keys: from the visual states | Compute a match score for each memory row. |
| Values: from the visual states | Supply the features that the weights will mix. |
Q decides how to read; K is compared; V supplies the returned content.
How does visual information become context for next-token prediction?
Original teaching example / diagram80 / 155
These phrases explain data flow, not interpretable concepts encoded by every learned coordinate. A query is a vector derived from the current text computation, not the entire question string. We now derive these three pieces on a zoomed-in view.
Ask: If we swap the image but keep this incoming language state fixed, what changes?
Answer: The visual keys and values change. The incoming query is fixed for this isolated lookup; in the full model later language states can also change.
03 · COULD VISION REMAIN SEPARATE?
The language state asks the question
Queries come from the language state.
How does visual information become context for next-token prediction?
Original teaching example / diagram81 / 155
H_L is the state entering this cross-attention sublayer, after causal self-attention and any normalization. Cross-attention has its own projection weights.
Ask: Queries come from the language state.
Answer: H_L is the state entering this cross-attention sublayer, after causal self-attention and any normalization. Cross-attention has its own projection weights.
03 · COULD VISION REMAIN SEPARATE?
The image supplies the keys
Keys describe what the visual memory can match.
How does visual information become context for next-token prediction?
Original teaching example / diagram82 / 155
Q and K share the per-head width d_k; their sequence lengths need not match.
Ask: Keys describe what the visual memory can match.
Answer: Q and K share the per-head width d_k; their sequence lengths need not match.
03 · COULD VISION REMAIN SEPARATE?
The image also supplies the values
Values provide the content that attention will mix.
How does visual information become context for next-token prediction?
Original teaching example / diagram83 / 155
W_K and W_V are separate learned matrices. Per-input K and V are computed activations.
Ask: Values provide the content that attention will mix.
Answer: W_K and W_V are separate learned matrices. Per-input K and V are computed activations.
03 · COULD VISION REMAIN SEPARATE?
Compare a query with every visual key
The result is one score per visual state.
How does visual information become context for next-token prediction?
Original teaching example / diagram84 / 155
Scaling controls dot-product magnitude before normalization. One attention head shown.
Ask: The result is one score per visual state.
Answer: Scaling controls dot-product magnitude before normalization. One attention head shown.
03 · COULD VISION REMAIN SEPARATE?
Turn the scores into mixing weights
Softmax normalizes across the visual memory positions.
How does visual information become context for next-token prediction?
Original teaching example / diagram85 / 155
These are attention weights, not probabilities of answer words.
Ask: Softmax normalizes across the visual memory positions.
Answer: These are attention weights, not probabilities of answer words.
03 · COULD VISION REMAIN SEPARATE?
Read a visual message
One query produces one context vector per head.
How does visual information become context for next-token prediction?
Original teaching example / diagram86 / 155
Heads are concatenated and projected in multi-head attention; the result returns to the language computation.
Ask: One query produces one context vector per head.
Answer: Heads are concatenated and projected in multi-head attention; the result returns to the language computation.
03 · COULD VISION REMAIN SEPARATE?
One tiny lookup: from a language query to a visual message
| Teaching label | Key | Value |
|---|---|---|
| face | [0.9, 0.1] | [1, 0] |
| fur | [0.6, 0.4] | [0, 2] |
| background | [0.1, 0.9] | [1, 1] |
Toy numbers and labels; no measured attention map is being shown.
How does visual information become context for next-token prediction?
Original teaching example / diagram87 / 155
Return to the visual-memory architecture. We are inside its teal cross-attention box, computing one message for one language position. The three row labels and all numbers are hand-set, not measured regions from the dog image. A large attention weight does not by itself prove grounding.
Ask: Toy numbers and labels; no measured attention map is being shown.
Answer: Return to the visual-memory architecture. We are inside its teal cross-attention box, computing one message for one language position. The three row labels and all numbers are hand-set, not measured regions from the dog image. A large attention weight does not by itself prove grounding.
03 · COULD VISION REMAIN SEPARATE?
Query × face key
The same language query is compared with another visual key.
How does visual information become context for next-token prediction?
Original teaching example / diagram88 / 155
Unscaled toy dot product. Earlier results remain on the board.
Ask: The same language query is compared with another visual key.
Answer: Unscaled toy dot product. Earlier results remain on the board.
03 · COULD VISION REMAIN SEPARATE?
Query × fur key
The same language query is compared with another visual key.
How does visual information become context for next-token prediction?
Original teaching example / diagram89 / 155
Unscaled toy dot product. Earlier results remain on the board.
Ask: The same language query is compared with another visual key.
Answer: Unscaled toy dot product. Earlier results remain on the board.
03 · COULD VISION REMAIN SEPARATE?
Query × background key
The same language query is compared with another visual key.
How does visual information become context for next-token prediction?
Original teaching example / diagram90 / 155
Unscaled toy dot product. Earlier results remain on the board.
Ask: The same language query is compared with another visual key.
Answer: Unscaled toy dot product. Earlier results remain on the board.
03 · COULD VISION REMAIN SEPARATE?
Scale, then normalize this row
The weights were calculated from Q and K.
How does visual information become context for next-token prediction?
Original teaching example / diagram91 / 155
Full precision: 0.4207287, 0.34030973, 0.23896158. They sum to one.
Ask: The weights were calculated from Q and K.
Answer: Full precision: 0.4207287, 0.34030973, 0.23896158. They sum to one.
03 · COULD VISION REMAIN SEPARATE?
Use those weights to mix the values
The result is a vector, not an output word.
How does visual information become context for next-token prediction?
Original teaching example / diagram92 / 155
Rounded display; the executable arithmetic check uses full precision. Values need not equal keys.
Ask: The result is a vector, not an output word.
Answer: Rounded display; the executable arithmetic check uses full precision. Values need not equal keys.
03 · COULD VISION REMAIN SEPARATE?
The message updates the language state; it is not a word
The projected attention output updates the language state.
How does visual information become context for next-token prediction?
Original teaching example / diagram93 / 155
Schematic residual update. Normalization, multi-head output projection and model-specific gates are omitted from the drawing; they do not turn the context vector directly into a word.
Ask: The projected attention output updates the language state.
Answer: Schematic residual update. Normalization, multi-head output projection and model-specific gates are omitted from the drawing; they do not turn the context vector directly into a word.
03 · COULD VISION REMAIN SEPARATE?
Check: has cross-attention already answered “dog”?
Lookup result
What happens next?
Does c name a word?
Or does it update the decoder’s current state?
Predict before advancing: where does a vocabulary token actually appear?
How does visual information become context for next-token prediction?
Original teaching example / diagram94 / 155
The numerical vector is not a word ID. Ask students to point to the vocabulary head in the whole architecture.
Ask: Does the cross-attention output directly name the answer token?
Answer: No. It updates a language hidden state. The vocabulary head later scores output token IDs.
03 · COULD VISION REMAIN SEPARATE?
The vocabulary head still produces the next-token scores
Visual lookup changes the language state; the language head chooses words.
How does visual information become context for next-token prediction?
Original teaching example / diagram95 / 155
Trace memory → cross-attention → updated decoder computation → vocabulary head. Cross-attention returns features, not an English answer. Model-specific blocks can repeat this lookup in depth.
Ask: Visual lookup changes the language state; the language head chooses words.
Answer: Trace memory → cross-attention → updated decoder computation → vocabulary head. Cross-attention returns features, not an English answer. Model-specific blocks can repeat this lookup in depth.
03 · COULD VISION REMAIN SEPARATE?
Two ways for language to access vision
In-sequence visual context or a separate visual memory.
How does visual information become context for next-token prediction?
Original teaching example / diagram96 / 155
Same image and prompt. Neither access pattern is universally better; training, model size, compute and task matter.
Ask: In-sequence visual context or a separate visual memory.
Answer: Same image and prompt. Neither access pattern is universally better; training, model size, compute and task matter.
03 · COULD VISION REMAIN SEPARATE?
Idea 2, end to end in one function
image
One image
question = “What animal is shown?”
answer_so_far = “It is a” (empty at the start)
def next_token(image, question, answer_so_far): memory = vision_encoder(image).featuresN × dV ids = tokenize(question, answer_so_far)T token IDs text = embed(ids)T × dLM for block in decoder_blocks: text = block.causal_self_attn(text)T × dLM text = block.cross_attn(text, memory)T × dLM text = block.mlp(text)T × dLM return argmax(lm_head(text[-1]))Next text token IDcross_attn: Q from text; K and V from image memory. The two sequences stay separate.
Sublayer helpers include residuals and norms. Append the returned token ID to the answer; repeat.
How does visual information become context for next-token prediction?
Transformer · Vaswani et al., 2017 · Oxford-IIIT Pet · CC BY-SA 4.097 / 155
A conceptual encoder-decoder computation for one image and one next-token prediction, batch omitted. vision_encoder includes image preprocessing and produces all N feature rows. For the ViT teaching example these are contextual patch states, with full image attention. The tokenizer helper formats question and answer prefix, as on the Idea 1 slide; T includes all supplied text positions. Each causal_self_attn helper returns updated text states using causal self-attention, residual addition and normalization; cross_attn and mlp also include their residual and normalization steps. Positional handling and the final stack normalization are abstracted. The cross-attention helper has separate learned projections: Q comes from text and K/V from image memory. Per-head Q and K widths must match, but dV and dLM need not match; the internal projections and output projection handle this. No visual/text concatenation occurs. The decoder text stream keeps T rows, and the final text position feeds the vocabulary head. Image memory is constant within this prediction. This is the familiar standard sequence-to-sequence ordering, not an exact implementation of Flamingo. Flamingo adds gated visual cross-attention at selected points among pretrained language blocks and uses a resampled visual memory. The later Flamingo paper slide maps those differences explicitly. For efficient generation, encode the image once and cache reusable attention keys and values. argmax is greedy decoding; display via the tokenizer and append the ID, stopping at EOS or the generation limit. All APIs and the answer prefix are illustrative.
Ask: Where do Q, K and V come from, and how many text states leave each block?
Answer: Cross-attention gets Q from text states and K/V from image memory. It updates T text states; the N image rows remain separate. The last text state supplies the next-token scores.
04 · DO WE NEED EVERY VISUAL TOKEN?
DO WE NEED EVERY VISUAL TOKEN?
04
Return to Idea 3: how can we make a shorter visual summary?
How does visual information become context for next-token prediction?
Original teaching example / diagram98 / 155
Return to Idea 3: how can we make a shorter visual summary?
Ask:
Answer: Return to Idea 3: how can we make a shorter visual summary?
04 · DO WE NEED EVERY VISUAL TOKEN?
Idea 3 in detail · make the visual context shorter
Replace the row-wise projector with a learned summary followed by projection.
How does visual information become context for next-token prediction?
Original teaching example / diagram99 / 155
Compare with the first complete prefix architecture. Only the left-hand connector changes. The decoder still consumes visual vectors followed by text. The same compression idea can instead provide the separate memory used by cross-attention.
Ask: Replace the row-wise projector with a learned summary followed by projection.
Answer: Compare with the first complete prefix architecture. Only the left-hand connector changes. The decoder still consumes visual vectors followed by text. The same compression idea can instead provide the separate memory used by cross-attention.
04 · DO WE NEED EVERY VISUAL TOKEN?
How long is the combined context?
576 visual states + 64 text positions = 640 positions.
How does visual information become context for next-token prediction?
Original teaching example / diagram100 / 155
Illustrative 24 by 24 patch grid, with no CLS state included. Token counts depend on image resolution and preprocessing.
Ask: 576 visual states + 64 text positions = 640 positions.
Answer: Illustrative 24 by 24 patch grid, with no CLS state included. Token counts depend on image resolution and preprocessing.
04 · DO WE NEED EVERY VISUAL TOKEN?
What if the image became 32 context vectors?
32 visual states + 64 text positions = 96 positions.
How does visual information become context for next-token prediction?
Original teaching example / diagram101 / 155
Both bars use the same scale. Dense attention score counts are 640² versus 96², about 44.4 times fewer pairs, not 44.4 times total speedup.
Ask: 32 visual states + 64 text positions = 96 positions.
Answer: Both bars use the same scale. Dense attention score counts are 640² versus 96², about 44.4 times fewer pairs, not 44.4 times total speedup.
04 · DO WE NEED EVERY VISUAL TOKEN?
Now build the summary: start with the image states
How can we read a fixed-size summary of a variable visual memory?
How does visual information become context for next-token prediction?
Original teaching example / diagram102 / 155
Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
Ask: How can we read a fixed-size summary of a variable visual memory?
Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
04 · DO WE NEED EVERY VISUAL TOKEN?
Use 32 learned query vectors as 32 readers
How can we read a fixed-size summary of a variable visual memory?
How does visual information become context for next-token prediction?
Original teaching example / diagram103 / 155
Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
Ask: How can we read a fixed-size summary of a variable visual memory?
Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
04 · DO WE NEED EVERY VISUAL TOKEN?
Each reader returns one image-dependent summary vector
The query count fixes the number of output vectors.
How does visual information become context for next-token prediction?
Original teaching example / diagram104 / 155
Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
Ask: The query count fixes the number of output vectors.
Answer: Recall the lookup we just computed: one query returns one mixture of values. Here the initial query vectors are learned parameters instead of being derived from the current answer position. 32 query slots therefore produce 32 output vectors. The initial slots persist across images; the returned summaries change with the image. Real resamplers may refine these slots across layers. The picture shows 4 of 32 outputs and 6 of 576 inputs.
04 · DO WE NEED EVERY VISUAL TOKEN?
Which detail survives the bottleneck?
32 outputs must carry the evidence needed by the question.
A shorter sequence can lose small text or spatial detail.
Compression is a separate choice from prefix versus cross-attention access.
How does visual information become context for next-token prediction?
Original teaching example / diagram105 / 155
A resampler can feed either access pattern. Do not equate fewer visual tokens with universally better performance.
Ask: Compression is a separate choice from prefix versus cross-attention access.
Answer: A resampler can feed either access pattern. Do not equate fewer visual tokens with universally better performance.
04 · DO WE NEED EVERY VISUAL TOKEN?
Attach names after the design choices
| Example | Visual states | Connector | Access |
|---|---|---|---|
| LLaVA-style | ViT patches | Linear / MLP | Visual prefix |
| Flamingo | Vision feature grid | Perceiver Resampler | Gated cross-attention |
| BLIP-2 | ViT patches | Q-Former + projection | Query-output prefix* |
*Decoder-only LLM variant. Other language backbones use different interfaces.
Flamingo and BLIP-2 use learned-query mechanisms in different systems.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 2023 · Flamingo · Alayrac et al., 2022 · BLIP-2 · Li et al., 2023106 / 155
This is a design comparison, not a ranking. BLIP-2 supports different language backbones; decoder-only examples use visual prefix conditioning. Training objectives and exact modules differ.
Ask: Flamingo and BLIP-2 use learned-query mechanisms in different systems.
Answer: This is a design comparison, not a ranking. BLIP-2 supports different language backbones; decoder-only examples use visual prefix conditioning. Training objectives and exact modules differ.
PAPER · Flamingo · NeurIPS 2022
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac · Jeff Donahue · Pauline Luc · et al.
Read the paper ↗1 · Idea 3 · compress visual features
Frozen NFNet features → Perceiver Resampler → fixed-size visual memory.
2 · Idea 2 · read separate memory
Gated cross-attention inserts visual information between LM blocks.
3 · Read the freeze symbols
Vision encoder and LM stay frozen; the added bridge modules are trained.
How does visual information become context for next-token prediction?
Flamingo · Alayrac et al., 2022107 / 155
Original Figure 3 from arXiv 2204.14198v1. Trace both visual paths into the language stack. The original vision backbone is an NFNet-F6 convolutional network, not the generic ViT used for our teaching diagram. The Perceiver Resampler compresses the visual features; gated cross-attention keeps them separate from the text stream. The diagram also shows interleaved image/text context for few-shot prediction. Snowflakes indicate frozen pretrained modules. The paper has an image-conditioning mask and other details beyond our generic cross-attention illustration. Its figure colors follow the paper, not the lecture palette.
Ask: Point to the modules that implement the connection we just derived.
Answer: Idea 3 · compress visual features: Frozen NFNet features → Perceiver Resampler → fixed-size visual memory. Idea 2 · read separate memory: Gated cross-attention inserts visual information between LM blocks. Read the freeze symbols: Vision encoder and LM stay frozen; the added bridge modules are trained.
PAPER · BLIP-2 · ICML 2023
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li · Dongxu Li · Silvio Savarese · Steven Hoi
Read the paper ↗1 · Idea 3 · learned-query summaries
32 learned queries produce 32 image-dependent output vectors.
2 · Idea 1 · prefix for OPT (top)
Project the summaries; prepend them to text embeddings.
3 · Bottom: encoder–decoder
Visual vectors + text enter the encoder; its decoder generates the answer.
How does visual information become context for next-token prediction?
BLIP-2 · Li et al., 2023108 / 155
Original Figure 3 from arXiv 2301.12597v3, showing the second generative-pretraining stage. The top OPT route maps to our compressed visual-prefix connection. The lower FlanT5 route conditions the language encoder, so it is not the same decoder-only interface. The fully connected layer changes vector width after Q-Former compression. Query outputs depend on the image, while the learned query parameters are shared. Image encoder and LLM remain frozen in this stage; Q-Former and the connecting projection are trained. The first representation-learning stage and its objectives are described elsewhere in the paper; the simple learned-reader diagram was not a full Q-Former specification.
Ask: Point to the modules that implement the connection we just derived.
Answer: Idea 3 · learned-query summaries: 32 learned queries produce 32 image-dependent output vectors. Idea 1 · prefix for OPT (top): Project the summaries; prepend them to text embeddings. Bottom: encoder–decoder: Visual vectors + text enter the encoder; its decoder generates the answer.
04 · DO WE NEED EVERY VISUAL TOKEN?
The learned readers stay; their image summaries change
Parameters persist across examples; activations depend on the input.
How does visual information become context for next-token prediction?
Original teaching example / diagram109 / 155
A pretrained but frozen weight is still a learned parameter. The frozen/trainable decision will be a separate overlay during training. Learned query embeddings are parameters; actual projected attention queries are activations.
Ask: Parameters persist across examples; activations depend on the input.
Answer: A pretrained but frozen weight is still a learned parameter. The frozen/trainable decision will be a separate overlay during training. Learned query embeddings are parameters; actual projected attention queries are activations.
04 · DO WE NEED EVERY VISUAL TOKEN?
Idea 3, end to end in one function
image
One image
question = “What animal is shown?”
answer_so_far = “It is a” (empty at the start)
def next_token(image, question, answer_so_far): patches = vit(image).patch_statesN × dV summary = summarizer(queries, patches)M × dS; M < N visual = projector(summary)M × dLM ids = tokenize(question, answer_so_far)T token IDs text = embed(ids)T × dLM inputs = concat([visual, text], axis="tokens")(M + T) × dLM hidden = causal_decoder(inputs)(M + T) × dLM return argmax(lm_head(hidden[-1]))Next text token IDsummarizer: M learned queries read all N patch rows through cross-attention; outputs depend on the image.
Shown: BLIP-2 / OPT-style prefix. For a Flamingo-style connection, use the summaries as Idea 2’s memory.
How does visual information become context for next-token prediction?
BLIP-2 · Li et al., 2023 · Flamingo · Alayrac et al., 2022 · Oxford-IIIT Pet · CC BY-SA 4.0110 / 155
This conceptual function demonstrates the compressed-prefix route, corresponding to the top OPT-style path in the BLIP-2 architecture figure. The image enters a ViT and all N contextual patch features enter the summarizer. queries is an M-row learned parameter bank shared across images, not an input text question. The summarizer forms attention queries from this bank and its evolving states; patch features supply image keys and values. Its M output rows depend on the image and are mixtures, not M selected original patches. dS is the summarizer output width. The summarizer can contain repeated self-attention, cross-attention and MLP blocks, with residuals and norms; this one call is not a full specification of Q-Former or Perceiver Resampler. BLIP-2 uses 32 learned queries. M < N expresses the compression example rather than a universal constraint at every possible resolution. The projector changes width dS to dLM, preserving M. Then the same tokenize, embed, concat, causal decoder and language head as Idea 1 produce the next answer token. For the other connection, keep summary as a separate memory and pass it to the Idea 2 cross-attention decoder with the T text embeddings; do not also concatenate. Flamingo exemplifies compression plus separate memory, with NFNet visual features rather than this ViT and a gated cross-attention architecture. Its memory need not be externally projected to dLM; cross-attention has its own K/V and output projections. BLIP-2 also has an encoder-decoder FlanT5 route, shown in the paper figure but not implemented by this decoder-only sketch. The input formatting, positional handling, masking, full decoder stack, final normalization and generation caching conventions are the same conceptual abstractions as Idea 1. The summarizer queries are learned parameters; its summaries are newly computed activations. Freeze/train choices are taught in the training section.
Ask: Which line changes the number of visual rows, and which line changes their width?
Answer: summarizer changes N patch rows into M image-dependent summaries. projector changes each summary’s width to dLM while retaining M. This sketch uses a prefix; the same summaries could instead become separate memory.
05 · FOLLOW ONE ANSWER TOKEN END TO END
FOLLOW ONE ANSWER TOKEN END TO END
05
Now follow the connection we built all the way to a word.
How does visual information become context for next-token prediction?
Original teaching example / diagram111 / 155
Now follow the connection we built all the way to a word.
Ask:
Answer: Now follow the connection we built all the way to a word.
05 · FOLLOW ONE ANSWER TOKEN END TO END
The same Newfoundland image enters the encoder
The encoder computes the visual state matrix.
How does visual information become context for next-token prediction?
Oxford-IIIT Pet · CC BY-SA 4.0112 / 155
Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
Ask: The encoder computes the visual state matrix.
Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
05 · FOLLOW ONE ANSWER TOKEN END TO END
The connector exposes usable visual context
The visual context is available to the language computation.
How does visual information become context for next-token prediction?
Oxford-IIIT Pet · CC BY-SA 4.0113 / 155
Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
Ask: The visual context is available to the language computation.
Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Combine image, question and answer prefix
The current language state will drive the vocabulary head.
How does visual information become context for next-token prediction?
Oxford-IIIT Pet · CC BY-SA 4.0114 / 155
Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
Ask: The current language state will drive the vocabulary head.
Answer: Keep the canonical layout fixed and fade already-introduced stages. The photo and question are unchanged.
05 · FOLLOW ONE ANSWER TOKEN END TO END
The LM head scores vocabulary entries
| Vocabulary entry | Logit |
|---|---|
| dog | ℓdog |
| cat | ℓcat |
| horse / … | ℓhorse/… |
Visual evidence can change hₜ and therefore every vocabulary score.
How does visual information become context for next-token prediction?
Original teaching example / diagram115 / 155
Scores are symbolic; the browser API does not expose reliable per-token logits here. No numerical probabilities are invented.
Ask: Visual evidence can change hₜ and therefore every vocabulary score.
Answer: Scores are symbolic; the browser API does not expose reliable per-token logits here. No numerical probabilities are invented.
05 · FOLLOW ONE ANSWER TOKEN END TO END
From vocabulary scores to a selected token
Selection can be greedy or sampled; the output is a vocabulary token.
How does visual information become context for next-token prediction?
Original teaching example / diagram116 / 155
“dog” is an illustrative selection explaining the computation, not a claim that the tiny model chose it. The measured counterfactual follows. Real tokenizers may split displayed words.
Ask: Selection can be greedy or sampled; the output is a vocabulary token.
Answer: “dog” is an illustrative selection explaining the computation, not a claim that the tiny model chose it. The measured counterfactual follows. Real tokenizers may split displayed words.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Start with image + prompt
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram117 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Append the first token
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram118 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Use the growing prefix
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram119 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Append the object word
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram120 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Finish the sentence
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram121 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Stop at the end token
Fixed prompt
Describe the animal.
Generated prefix
Illustrative trace: the image and prompt remain fixed.
How does visual information become context for next-token prediction?
Original teaching example / diagram122 / 155
The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
Ask: Illustrative trace: the image and prompt remain fixed.
Answer: The same photo remains fixed. Reuse cached context and feed the selected token to the next step. These are word-level teaching steps, not measured tokenizer boundaries or probabilities.
05 · FOLLOW ONE ANSWER TOKEN END TO END
Keep the prompt fixed. Change only the image.
What animal is shown?
Dog
There is a black dog in the foreground.
Cat
A cat is visible.
Blank
A squirrel.
Actual saved answers from the pinned browser model; probabilities are not exposed.
How does visual information become context for next-token prediction?
SmolVLM-256M · model card123 / 155
Blank means a white image passed through the same image processor, not omission of the image input. Responses test sensitivity to image changes; they do not by themselves establish reliable grounding. Full exact records and metadata are in results/vlm_lab_outputs.json.
Ask: Actual saved answers from the pinned browser model; probabilities are not exposed.
Answer: Blank means a white image passed through the same image processor, not omission of the image input. Responses test sensitivity to image changes; they do not by themselves establish reliable grounding. Full exact records and metadata are in results/vlm_lab_outputs.json.
06 · HOW DO WE TRAIN THE CONNECTION?
HOW DO WE TRAIN THE CONNECTION?
06
Make the visual connection useful by training on reference answers.
How does visual information become context for next-token prediction?
Original teaching example / diagram124 / 155
Make the visual connection useful by training on reference answers.
Ask:
Answer: Make the visual connection useful by training on reference answers.
06 · HOW DO WE TRAIN THE CONNECTION?
The wiring fits. How does it learn to use the image?
Use the correct answer tokens to train the visual-to-language connection.
How does visual information become context for next-token prediction?
Original teaching example / diagram125 / 155
Shape matching made the forward computation possible. It did not teach what the image features mean to language. Use image/answer pairs, next-token loss and gradients. The next slides identify which parameters receive updates in example training recipes.
Ask: Use the correct answer tokens to train the visual-to-language connection.
Answer: Shape matching made the forward computation possible. It did not teach what the image features mean to language. Use image/answer pairs, next-token loss and gradients. The next slides identify which parameters receive updates in example training recipes.
06 · HOW DO WE TRAIN THE CONNECTION?
Start from two pretrained components
The connector is new; vision and language already have learned weights.
How does visual information become context for next-token prediction?
Original teaching example / diagram126 / 155
The snowflake denotes frozen parameters; the flame in later slides denotes parameters being optimized.
Ask: The connector is new; vision and language already have learned weights.
Answer: The snowflake denotes frozen parameters; the flame in later slides denotes parameters being optimized.
06 · HOW DO WE TRAIN THE CONNECTION?
Stage 1 · align visual features with language
Example recipe: caption alignment; freezing choices depend on the model.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 2023127 / 155
Simplified LLaVA-style recipe. Frozen vision encoder and frozen LLM; connector trainable. Exact objectives and freeze choices are architecture-dependent.
Ask: Example recipe: caption alignment; freezing choices depend on the model.
Answer: Simplified LLaVA-style recipe. Frozen vision encoder and frozen LLM; connector trainable. Exact objectives and freeze choices are architecture-dependent.
06 · HOW DO WE TRAIN THE CONNECTION?
Supply the correct previous caption tokens
Reference caption
Teacher forcing supplies the reference prefix during training.
How does visual information become context for next-token prediction?
Original teaching example / diagram128 / 155
Targets and logits are shifted: each state predicts the next target. With a causal mask all known reference positions can be evaluated in parallel.
Ask: Teacher forcing supplies the reference prefix during training.
Answer: Targets and logits are shifted: each state predicts the next target. With a causal mask all known reference positions can be evaluated in parallel.
06 · HOW DO WE TRAIN THE CONNECTION?
Stage 2 · learn to follow visual instructions
Example recipe: image + question + answer; train the LLM or its adapters.
How does visual information become context for next-token prediction?
LLaVA · Liu et al., 2023129 / 155
Example recipe: frozen vision, trainable connector and LLM, or LLM adapters. Some models also unfreeze vision. These choices are model-dependent.
Ask: Example recipe: image + question + answer; train the LLM or its adapters.
Answer: Example recipe: frozen vision, trainable connector and LLM, or LLM adapters. Some models also unfreeze vision. These choices are model-dependent.
06 · HOW DO WE TRAIN THE CONNECTION?
The same image can support different questions
USER: What colour is the dog?
ASSISTANT: Black.
The loss supervises the answer targets selected by the training recipe.
How does visual information become context for next-token prediction?
Original teaching example / diagram130 / 155
A denotes selected assistant answer targets. This answer is an authored training example. Dataset coverage matters: captions alone may not teach counting, OCR or justified rejection.
Ask: The loss supervises the answer targets selected by the training recipe.
Answer: A denotes selected assistant answer targets. This answer is an authored training example. Dataset coverage matters: captions alone may not teach counting, OCR or justified rejection.
06 · HOW DO WE TRAIN THE CONNECTION?
Mark the target at every position
Only Black, period and EOS are targets; preceding-position logits predict them.
How does visual information become context for next-token prediction?
Original teaching example / diagram131 / 155
Every displayed token has a target/no-target marker. The loss uses logits at the preceding positions to predict these targets; a prompt-position logit can predict the first answer. Masked context still influences answer gradients.
Ask: Only Black, period and EOS are targets; preceding-position logits predict them.
Answer: Every displayed token has a target/no-target marker. The loss uses logits at the preceding positions to predict these targets; a prompt-position logit can predict the first answer. Masked context still influences answer gradients.
06 · HOW DO WE TRAIN THE CONNECTION?
Two masks answer different questions
Attention mask
↑ available to the current query
Which earlier states may this query read?
Loss mask
Which next-token targets contribute to the loss?
Removing a target loss does not hide a context token.
How does visual information become context for next-token prediction?
Original teaching example / diagram132 / 155
The attention mask controls information flow; the loss mask selects supervised target positions. They are not interchangeable.
Ask: Removing a target loss does not hide a context token.
Answer: The attention mask controls information flow; the loss mask selects supervised target positions. They are not interchangeable.
06 · HOW DO WE TRAIN THE CONNECTION?
A frozen LLM still passes the gradient
Frozen ≠ stop-gradient.
How does visual information become context for next-token prediction?
Original teaching example / diagram133 / 155
Differentiate the answer loss through the frozen language operations with respect to their input, so connector weights receive a gradient. The LLM weights are not updated. Wrapping that LLM forward in no_grad would sever the required path.
Ask: Frozen ≠ stop-gradient.
Answer: Differentiate the answer loss through the frozen language operations with respect to their input, so connector weights receive a gradient. The LLM weights are not updated. Wrapping that LLM forward in no_grad would sever the required path.
07 · WHAT HAPPENS AT INFERENCE?
WHAT HAPPENS AT INFERENCE?
07
At inference: prepare the image once, then grow the answer.
How does visual information become context for next-token prediction?
Original teaching example / diagram134 / 155
At inference: prepare the image once, then grow the answer.
Ask:
Answer: At inference: prepare the image once, then grow the answer.
07 · WHAT HAPPENS AT INFERENCE?
The whole generation loop before its individual steps
Prepare the fixed context once; repeatedly predict and append a token.
How does visual information become context for next-token prediction?
Original teaching example / diagram135 / 155
This is the same autoregressive loop from the earlier language-model lecture, with one new preparation step for the image. The next frames walk through each step in this complete loop. A cache stores computed keys and values; it does not store the future answer.
Ask: Prepare the fixed context once; repeatedly predict and append a token.
Answer: This is the same autoregressive loop from the earlier language-model lecture, with one new preparation step for the image. The next frames walk through each step in this complete loop. A cache stores computed keys and values; it does not store the future answer.
07 · WHAT HAPPENS AT INFERENCE?
Encode the fixed image once
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram136 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Prefill the image and question context
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram137 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Decode one new token
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram138 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Append the selected token
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram139 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Reuse the cached keys and values
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram140 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Repeat until EOS or the generation limit
Only the generated prefix changes at each decoding step.
How does visual information become context for next-token prediction?
Original teaching example / diagram141 / 155
This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
Ask: Only the generated prefix changes at each decoding step.
Answer: This is one generation call with fixed preprocessing and weights. The generate API handles the loop. Prefix models cache visual positions; cross-attention models may also cache per-layer visual K,V. Caching avoids recomputing states, not all attention work.
07 · WHAT HAPPENS AT INFERENCE?
Same conditional model, different available prefix
Training
Reference prefix supplied.
Target loss and gradients.
Inference
Generated prefix grows.
No reference answer or backward pass.
Known targets allow parallel teacher-forced positions; generation is sequential.
How does visual information become context for next-token prediction?
Original teaching example / diagram142 / 155
The asterisk marks reference tokens. At inference each selected output becomes part of the next step’s input.
Ask: Known targets allow parallel teacher-forced positions; generation is sequential.
Answer: The asterisk marks reference tokens. At inference each selected output becomes part of the next step’s input.
07 · WHAT HAPPENS AT INFERENCE?
Reuse the image; keep the conversation context
User: What animal is this?
Assistant: A dog.
User: What colour is its fur?
Assistant: Black.
Illustrative conversation, not a captured model transcript.
Visual features can persist. Decoder cache reuse also depends on the exact prefix.
How does visual information become context for next-token prediction?
Oxford-IIIT Pet · CC BY-SA 4.0143 / 155
69–71 min. Reprocessing the full conversation is valid but less efficient. Image feature caching and decoder KV caching are distinct. The included classroom lab evaluates independent questions; it does not claim multi-turn KV reuse.
Ask: Visual features can persist. Decoder cache reuse also depends on the exact prefix.
Answer: 69–71 min. Reprocessing the full conversation is valid but less efficient. Image feature caching and decoder KV caching are distinct. The included classroom lab evaluates independent questions; it does not claim multi-turn KV reuse.
08 · CAN WE TRUST THE ANSWER?
CAN WE TRUST THE ANSWER?
08
Check evidence use, not just fluent language.
How does visual information become context for next-token prediction?
Original teaching example / diagram144 / 155
Check evidence use, not just fluent language.
Ask:
Answer: Check evidence use, not just fluent language.
08 · CAN WE TRUST THE ANSWER?
A fluent sentence can have no visual support
What color is the bicycle?
Reference: No bicycle is visible.
Recorded baseline: “The bicycle is red.”
Fluent ≠ grounded.
How does visual information become context for next-token prediction?
Original teaching example / diagram145 / 155
The displayed raw answer was measured in the earlier six-case browser run and is preserved in output/verification/live-model-results.json. The rebuilt lab also records this preset with the fixed decoding settings.
Ask: Fluent ≠ grounded.
Answer: The displayed raw answer was measured in the earlier six-case browser run and is preserved in output/verification/live-model-results.json. The rebuilt lab also records this preset with the fixed decoding settings.
08 · CAN WE TRUST THE ANSWER?
Keep a small diagnostic suite
| Image | Fixed question | Reference |
|---|---|---|
| Dog | What animal is this? | Dog |
| Desk | Is the laptop plugged in? | Yes, visible cable connected |
| Shapes | What is immediately left of the mug? | Red ball |
| Invoice | What is the invoice number? | ES667-042 |
| Chart | Which month has the highest value? | March |
| Coffee | What color is the bicycle? | No bicycle visible |
Preserve the raw answer, reveal the reference, then judge support.
How does visual information become context for next-token prediction?
Original teaching example / diagram146 / 155
Dog, desk, spatial, invoice, chart and false-premise questions span distinct failure modes. Mark supported / partial / unsupported in the lab. Six items are a classroom diagnostic, not benchmark performance.
Ask: Preserve the raw answer, reveal the reference, then judge support.
Answer: Dog, desk, spatial, invoice, chart and false-premise questions span distinct failure modes. Mark supported / partial / unsupported in the lab. Six items are a classroom diagnostic, not benchmark performance.
09 · SUMMARY
SUMMARY
09
Represent the image, expose context, predict the next token.
How does visual information become context for next-token prediction?
Original teaching example / diagram147 / 155
Represent the image, expose context, predict the next token.
Ask:
Answer: Represent the image, expose context, predict the next token.
09 · SUMMARY
ViT → CLIP → generative VLM
Visual representation → cross-modal matching → conditioned generation.
How does visual information become context for next-token prediction?
Original teaching example / diagram148 / 155
CLIP remains a vision-language model in the broad sense. The progression distinguishes what each system is trained and used to produce.
Ask: Visual representation → cross-modal matching → conditioned generation.
Answer: CLIP remains a vision-language model in the broad sense. The progression distinguishes what each system is trained and used to produce.
09 · SUMMARY
One computation to remember
Visual information becomes context for next-token prediction.
How does visual information become context for next-token prediction?
Original teaching example / diagram149 / 155
Exit ticket: distinguish width matching, token compression, and access pattern. Then explain how an answer target can train a connector through a frozen language model.
Ask: Visual information becomes context for next-token prediction.
Answer: Exit ticket: distinguish width matching, token compression, and access pattern. Then explain how an answer target can train a connector through a frozen language model.
OPTIONAL · NEXT LECTURE / SOURCES
Different benchmarks probe different capabilities
Choose the evaluation from the capability you need.
How does visual information become context for next-token prediction?
OCRBench · DocVQA · ChartQA · MathVista · MMMU · MMBench · MM-Vet · HallusionBench150 / 155
71–73 min. Representative benchmark families, not current rankings. MMMU covers 30 subjects; MathVista emphasizes visual mathematical reasoning. MMBench and MM-Vet test different broad or integrated capabilities. Multi-turn and multi-image interaction need additional protocols.
Ask: Would high OCR performance establish good spatial reasoning?
Answer: No. The capabilities and failure modes differ.
OPTIONAL · NEXT LECTURE / SOURCES
Inspect the questions before interpreting a score
Read: What is the invoice ID?
Calculate: March minus January?
Reason: What is x?
Reject: What colour is the bicycle?
Authored classroom examples, not copied benchmark items.
A wrong answer may come from perception, reasoning, or accepting a false premise.
How does visual information become context for next-token prediction?
Original teaching example / diagram151 / 155
73–74 min. Ask students to localize the error: wrong extracted number versus wrong subtraction, for example. The benchmark names motivate task families; these four examples are not official items.
Ask: How could you separate a chart-reading error from an arithmetic error?
Answer: First ask for the two values, then ask for their difference.
OPTIONAL · NEXT LECTURE / SOURCES
Use a metric that matches the answer format
| Output / use | Measure | What still needs inspection |
|---|---|---|
| Fixed answer | Accuracy | Ambiguity and answer normalization |
| Document text | Exact / normalized string match; ANLS | OCR errors and layout dependence |
| Open explanation | Human or specified model rubric | Judge bias and unsupported details |
| Grounding coordinates | IoU / pointing accuracy | Object identity and spatial accuracy |
| Deployment | Latency, memory, output tokens/s | Hardware and decoding settings |
Report the evaluation protocol along with the number.
How does visual information become context for next-token prediction?
DocVQA152 / 155
74–75 min. Recall@K is relevant to retrieval systems such as CLIP, not the default metric for generated answers. ANLS is average normalized Levenshtein similarity; DocVQA uses a thresholded normalized edit-similarity convention. Do not treat model grading as ground truth.
Ask: Report the evaluation protocol along with the number.
Answer: 74–75 min. Recall@K is relevant to retrieval systems such as CLIP, not the default metric for generated answers. ANLS is average normalized Levenshtein similarity; DocVQA uses a thresholded normalized edit-similarity convention. Do not treat model grading as ground truth.
OPTIONAL · NEXT LECTURE / SOURCES
Further training can change how the model answers
Behaviour
Instruction following
Preference optimisation
Safer responses
Evidence use
Better grounding data
Hard negative examples
Missing-evidence responses
The basic inference target remains image-conditioned next-token prediction.
How does visual information become context for next-token prediction?
Original teaching example / diagram153 / 155
65–66 min. Optional post-training is a single-slide preview. Do not claim every model uses every method. Return to the original conditional distribution.
Ask: The basic inference target remains image-conditioned next-token prediction.
Answer: 65–66 min. Optional post-training is a single-slide preview. Do not claim every model uses every method. Return to the original conditional distribution.
OPTIONAL · NEXT LECTURE / SOURCES
Sources · architectures and browser implementation
CLIP · Radford et al., 2021
https://arxiv.org/abs/2103.00020
LLaVA · Liu et al., 2023
https://arxiv.org/abs/2304.08485
Flamingo · Alayrac et al., 2022
https://arxiv.org/abs/2204.14198
BLIP-2 · Li et al., 2023
https://arxiv.org/abs/2301.12597
SmolVLM-256M · model card
https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct
Official Transformers.js browser example
https://github.com/huggingface/transformers.js-examples/tree/main/smolvlm-webgpu
Oxford-IIIT Pet · CC BY-SA 4.0
https://www.robots.ox.ac.uk/~vgg/data/pets/
Primary sources; no leaderboard scores are used in this lecture.
How does visual information become context for next-token prediction?
Original teaching example / diagram154 / 155
Source links checked 1 October 2026. Full provenance and limitations are in SOURCES.md. Sources describe named architectures; drawings are original simplifications.
Ask: Primary sources; no leaderboard scores are used in this lecture.
Answer: Source links checked 1 October 2026. Full provenance and limitations are in SOURCES.md. Sources describe named architectures; drawings are original simplifications.
OPTIONAL · NEXT LECTURE / SOURCES
Sources · evaluation
OCRBench
https://arxiv.org/abs/2305.07895
DocVQA
https://arxiv.org/abs/2007.00398
ChartQA
https://arxiv.org/abs/2203.10244
MathVista
https://arxiv.org/abs/2310.02255
MMMU
https://arxiv.org/abs/2311.16502
MMBench
https://arxiv.org/abs/2307.06281
MM-Vet
https://arxiv.org/abs/2308.02490
HallusionBench
https://arxiv.org/abs/2310.14566
The local six-image suite is an instructional diagnostic, not one of these benchmarks.
How does visual information become context for next-token prediction?
Original teaching example / diagram155 / 155
Benchmark examples in the deck are authored analogues and are not official samples. Do not compare the six-case result against published benchmark scores.
Ask: The local six-image suite is an instructional diagnostic, not one of these benchmarks.
Answer: Benchmark examples in the deck are authored analogues and are not official samples. Do not compare the six-case result against published benchmark scores.