The question we want to answer
We can reuse the encoder once we turn the image into tokens.
The previous lecture, Transformers beyond next-token prediction, compared these three model families. The encoder on the left uses full attention to read the complete input. Here we work out how to give it an image.
Main lecture PDF · 55 pages · Searchable transcript · Optional labs and references · Reference PDF · 151 pages
This photograph comes from Oxford-IIIT Pet. The model chooses from 1,000 ImageNet classes.
Our photograph is newfoundland_31 from Oxford-IIIT Pet. We use vit_tiny_patch16_224.augreg_in21k_ft_in1k throughout; it assigns Newfoundland a probability of 0.9572675228. Oxford-IIIT Pet provides the image, while ImageNet defines the classifier’s output labels. The optional extension shows how to adapt the model to dog/cat classification.
The encoder takes a sequence of vectors. We need to turn the pixels into that sequence.
The image is an RGB array arranged in a spatial grid. The encoder takes rows of features. Our first job is to turn that grid into feature rows; Q, K and V come later.
Each question gives us a reason to add the next part of the model.
We will answer each question before introducing the component that solves it. By the end, we will have followed one photograph all the way to a class prediction and seen what a fixed set of output classes leaves out.
How can an image become a token sequence?
Using patches also keeps the number of attention comparisons manageable.
Pixel tokens would give 50,176 spatial positions, each initially holding RGB. The fine grid is schematic. Patch tokens give 196 positions before CLS. Attention score counts scale quadratically with sequence length. This comparison ignores CLS to emphasize the scale, not memory implementation details.
Each non-overlapping crop is 16 × 16 pixels, with all three colour channels.
The photograph has already been resized and centre-cropped using the checkpoint transform. We now work with its 224 by 224 RGB input. Patches are ordered row by row.
A patch has 768 pixel values. We will add its location separately.
Patch P63 lies at zero-based grid row 4, column 6. The crop illustration is magnified. Before projection, use the checkpoint normalization: RGB to [0,1], subtract 0.5 and divide by 0.5 per channel.
Flattening only rearranges the values; it does not learn features.
Read pixels row by row within each channel, then concatenate R, G and B. This is channel-major ordering, matching PyTorch Unfold and flattened Conv2d kernels. Pixel-major ordering is also valid if the weight columns are permuted to match.
The learned projection turns those 768 pixel values into 192 features.
Use row-vector notation: x_i is 1 by 768, W_E is 768 by 192, and b_E has 192 entries. This affine map has no following activation in the patch layer. PyTorch stores the Linear weight transposed, as 192 by 768.
Every patch passes through the same projection with its own pixel values. We learn one set of weights and biases, shared across all 196 patches. We still need a way to tell the model where each patch came from.
The projection changes 768 features to 192. We still have 196 patch rows.
X contains all flattened patches. C = X W_E + b_E broadcasts the same bias across every row. The batch axis is omitted in the main lecture.
These are the first three of P63’s 192 features.
The first three values come from real-patch-path.json for our saved checkpoint. The dots stand for the other 189 coordinates. These features describe the patch before we add position. We have not assigned a meaning to each individual coordinate.
The image is now 196 rows, each with 192 learned features.
This diagram tracks our progress through the model. Completed stages fade, and the current stage is larger. We have patch features now; next we need to add where each patch came from.
How does the model know where a patch came from?
The content vector describes the crop. A position vector tells the model where it came from.
P63 is in row 5, column 7 when we count from one. The shared projection has no explicit position index. The picture itself may offer clues about location, but the learned position embedding gives the model that information explicitly.
Adding position preserves the 196 × 192 shape.
Content vectors are blue and learned positional vectors amber. We add corresponding coordinates, not concatenate them. The vectors remain 192 features wide. The CLS position is added when CLS is included.
Every image with this grid uses the same learned position table.
Training updates the position table along with the rest of the model. At inference, the table stays fixed. Our checkpoint has 197 position vectors, including one for the CLS slot.
We have 196 rows. Which one should we read?
Reuse the CLS readout we already know from text.
Averaging final patch features is a valid alternative when the model is trained for that readout. We keep the actual checkpoint’s CLS design, rather than substituting a different architecture.
We use patch tokens where the text model used word tokens.
As in the text model, the starting CLS vector is a learned parameter. It is shown in amber. After the encoder, CLS depends on the input; its colour matches the text or image branch.
Both photographs start with exactly the same CLS vector.
Both photographs use the same trained model, starting CLS vector and CLS position vector. Only their patch inputs differ. The cat helps us see how the input changes the final representation; we are still using the original classifier.
The same model produces a different CLS representation for each photograph.
The arrows show the computation. The dog and cat produce different final CLS states even though they use the same model. The reference deck includes measured traces for both images; this diagram does not display numerical predictions.
CLS and patch rows can exchange information in both directions.
Each block computes all updates from its incoming states, in parallel. The next block reads the updated CLS and patch states. Arrows show allowed dependencies, including patch-to-CLS and CLS-to-patch, not guaranteed large weights.
196 projected patches plus one CLS row, each with 192 features.
Prepend CLS, then add a position vector to each of the 197 rows. This gives us the complete image sequence. We can now use that sequence to compute Q, K and V.
How do the patches exchange information?
We can now apply the same encoder operations we used for text.
Normalization, self-attention, residual additions and the MLP all work on feature rows. The small pictures in our diagrams identify patches; the encoder itself receives their vectors.
ViT gets Q, K and V from one image sequence. Translation cross-attention gets them from two sequences.
In ViT, three learned projections read the same normalized image sequence. In translation cross-attention, queries come from target decoder states and keys/values from source encoder states. Translation also contains self-attention; the comparison concerns its cross-attention operation.
Every token makes a query, a key and a value through three learned projections.
Recall the same three roles from text. In this head q_i = LN(e_i) W_Q, k_i = LN(e_i) W_K and v_i = LN(e_i) W_V (with the checkpoint biases). Each is 64 features wide. These are learned numeric vectors; the questions are an analogy for their roles. P74 can receive information using its query and send information using its key and value.
Each query can gather different information from the same image.
These questions illustrate what a query might do; they are not translations of the trained vectors. A key has features that can match a query. The value carries information from the same source through a separate projection. An ear need not attend to another ear, and a head has no fixed semantic role. We will follow the calculation and then inspect a measured attention map.
The weight for P60 scales all 64 coordinates of P60’s value vector.
The pictured source provides both k_60 and v_60, using different projections of its current normalized state. q_74 dot k_60 divided by 8 is one score. Its attention weight comes from softmax against every source score in row P74, including CLS. The weight is a scalar; the value and weighted contribution are 64-feature vectors. This diagram illustrates the roles of Q, K and V; it does not show measured attention weights.
Multiply each value vector by its weight, then add the results to get one message.
These chosen toy weights sum to one; they are not from the dog checkpoint. Reveal the second source, third source, then their sum: 0.6[2,0] + 0.3[0,1] + 0.1[1,1] = [1.3,0.4]. The real head sums 197 source contributions with 64 features each. It produces one message per receiving query. The multihead output projection converts the joined message into the update added to that receiver’s original row.
Each row is a query; each column is a key. Every pair is allowed, but the weights can differ.
The small matrices show permitted query-key pairs, not learned weights. In a next-token decoder, future source positions are masked; a ViT encoder receives the complete image, so all 197 by 197 pairs are allowed, including CLS and self-pairs. Actual scores and row-softmax weights are generally asymmetric because Q and K differ. Patch order in the sequence does not impose a temporal or reading direction.
We used the CLS query, so this message updates CLS.
In one head, all 197 weighted value rows contribute to a 64-feature message. We join the heads, project the result to 192 features, and add it to the incoming CLS embedding. The optional attention appendix works through each source’s contribution.
Join each row’s three 64-feature messages, then project them back to 192 features.
The heads run in parallel on the same normalized input, each with its own learned projections. They can gather different messages without having fixed roles such as an "ear head". Joining them increases the feature width while keeping the same token rows.
Choose one query and look at its source weights. Another head or block may gather a different mixture.
The maps use saved Q/K from trained ViT-Tiny block 4, head 1, query P74. P60 receives 9.088265% and CLS receives 2.089876%. All 197 weights sum to one; the displayed 196 patch weights are not renormalized. Teal intensity is scaled to this map’s maximum. The violet outline marks the query and white marks source P60. An attention map visualizes mixing, not a segmentation or a complete explanation of the class prediction.
Gather a message, project it to 192 features, then add it to the receiving token’s original embedding.
Follow P74 as one receiver. Its attention heads gather weighted values from the current image sequence. Concatenate the three 64-feature messages and apply the learned output projection to obtain a 192-feature update. Add that update to P74’s incoming embedding; the message alone is not the new embedding. All other patch rows and CLS receive their own updates in parallel. Normalization is omitted from this conceptual diagram; the actual checkpoint uses U = E + MSA(LN(E)). The complete pre-LN block remains in the optional reference deck.
Attention gathers context from other rows; the MLP then processes each row’s features.
The MLP receives the attention-updated representation and computes another 192-feature update. Add it to that same attention-updated representation, not to the original input from before attention. The same MLP parameters are used independently for every patch and CLS. This completes one block. Normalization and the internal 192 → 768 → 192 layers with GELU are omitted from this conceptual figure; the exact formula is E_next = U + MLP(LN(U)). Implementation details remain in the reference deck.
The next block starts from these new embeddings. After block 12, we read CLS.
All 197 rows continue through the stack: 196 patch embeddings and one CLS embedding. Each block reads the states produced by the preceding block and performs attention and MLP updates. Blocks have their own learned weights. The feature width remains 192; what each row represents changes. After the last block, the checkpoint’s final normalization and CLS readout give one image representation for classification.
Training learns W_Q, W_K and W_V. Each new image gets its own attention weights.
At inference, learned parameters stay fixed and the model recomputes its activations for each image. Training uses gradients to update the parameters. The query and key projection weights are learned parameters; the attention matrix is computed from the current input.
What does the model predict, and what changes its answer?
The classifier reads the final CLS embedding as its image summary.
After the final LayerNorm, we select the CLS row to get a 192-feature image embedding. Selecting the row does not update it. This diagram leaves out normalization so we can focus on the readout. All 196 final patch embeddings still exist, but the trained ImageNet head uses CLS. Next, the head turns those features into 1,000 class scores.
The class head learns how each image feature contributes to each class score.
The head applies a learned affine map. Softmax turns its class scores into probabilities across the output classes. Attention softmax instead normalizes weights across source tokens. We predict the class with the highest score.
The model assigns Newfoundland a probability of 95.73% among its 1,000 classes.
These are saved measured predictions from real-inference.json: Newfoundland 95.726752%, Tibetan mastiff 1.625724%, briard 0.674269%. They are not performance or accuracy estimates. The remaining probability belongs to the other ImageNet classes.
What do you expect to change when we pass the covered image through the same model?
Keep Newfoundland as the target class throughout this experiment. Compare the original photograph with a fresh copy whose top-left 112 by 112 pixels are replaced with mid-gray. We change the pixels, keeping all tokens and the attention mask unchanged. Here we measure the final prediction; the earlier attention map showed weights inside the model.
The model keeps its learned weights and recomputes patch features, attention weights and class probabilities.
The checkpoint transform normalizes each channel as (RGB − 0.5) / 0.5, so normalized zero means gray RGB 0.5, not black. Index 256 is Newfoundland. Both calls use the same frozen model in evaluation mode. Image shape, all 196 patch tokens and CLS remain; 49 patch inputs now contain gray pixels. Each test uses a fresh clone, so covers never accumulate.
Covering the top right gives the biggest drop: 16.11 percentage points. All four covered images still predict Newfoundland.
Choose the four quadrants before viewing results. Run each covered copy separately and measure the same target-class probability. Top-right drops most among these four tests. This shows sensitivity to a specified pixel replacement on this photograph, not unique importance of an ear, face or other semantic part. The covers add gray edges and change the input distribution. Probability drops are not additive.
Each coloured square records a separate forward pass with one 16 × 16 patch covered.
The model and its patch size are unchanged; only the covered area shrinks. A fresh input is used for each of 196 tests. The largest drop is P78, 3.836018 points. This is a probability-change map, not attention. Small drops can occur because other regions preserve related clues; they do not prove a patch is useless. Detailed experimental protocol and all measurements remain in the reference deck.
What have we built, and what is still missing?
CNNs gather wider context through local layers. Global ViT attention can connect distant patches in a single block.
These drawings illustrate receptive fields and attention connections; they are not measured maps. Both families can use information from the whole image. We are comparing ordinary local CNNs with this ViT’s global attention. The diagrams do not tell us which model will perform better on a particular task.
An inductive bias is an assumption built into the model before it learns from data.
Convolution builds in local neighborhoods and shared filters. ViT also has built-in assumptions: it uses patches and shared projections. Its global attention weights depend on the content, and it receives position information explicitly. The data, training setup and pretrained weights all affect performance; the optional material discusses model choice.
CNN filters assume that nearby pixels matter and that the same pattern can occur in different places.
The boxes and arrows are schematic. A learned local CNN filter is reused across locations; standard convolution mixes a fixed neighborhood with the same kernel weights, although its activations depend on the image. ViT query-key matching computes different source weights for each receiver and image, using shared projection parameters. With limited task data, useful priors or pretrained representations can help; there is no universal winner. Compare models under the actual data, accuracy and compute constraints.
Smaller patches give us a finer grid, with more tokens to process.
For a 224 by 224 image: 32-pixel patches give 49 patch tokens, 16 gives 196, and 8 gives 784. Add one CLS token for this architecture. These are architectural comparisons, not a claim that the saved checkpoint accepts all three patch sizes without adaptation.
Each token compares with every token, so doubling the token count gives four times as many scores.
Including CLS, 197 squared is 38,809 and 785 squared is 616,225, about 15.88 times larger. The approximately 16-fold statement is exact for the patch-only counts. This concerns attention scores, not a 16-fold claim about total runtime or all model operations.
Turn patches into tokens, update them with the encoder, then use final CLS to predict the class.
Read the three panels from left to right, as in the opening model-family recap. Split the image into 196 RGB patches and apply one shared projection; prepend CLS and add position embeddings. All 197 rows pass through 12 encoder blocks. Within each block, full attention gathers messages, joins and projects the head outputs, and adds an update to each incoming row; the MLP then adds its own update to that attention-updated row. Each block has its own parameters. The shape remains 197 by 192. After final normalization, select only the CLS row for the trained 1,000-class head. The patch rows still exist. LayerNorm and the output projection are omitted from this overview; the nearby code and the detailed block slides make them explicit. The final label is the measured prediction for our photograph.
These shapes describe one photograph. The batch dimension is left out.
The patch projection sets the feature width. Adding CLS takes us from 196 rows to 197, and all twelve encoder blocks keep the 197 by 192 shape. Reading CLS selects one row. The classifier then maps its 192 features to 1,000 logits.
B is the batch size. Token 0 is CLS; the head returns scores for the 1,000 ImageNet classes.
This is readable pseudocode for the same pre-LN encoder classifier. The input image tensor has shape B by 3 by 224 by 224. patch_embed includes patch extraction, shared projection and conversion to patch rows. pos has shape 1 by 197 by 192 and broadcasts across the batch. Starting CLS is shared across the batch and prepended at token index zero. Each block updates every patch row and CLS through attention and MLP residual branches. norm(x) normalizes all final rows; [:, 0] selects CLS for every image in the batch, yielding B by 192. The linear head maps that image summary to 1,000 unnormalized class scores. No softmax is performed in these six executable lines.
The reference deck has the full calculations, code, diagrams and exercises.
The reference deck at vision1-reference.html has the full detail. Section 6 covers implementation; sections 7 to 9 contain the optional extensions. You can also find initialization, optimization and CLS-versus-pooling comparisons there. Feature similarity, attention weights and the effect of covering pixels measure different things.
Implementation lab · Optional extensions · Opening the patch-projection box · Opening attention again with real ViT numbers · Interactive lab
Its 1,000 output classes form a fixed vocabulary.
The 192-feature CLS is compared with a learned weight vector for every ImageNet class, with one bias per class. Persian cat and sports car name actual ImageNet classes. The equation uses column-vector notation for h. The classifier cannot accept an arbitrary new class description without changing its readout or model.
We will pick up this question in the CLIP lecture.
The missing piece is a way to turn a class description into a vector we can compare with the image. Our checkpoint cannot do that for arbitrary text. The next lecture explains how CLIP learns image and text representations that can be compared.