Notation used on this pagesymbols, meanings, shapes
    01

    The question we want to answer

    We already know how an encoder builds contextENCODER. Encode the input. DECODER. Predict the next token. ENCODER-DECODER. Generate from a source. . . . . . . . Encoder. . . . . . . one vector per token. labels · retrieval · readouts. . . . . Decoder. next. append → repeat. chat · code · continuation. Encoder. . . . target prefix. source K,V. Decoder. next. source + target prefix. translation · summarizationENCODEREncode the inputDECODERPredict the next tokenENCODER-DECODERGenerate from a sourceEncoderone vector per tokenlabels · retrieval · readoutsDecodernextappend → repeatchat · code · continuationEncodertarget prefixsource K,VDecodernextsource + target prefixtranslation · summarization

    We can reuse the encoder once we turn the image into tokens.

    The previous lecture, Transformers beyond next-token prediction, compared these three model families. The encoder on the left uses full attention to read the complete input. Here we work out how to give it an image.

    Route: ImageImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What should the model predict for this photograph?Pretrained ViT-TinyImageNet-1k1,000 possible classesNewfoundland: 95.73%Our input photographOne measured forward pass

    This photograph comes from Oxford-IIIT Pet. The model chooses from 1,000 ImageNet classes.

    Our photograph is newfoundland_31 from Oxford-IIIT Pet. We use vit_tiny_patch16_224.augreg_in21k_ft_in1k throughout; it assigns Newfoundland a probability of 0.9572675228. Oxford-IIIT Pet provides the image, while ImageNet defines the classifier’s output labels. The optional extension shows how to adapt the model to dog/cat classification.

    Route: ImageImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    An encoder expects vectors. How can an image supply them??vector 1vector 2…vector N

    The encoder takes a sequence of vectors. We need to turn the pixels into that sequence.

    The image is an RGB array arranged in a spatial grid. The encoder takes rows of features. Our first job is to turn that grid into feature rows; Q, K and V come later.

    Five questions build the architecture1Make tokensPixels → vectors2Keep locationWhich patch goes where?3Choose a readoutMany rows → one label4Share informationUse the known encoder5Predict a classSummary → class scores

    Each question gives us a reason to add the next part of the model.

    We will answer each question before introducing the component that solves it. By the end, we will have followed one photograph all the way to a class prediction and seen what a fixed set of output classes leaves out.

    02

    How can an image become a token sequence?

    Route: PatchesImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Why not make every pixel a token?Pixel tokens224 × 22450,176 tokens16 × 16 patches14 × 14196 tokensAttention compares token pairs: its score matrix grows as N².

    Using patches also keeps the number of attention comparisons manageable.

    Pixel tokens would give 50,176 spatial positions, each initially holding RGB. The fine grid is schematic. Patch tokens give 196 positions before CLS. Attention score counts scale quadratically with sequence length. This comparison ignores CLS to emphasize the scale, not memory implementation details.

    Route: PatchesImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Split the same photograph into 196 patches224 ÷ 16 = 14 patches per side14 × 14 = 196 patches

    Each non-overlapping crop is 16 × 16 pixels, with all three colour channels.

    The photograph has already been resized and centre-cropped using the checkpoint transform. We now work with its 224 by 224 RGB input. Patches are ordered row by row.

    Route: PatchesImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What is inside one patch?P6316 × 16 × 3768 input numbers

    A patch has 768 pixel values. We will add its location separately.

    Patch P63 lies at zero-based grid row 4, column 6. The crop illustration is magnified. Before projection, use the checkpoint normalization: RGB to [0,1], subtract 0.5 and divide by 0.5 per channel.

    Route: PatchesImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Flatten one channel at a time: R, then G, then BR channel16 × 16 values256 numbersG channel16 × 16 values256 numbersB channel16 × 16 values256 numbersxᵢ = [ all R pixels | all G pixels | all B pixels ] ∈ ℝ⁷⁶⁸

    Flattening only rearranges the values; it does not learn features.

    Read pixels row by row within each channel, then concatenate R, G and B. This is channel-major ordering, matching PyTorch Unfold and flattened Conv2d kernels. Pixel-major ordering is also valid if the weight columns are permuted to match.

    Route: ProjectionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    How do 768 pixel values become 192 features?xᵢ768 valuesShared learned projectionW_E: 768 × 192b_E: 192cᵢ192 featurescᵢ = xᵢ W_E + b_E

    The learned projection turns those 768 pixel values into 192 features.

    Use row-vector notation: x_i is 1 by 768, W_E is 768 by 192, and b_E has 192 entries. This affine map has no following activation in the patch layer. PyTorch stores the Linear weight transposed, as 192 by 768.

    Route: ProjectionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Should every patch get a different projection?x1: 768 valuessame W_E, b_Ec1: 192x63: 768 valuessame W_E, b_Ec63: 192x196: 768 valuessame W_E, b_Ec196: 192

    Every patch uses the same W_E and b_E.

    Every patch passes through the same projection with its own pixel values. We learn one set of weights and biases, shared across all 196 patches. We still need a way to tell the model where each patch came from.

    Route: ProjectionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Stack the patches into one feature matrixx₁: 768x₂: 768…x₁₉₆: 768W_E, b_Esharedc₁: 192c₂: 192…c₁₉₆: 192X: 196 × 768C: 196 × 192

    The projection changes 768 features to 192. We still have 196 patch rows.

    X contains all flattened patches. C = X W_E + b_E broadcasts the same bias across every row. The batch axis is omitted in the main lecture.

    Route: ProjectionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What does the trained projection return?Measured P63768 inputstrained projection192 output featuresc₆₃[−0.852, 1.339, 0.504, …]

    These are the first three of P63’s 192 features.

    The first three values come from real-patch-path.json for our saved checkpoint. The dots stand for the other 189 coordinates. These features describe the patch before we add position. We have not assigned a meaning to each individual coordinate.

    Route: ProjectionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    We now have a token for each patchIMAGE3 × 224 × 224PATCHES196 cropsPATCHPROJECTION196 × 192+ POSITION+ CLS197 × 192ENCODER× 12197 × 192FINALCLS192CLASSHEAD1,000 logitsHERE

    The image is now 196 rows, each with 192 learned features.

    This diagram tracks our progress through the model. Completed stages fade, and the current stage is larger. We have patch features now; next we need to add where each patch came from.

    03

    How does the model know where a patch came from?

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    The projection is shared. Where does location enter?content c₆₃Where?row 5, column 7

    The content vector describes the crop. A position vector tells the model where it came from.

    P63 is in row 5, column 7 when we count from one. The shared projection has no explicit position index. The picture itself may offer clues about location, but the learned position embedding gives the model that information explicitly.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Add WHAT and WHEREcontent cᵢWHAT+position pᵢWHEREeᵢeᵢ = cᵢ + pᵢ · 192 features

    Adding position preserves the 196 × 192 shape.

    Content vectors are blue and learned positional vectors amber. We add corresponding coordinates, not concatenate them. The vectors remain 192 features wide. The CLS position is added when CLS is included.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Which positional vector goes with each row?P1 contentP63 contentP196 content+position 1position 63position 196P1 + locationP63 + locationP196 + locationLearned table: one 192-number row per token position

    Every image with this grid uses the same learned position table.

    Training updates the position table along with the rest of the model. At inference, the table stays fixed. Our checkpoint has 197 position vectors, including one for the CLS slot.

    04

    We have 196 rows. Which one should we read?

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Many patch representations, one image labelpatch 1patch 63…patch 196Which readout?one image labelMean pooling is another valid readout; this checkpoint was trained with CLS.

    Reuse the CLS readout we already know from text.

    Averaging final patch features is a valid alternative when the model is trained for that readout. We keep the actual checkpoint’s CLS design, rather than substituting a different architecture.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Same CLS idea, different input tokensTEXTCLSsentence tokensencoderfinal CLSclassIMAGECLSpatch tokensencoderfinal CLSclass

    We use patch tokens where the text model used word tokens.

    As in the text model, the starting CLS vector is a learned parameter. It is shown in amber. After the encoder, CLS depends on the input; its colour matches the text or image branch.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Every image starts with the same learned CLS vectorOne learned CLS start192 stored numberssame CLS startsame CLS start

    Both photographs start with exactly the same CLS vector.

    Both photographs use the same trained model, starting CLS vector and CLS position vector. Only their patch inputs differ. The cat helps us see how the input changes the final representation; we are still using the original classifier.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    After the encoder, CLS depends on the imagesame CLS startsame encoder+ these patch rowsdog image summarysame CLS startsame encoder+ these patch rowscat image summary

    The same model produces a different CLS representation for each photograph.

    The arrows show the computation. The dog and cat produce different final CLS states even though they use the same model. The reference deck includes measured traces for both images; this diagram does not display numerical predictions.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    CLS reads the current patch states at every blockCLSstate 0196 patch rowsstate 0CLSstate 1196 patch rowsstate 1CLSstate 2196 patch rowsstate 2blocknext block

    CLS and patch rows can exchange information in both directions.

    Each block computes all updates from its incoming states, in parallel. The next block reads the updated CLS and patch states. Arrows show allowed dependencies, including patch-to-CLS and CLS-to-patch, not guaranteed large weights.

    Route: Prepare rowsImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    We now have the token sequence the encoder needsIMAGE3 × 224 × 224PATCHES196 cropsPATCHPROJECTION196 × 192+ POSITION+ CLS197 × 192ENCODER× 12197 × 192FINALCLS192CLASSHEAD1,000 logitsHERE[ CLS ; P1 ; P2 ; … ; P196 ] + position → 197 × 192

    196 projected patches plus one CLS row, each with 192 features.

    Prepend CLS, then add a position vector to each of the 197 rows. This gives us the complete image sequence. We can now use that sequence to compute Q, K and V.

    05

    How do the patches exchange information?

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    From here, reuse the encoder we already knowCLSP1…P196The encoderwe already knowCLS updatedP1 updated…P196 updated

    We can now apply the same encoder operations we used for text.

    Normalization, self-attention, residual additions and the MLP all work on feature rows. The small pictures in our diagrams identify patches; the encoder itself receives their vectors.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Where do Q, K and V come from?TRANSLATION · cross-attentiontarget-prefix states → Qsource-encoder states → K, VViT · self-attentionCLS + patch statesone image sequenceQ, K and V

    ViT gets Q, K and V from one image sequence. Translation cross-attention gets them from two sequences.

    In ViT, three learned projections read the same normalized image sequence. In translation cross-attention, queries come from target decoder states and keys/values from source encoder states. Translation also contains self-attention; the comparison concerns its cross-attention operation.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    One patch representation has three jobsReceiver: P74current patch state192 featuresQUERY · QContext I seekKEY · KWhat can matchVALUE · VInformation I send

    Every token makes a query, a key and a value through three learned projections.

    Recall the same three roles from text. In this head q_i = LN(e_i) W_Q, k_i = LN(e_i) W_K and v_i = LN(e_i) W_V (with the checkpoint biases). Each is 64 features wide. These are learned numeric vectors; the questions are an analogy for their roles. P74 can receive information using its query and send information using its key and value.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What might different queries try to gather?Illustrative questions: the model uses vectors, not sentencesA patch on the dogWhich other regions help interpret this texture?Other fur / face regionsA background patchWhich other regions provide scene context?Branches / bright backgroundCLSCLSWhich image features help classify the whole image?Useful parts across the photograph

    Each query can gather different information from the same image.

    These questions illustrate what a query might do; they are not translations of the trained vectors. A key has features that can match a query. The value carries information from the same source through a separate projection. An ear need not attend to another ear, and a head has no fixed semantic role. We will follow the calculation and then inspect a measured attention map.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    A source supplies both a key and a valueP74 query q₇₄64 featuresSource P60key k₆₀64 featuresvalue v₆₀64 featuresq₇₄ · k₆₀ / √64one score → row softmaxa₇₄,₆₀ × v₆₀Q and K set a scalar weight. V supplies the vector being weighted.

    The weight for P60 scales all 64 coordinates of P60’s value vector.

    The pictured source provides both k_60 and v_60, using different projections of its current normalized state. q_74 dot k_60 divided by 8 is one score. Its attention weight comes from softmax against every source score in row P74, including CLS. The weight is a scalar; the value and weighted contribution are 64-feature vectors. This diagram illustrates the roles of Q, K and V; it does not show measured attention weights.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Several source values form one receiver’s messageToy example · one receiver · three sources · two value featuressource A0.6×[2, 0][1.2, 0]source B0.3×[0, 1][0, 0.3]source C0.1×[1, 1][0.1, 0.1]Add the contributions: m = [1.3, 0.4]This message goes to the receiver whose query chose the weights.

    Multiply each value vector by its weight, then add the results to get one message.

    These chosen toy weights sum to one; they are not from the dog checkpoint. Reveal the second source, third source, then their sum: 0.6[2,0] + 0.3[0,1] + 0.1[1,1] = [1.3,0.4]. The real head sums 197 source contributions with 64 features each. It produces one message per receiving query. The multihead output projection converts the joined message into the update added to that receiver’s original row.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    ViT uses full attention, not causal attentionTEXT DECODER · causalsource keys →queries↓Earlier + current tokensViT ENCODER · fullsource keys →queries↓CLS + every patch tokenFilled = allowed. ViT sees the complete image: no causal mask.

    Each row is a query; each column is a key. Every pair is allowed, but the weights can differ.

    The small matrices show permitted query-key pairs, not learned weights. In a next-token decoder, future source positions are masked; a ViT encoder receives the complete image, so all 197 by 197 pairs are allowed, including CLS and self-pairs. Actual scores and row-softmax weights are generally asymmetric because Q and K differ. Patch order in the sequence does not impose a temporal or reading direction.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    How does one CLS query collect one message?CLS queryone receivercompare with 197 keys197 scoresrow softmax197 weightsweight the 197 valuessum their contributionsone CLS message64 features in this headOther query rows receive their own messages in parallel.

    We used the CLS query, so this message updates CLS.

    In one head, all 197 weighted value rows contribute to a 64-feature message. We join the heads, project the result to 192 features, and add it to the incoming CLS embedding. The optional attention appendix works through each source’s contribution.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Three heads gather three messages for each tokensame sequence197 × 192Head 1197 × 64Head 2197 × 64Head 3197 × 64concatenate messages197 × 192then output projection

    Join each row’s three 64-feature messages, then project them back to 192 features.

    The heads run in parallel on the same normalized input, each with its own learned projections. They can gather different messages without having fixed roles such as an "ear head". Joining them increases the feature width while keeping the same token rows.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Where does the trained query look?Query: P74Measured source weightsP60Weight: 9.09%P60’s value contributesto P74’s message.Trained ViT · block 4 · head 1Teal: 0 → 10.35%Explore the nine examples ↗

    Choose one query and look at its source weights. Another head or block may gather a different mixture.

    The maps use saved Q/K from trained ViT-Tiny block 4, head 1, query P74. P60 receives 9.088265% and CLS receives 2.089876%. All 197 weights sum to one; the displayed 196 patch weights are not renormalized. Teal intensity is scaled to this map’s maximum. The violet outline marks the query and white marks source P60. An attention map visualizes mixing, not a segmentation or a complete explanation of the class prediction.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Add the attention update to the original embeddingP74 input192 featuresAttentiongather messagesProject message192 features+Updated P74192 featuresKeep the original embeddingoriginal embedding + projected message = updated embeddingEvery patch gets its own update. So does CLS.

    Gather a message, project it to 192 features, then add it to the receiving token’s original embedding.

    Follow P74 as one receiver. Its attention heads gather weighted values from the current image sequence. Concatenate the three 64-feature messages and apply the learned output projection to obtain a 192-feature update. Add that update to P74’s incoming embedding; the message alone is not the new embedding. All other patch rows and CLS receive their own updates in parallel. Normalization is omitted from this conceptual diagram; the actual checkpoint uses U = E + MSA(LN(E)). The complete pre-LN block remains in the optional reference deck.

    Route: MLPImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    The MLP adds one more update to each embeddingAfter attention192 featuresMLP192-feature update+New embedding192 featuresKeep the embedding after attentionThe MLP processes each token separately.This new embedding is ready for the next block.

    Attention gathers context from other rows; the MLP then processes each row’s features.

    The MLP receives the attention-updated representation and computes another 192-feature update. Add it to that same attention-updated representation, not to the original input from before attention. The same MLP parameters are used independently for every patch and CLS. This completes one block. Normalization and the internal 192 → 768 → 192 layers with GELU are omitted from this conceptual figure; the exact formula is E_next = U + MLP(LN(U)). Implementation details remain in the reference deck.

    Route: Attention, MLPImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Every block updates the patches and CLS againEach block: attention update, then MLP update.Block 1Block 2Block 3Block 4Block 5Block 6Block 7Block 8Block 9Block 10Block 11Block 12197 token embeddings · still 192 features each

    The next block starts from these new embeddings. After block 12, we read CLS.

    All 197 rows continue through the stack: 196 patch embeddings and one CLS embedding. Each block reads the states produced by the preceding block and performs attention and MLP updates. Blocks have their own learned weights. The feature width remains 192; what each row represents changes. After the last block, the checkpoint’s final normalization and CLS readout give one image representation for classification.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What is stored, and what changes with the image?LEARNED PARAMETERSCOMPUTED FOR EACH IMAGEPatch projection W, bPatch embeddingsPosition table + starting CLSQ, K, V and attention weightsTransformer weights, including W_Q/K/VContextual patch states + final CLSClassifier weights + biasesLogits + probabilities

    Training learns W_Q, W_K and W_V. Each new image gets its own attention weights.

    At inference, learned parameters stay fixed and the model recomputes its activations for each image. Training uses gradients to update the parameters. The query and key projection weights are learned parameters; the attention matrix is computed from the current input.

    06

    What does the model predict, and what changes its answer?

    Route: Read summaryImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    After the blocks, read the updated CLS embeddingAfter all 12 blocksUpdated CLS192 featuresUpdated patch embeddings196 rows still existRead final CLSone image embedding192 featuresNext: the class headThe classifier uses the image summary carried by CLS.

    The classifier reads the final CLS embedding as its image summary.

    After the final LayerNorm, we select the CLS row to get a 192-feature image embedding. Selecting the row does not update it. This diagram leaves out normalization so we can focus on the readout. All 196 final patch embeddings still exist, but the trained ImageNet head uses CLS. Next, the head turns those features into 1,000 class scores.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    How do 192 features score 1,000 classes?h_CLS192 featuresLinear(192, 1000)1,000 logitssoftmax across classes1,000 probabilitiesOne logit and one probability for every ImageNet class

    The class head learns how each image feature contributes to each class score.

    The head applies a learned affine map. Softmax turns its class scores into probabilities across the output classes. Attention softmax instead normalizes weights across source tokens. We predict the class with the highest score.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What does the model predict for our photograph?Newfoundland95.73%Tibetan mastiff1.63%briard0.67%

    The model assigns Newfoundland a probability of 95.73% among its 1,000 classes.

    These are saved measured predictions from real-inference.json: Newfoundland 95.726752%, Tibetan mastiff 1.625724%, briard 0.674269%. They are not performance or accuracy estimates. The remaining probability belongs to the other ImageNet classes.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What happens if we hide one quarter of the image?Original: 95.73%Cover one quarterSame model.Changed pixels.Same label?Same probability?

    What do you expect to change when we pass the covered image through the same model?

    Keep Newfoundland as the target class throughout this experiment. Compare the original photograph with a fresh copy whose top-left 112 by 112 pixels are replaced with mid-gray. We change the pixels, keeping all tokens and the attention mask unchanged. Here we measure the final prediction; the earlier attention map showed weights inside the model.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Change the pixels, then run the model againcovered = x.clone() # x: (1, 3, 224, 224), normalizedcovered[:, :, :112, :112] = 0 # gray RGB; 112 × 112 pixelsmodel.eval()with torch.inference_mode(): before = model(x).softmax(-1)[0, 256] after = model(covered).softmax(-1)[0, 256]same ViT95.73%ViT83.00%

    The model keeps its learned weights and recomputes patch features, attention weights and class probabilities.

    The checkpoint transform normalizes each channel as (RGB − 0.5) / 0.5, so normalized zero means gray RGB 0.5, not black. Index 256 is Newfoundland. Both calls use the same frozen model in evaluation mode. Image shape, all 196 patch tokens and CLS remain; 49 patch inputs now contain gray pixels. Each test uses a fresh clone, so covers never accumulate.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Cover each quarter, then compare predictionsOriginal P(Newfoundland) = 95.73%Top left83.00%−12.73 pointsTop right79.62%−16.11 pointsBottom left84.54%−11.19 pointsBottom right82.16%−13.57 points

    Covering the top right gives the biggest drop: 16.11 percentage points. All four covered images still predict Newfoundland.

    Choose the four quadrants before viewing results. Run each covered copy separately and measure the same target-class probability. Top-right drops most among these four tests. This shows sensitivity to a specified pixel replacement on this photograph, not unique importance of an ear, face or other semantic part. The covers add gray edges and change the input distribution. Probability drops are not additive.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    Would smaller covers tell us more?Cover one patch: P78196 separate testsLargest local drop95.73% → 91.89%P78: −3.84 pointsRed: probability fallsBlue: probability risesScale: −4 to +4 pointsAll 196 tests keep the same label. Smaller covers probe more local sensitivity.

    Each coloured square records a separate forward pass with one 16 × 16 patch covered.

    The model and its patch size are unchanged; only the covered area shrinks. A fresh input is used for each of 196 tests. The largest drop is P78, 3.836018 points. This is a probability-change map, not attention. Small drops can occur because other regions preserve related clues; they do not prove a patch is useless. Detailed experimental protocol and all measurements remain in the reference deck.

    07

    What have we built, and what is still missing?

    Two ways to gather image contextCNNViTNearby first → wider contextDirect access to distant patches

    CNNs gather wider context through local layers. Global ViT attention can connect distant patches in a single block.

    These drawings illustrate receptive fields and attention connections; they are not measured maps. Both families can use information from the whole image. We are comparing ordinary local CNNs with this ViT’s global attention. The diagrams do not tell us which model will perform better on a particular task.

    Which assumptions are built into the architecture?CNNViTBasic unitsLocal feature mapsPatch-token rowsMixing ruleShared local convolutionContent-dependent attentionLocalityBuilt into each filterLess spatial structure built inPositionGrid structureExplicit position embeddings

    An inductive bias is an assumption built into the model before it learns from data.

    Convolution builds in local neighborhoods and shared filters. ViT also has built-in assumptions: it uses patches and shared projections. Its global attention weights depend on the content, and it receives position information explicitly. The data, training setup and pretrained weights all affect performance; the optional material discusses model choice.

    How can these built-in assumptions help?CNN: reuse a local detectorViT: choose context from contentSame filter, different locationsDifferent weights for each queryA built-in image prior can help with limited data; pretraining changes the comparison.

    CNN filters assume that nearby pixels matter and that the same pattern can occur in different places.

    The boxes and arrows are schematic. A learned local CNN filter is reused across locations; standard convolution mixes a fixed neighborhood with the same kernel weights, although its activations depend on the image. ViT query-key matching computes different source weights for each receiver and image, using shared projection parameters. With limited task data, useful priors or pretrained representations can help; there is no universal winner. Compare models under the actual data, accuracy and compute constraints.

    Route: PatchesImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    How much detail should one token cover?32 × 32 patches49 patch tokens16 × 16 patches196 patch tokens8 × 8 patches784 patch tokens

    Smaller patches give us a finer grid, with more tokens to process.

    For a 224 by 224 image: 32-pixel patches give 49 patch tokens, 16 gives 196, and 8 gives 784. Add one CLS token for this architecture. These are architectural comparisons, not a claim that the saved checkpoint accepts all three patch sizes without adaptation.

    Route: AttentionImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    What does a finer grid cost?16 × 16 patches196 patch tokens197 tokens with CLS8 × 8 patches784 patch tokens785 tokens with CLSPatch tokens × 4 → attention scores ≈ × 1638,809 scores → 616,225 scores per head

    Each token compares with every token, so doubling the token count gives four times as many scores.

    Including CLS, 197 squared is 38,809 and 785 squared is 616,225, about 15.88 times larger. The approximately 16-fold statement is exact for the patch-only counts. This concerns attention scores, not a 16-fold claim about total runtime or all model operations.

    Once an image becomes tokens, the encoder is familiar1 · Make tokens2 · Build context3 · Read CLSSharedprojection224 × 224 RGB+ CLS + positionCLSP1P2…197 rows × 192 features12 encoder blocksFull attentiongather messages + addMLPupdate each row + add197 updated rows × 192Final CLS · 192Updated patch rowsLinear class head1,000 class scoresNewfoundland

    Turn patches into tokens, update them with the encoder, then use final CLS to predict the class.

    Read the three panels from left to right, as in the opening model-family recap. Split the image into 196 RGB patches and apply one shared projection; prepend CLS and add position embeddings. All 197 rows pass through 12 encoder blocks. Within each block, full attention gathers messages, joins and projects the head outputs, and adds an update to each incoming row; the MLP then adds its own update to that attention-updated row. Each block has its own parameters. The shape remains 197 by 192. After final normalization, select only the CLS row for the trained 1,000-class head. The patch rows still exist. LayerNorm and the output projection are omitted from this overview; the nearby code and the detailed block slides make them explicit. The final label is the measured prediction for our photograph.

    Follow the whole model through its shapesPixels3 × 224 × 224Patch projection196 × 192+ CLS + position197 × 19212 encoder blocks197 × 192Final LN → read CLS192Class head1,000 logitsThe blocks keep the same shape.Each row now includes context.

    These shapes describe one photograph. The batch dimension is left out.

    The patch projection sets the feature width. Adding CLS takes us from 196 rows to 197, and all twelve encoder blocks keep the 197 by 192 shape. Reading CLS selects one row. The classifier then maps its 192 features to 1,000 logits.

    The whole ViT in six lines# Project each image patch to 192 features.x = patch_embed(image) # (B, 196, 192)# Add CLS at index 0 and position to every token.x = prepend_cls(x) + pos # (B, 197, 192)# Update all patch embeddings and CLS in each block.for block in blocks: x = block(x) # (B, 197, 192)# Normalize, then extract token 0: final CLS.h = norm(x)[:, 0] # (B, 192)# Turn the image summary into 1,000 class scores.logits = head(h) # (B, 1000)

    B is the batch size. Token 0 is CLS; the head returns scores for the 1,000 ImageNet classes.

    This is readable pseudocode for the same pre-LN encoder classifier. The input image tensor has shape B by 3 by 224 by 224. patch_embed includes patch extraction, shared projection and conversion to patch rows. pos has shape 1 by 197 by 192 and broadcasts across the batch. Starting CLS is shared across the batch and prepended at token index zero. Each block updates every patch row and CLS through attention and MLP residual branches. norm(x) normalizes all final rows; [:, 0] selects CLS for every image in the batch, yielding B by 192. The linear head maps that image summary to 1,000 unnormalized class scores. No softmax is performed in these six executable lines.

    The reference deck has the full calculations, code, diagrams and exercises.

    The reference deck at vision1-reference.html has the full detail. Section 6 covers implementation; sections 7 to 9 contain the optional extensions. You can also find initialization, optimization and CLS-versus-pooling comparisons there. Feature similarity, attention weights and the effect of covering pixels measure different things.

    Route: Class scoresImagePatchesProjectionPrepare rowsAttentionMLPRead summaryClass scores
    The classifier stores a learned vector per known classh_CLS192 featuresw_Newfoundlandclass scorew_Persian catclass scorew_sports carclass score…class scores_c = h_CLSᵀ w_c + b_c

    Its 1,000 output classes form a fixed vocabulary.

    The 192-feature CLS is compared with a learned weight vector for every ImageNet class, with one bias per class. Persian cat and sports car name actual ImageNet classes. The equation uses column-vector notation for h. The classifier cannot accept an arbitrary new class description without changing its readout or model.

    What if our class vocabulary could come from words?“a photo of a dog”? ? ?v_dogWhat if a class vector came from language?NEXT: CLIP

    We will pick up this question in the CLIP lecture.

    The missing piece is a way to turn a class description into a vector we can compare with the image. Our checkpoint cannot do that for arbitrary text. The next lecture explains how CLIP learns image and text representations that can be compared.