Optional recap: earlier intuition
The main 54-slide lecture is complete without this material.
The main 54-slide lecture is complete without this material.
The main 54-slide lecture is complete without this material.
Today we reuse the encoder on the left: image patches become tokens, and the final image representation predicts a label.
The encoder, shown on the left. The full attention square lets every token read every token. The decoder uses causal attention and repeats next-token prediction. The encoder–decoder adds a source pathway into the target decoder through cross-attention. For ViT, replace text tokens with image patches, then use a classification head.
Today we reuse the encoder on the left: image patches become tokens, and the final image representation predicts a label.
Summary diagram reused from Transformers beyond next-token prediction (“Encoder, decoder-only and encoder-decoder models”). The same token rows, attention patterns and source-to-target pathway connect this lecture to that recap.
Attention builds context in either sequence. The input representation and the task head determine how we use the encoder.
The input embedding and task head change. The encoder machinery remains: attention, MLPs, residuals and normalization. Both pictured classifiers read final CLS.
Attention builds context in either sequence. The input representation and the task head determine how we use the encoder.
Both routes produce a row of D features. Text learns a vocabulary table; vision learns a shared pixel projection. Our image checkpoint chooses D = 192.
No. Its pixel values go through a learned affine layer. Once position is added, the encoder works with feature rows in either case.
Both routes produce a row of D features. Text learns a vocabulary table; vision learns a shared pixel projection. Our image checkpoint chooses D = 192.
Section 1
The image classification task
A text encoder built context from supplied tokens.
What should a model predict from a photograph?
Labeled photos → The image task → Useful clues
Start with the dataset, choose an image label, then ask which parts of the photograph help us decide.
Oxford-IIIT Pet: six examples. Each photo has a species label and a breed label.
Which parts of the photograph helped you decide?
For the rest of this lecture, we observe the whole photograph and predict one class label.
Attention uses the same Q, K, V roles as in text.
| Vector we follow | Role |
|---|---|
| Query from the dark receiver | What to look for |
| Key from a source such as the face | What can match the query |
| Value from that source | Information to send |
Receiver: this dark crop is ambiguous.
Source: a face patch elsewhere in the same image.
Face information → the dark patch’s numbers. Context arrives; the original pixels stay fixed.
Next: build the image queries, keys and values, then work through a numerical example.
Queries and keys determine weights; values supply the information to combine. The dark patch is our receiver, and the face patch is one possible source.
The receiver is the same dark patch.
| Source | Illustrative contribution |
|---|---|
| Face patch | Larger share of the message |
| Branches | Smaller share of the message |
Attention learns the weights. These shares illustrate the idea; they are not measured results.
Here, face clues could help more than branches. Attention weights control how much each source contributes to the message.
The dark crop’s pixels stay unchanged.
| Source | Contribution |
|---|---|
| Face patch representation | face value × its attention weight |
| Branch patch representation | branch value × its attention weight |
Sum the weighted values, then apply the output projection to form the context message. Two sources are shown; all image rows can contribute.
Current patch row + context message → updated patch row.
Later layers use the updated rows to predict one label for the whole image.
The face and branches contribute value vectors. Attention weights mix those vectors, and an output projection forms the message added to the dark patch’s row. The pixels stay fixed.
| Text receiver | Image receiver |
|---|---|
| bank row → W_Q → q_bank | dark-patch row → W_Q → q_dark |
| Seek context for bank | Seek context for the dark texture |
The row being updated makes a query. Its learned projection determines what kinds of source information it can match.
| Text source | Image source |
|---|---|
| river row → W_K → k_river | face-patch row → W_K → k_face |
| q_bank compares with source keys | q_dark compares with source keys |
Each source row makes a key. Comparing the receiver’s query with source keys gives the scores used to choose attention weights.
| Text | Image |
|---|---|
| river row → W_V → v_river | face-patch row → W_V → v_face |
| Sum weighted source values | Sum weighted source values |
Q and K choose the weights. V supplies the information to combine using those weights; the combined message updates the receiving representation.
Attention works on rows of numbers. Next, we choose image patches and turn their pixels into those rows.
Patches become feature rows. Learned CLS and position prepare the input. Encoder blocks build context; the final CLS feeds the classifier.
The shared patch projection creates 196 rows. CLS adds row 197; position adds location information without changing the shape.
Patches become feature rows. Learned CLS and position prepare the input. Encoder blocks build context; the final CLS feeds the classifier.
Architecture adapted from Dosovitskiy et al., 2020. The original paper figure is retained in the optional reference material below.
Optional reference and extra examples
To images: An Image is Worth 16 × 16 Words

Image → patches → projected rows + positions + CLS → encoder → image class scores.
16×16 is the pixel size of one patch. The nine patches drawn in this figure are schematic.
Image patches become tokens. The encoder builds a summary for classification.
Optional recap and extra examples
From text: Attention Is All You Need

Follow the encoder: text embeddings + positions → self-attention and feed-forward layers → contextual rows.
Next: give an encoder image patches and classify the image.
Follow the encoder: text tokens become rows, then attention gives them context.
How many images and classes are there?
| Dataset quantity | Count |
|---|---|
| Images | 7,349 |
| Train / validation | 3,680 |
| Test | 3,669 |
| Species classes | 2: cat and dog |
| Breed classes | 37 |
Our opening question uses 2 classes: cat or dog. Predicting the breed would use 37 classes.
What shape is one image?
Height × width × channels
Original: 334 × 500 × 3
Resize and center-crop ↓
Model input: 224 × 224 × 3 = 150,528 numbers
Original sizes vary. Our demo uses 224 × 224 pixels, each with red, green and blue values: 150,528 numbers.
Classification: name the animal
Output: one class label for each photograph. Our example labels are dog and cat.
What were we asking the text model to predict?
Part I · character tokens
a a b → window embeddings + MLP → character scores → next character.
Choose one, append it, slide the window and run again.
A name grows one character at a time. Our fixed-window model scores the possible next characters.
And what were we predicting in the bank example?
Part II · word tokens
The fisherman sat beside the river bank and watched the ___
Updated final the row → vocabulary scores → a possible next word: water.
Append the chosen word and run again. Bank is an earlier contextual row; the final the row supplies this prediction.
Here one token is a word. The final “the” row uses the earlier context to predict what comes next; “water” is one plausible continuation.
What changed, and what stayed the same?
Both models score possible answers. What they observe, summarize and predict differs.
Would you recognize this crop on its own?
Fur, shadow, or background? Seeing the face elsewhere in the photo helps us interpret this crop.
What can we carry over from our text models?
We still turn inputs into rows, read useful information, and predict an answer.
Part I built a character predictor: embeddings, scores, probabilities, loss and learning. Part II gave “bank” a contextual representation. Part III let “coat” read multiple kinds of information. We reuse their row-vector notation, seven colours, and calculation sequence.
Back to text: what did attention update?
The fisherman sat beside the river bank and watched the …
| Step | Text example |
|---|---|
| Token | bank at position 7 |
| Initial row | token embedding + position |
| Read context | Q and K set weights; V supplies information |
| Update | e7 + attention update → e7′ |
| Predict after the full prefix | Use the updated final the row, e10′ |
Each token starts with an embedding plus position. Attention supplies context; adding its update gives a new representation of the same token.
What could be the image equivalent of a token?
| Text | Image |
|---|---|
| One token in the word-level toy | One fixed-size image patch |
| One row per token | One row per patch |
We can use one fixed-size patch as an image token. Each patch gets its own row, just as each text token did.
What could be the image equivalent of an embedding?
| Text | Image |
|---|---|
| bank → [0.7, 0.7, 0, 0.7] | patch pixels → learned projection → patch row |
| Add token position | Add patch position |
Text uses a learned lookup table. A patch projection produces D coordinates per patch. D is shared across patches and can differ from the number of pixel values.
Could this dark texture belong to the animal?
Receiver P10
A possible query: Which patches could connect this texture to the animal?
q₁₀ = x₁₀ W_Q. These questions express intuition; the model computes vectors.
Face and coat clues could help this patch represent part of the animal. The classifier will still predict one label for the whole photo.
Where is the rest of this face?
Receiver P7
A possible query: Which patches could complete the face around this eye?
q₇ = x₇ W_Q. These questions express intuition; the model computes vectors.
One patch contains only part of a face. Context from other patches could help its representation describe a larger visual structure.
Where does this branch continue?
Receiver P4
A possible query: Which patches could continue this branch beyond the crop?
q₄ = x₄ W_Q. These questions express intuition; the model computes vectors.
Background patches can also gather context. A branch representation could use nearby edge continuity; the task still predicts the image label.
What could each source offer for matching?
Receiver: dark patch P10, with query q₁₀.
Possible matching features: face-like shape and appearance. k₇ = x₇ W_K.
Possible matching features: fur-like texture and appearance. k₁₁ = x₁₁ W_K.
Possible matching features: branch-like edges and appearance. k₈ = x₈ W_K.
A different receiver can find different keys relevant. The words describe intuition for numerical vectors.
A source key can match one query well and another poorly. Comparing this receiver’s query with all source keys sets its attention weights.
What information could these values carry?
Possible information to send: shape around the eye and muzzle. v₇ = x₇ W_V.
Possible information to send: coat texture and colour. v₁₁ = x₁₁ W_V.
Possible information to send: background edges and context. v₈ = x₈ W_V.
Message for P10 = a₁₀,₇ v₇ + a₁₀,₁₁ v₁₁ + a₁₀,₈ v₈ + …
a₁₀,₇ is the weight for P10 reading P7. Other source values also contribute. Every receiver uses its own weights.
Each source supplies one value vector. Different receivers mix these values with different weights. Here a₁₀,₇ is the weight for P10 reading P7.
What is the “next token” for this image?
| Text generation | Image classification |
|---|---|
| Prefix tokens | Whole image |
| Updated final token row | One image summary |
| Vocabulary scores → next token | Class scores → image label |
Here the target is an image label. We reuse attention to build representations; our classification task changes what we read out and predict.
Opening the patch-projection box
Section 2
From pixels to patch embeddings
Attention updates a row of numbers for each token.
How do we make those rows from an image?
RGB pixels → Shared linear layer → Patch embeddings
Read one small patch, calculate its embedding, then scale the same operation to a real photograph.
- Image
- Patches
- Projection
- Prepare rows
- Attention
- MLP
- Read summary
- Class scores
Make patch features and prepare the rows. Inside a block, attention shares information and the MLP transforms each row.
Read one image summary and score the classes. We will open each box as we reach it.
First turn pixels into patch features. Prepare those rows, let them exchange information, then read one image summary to score the classes. We will open each box as we reach it.
The grid cuts through the photograph before the model knows where the dog is.
| Pixel | RGB / 255 |
|---|---|
| A | [1, 0, 0] |
| B | [0, 1, 0] |
| C | [0, 0, 1] |
| D | [1, 1, 1] |
For this calculation, divide RGB values by 255. Read A, B, C, D: top-left, top-right, bottom-left, bottom-right.
| Channel group | Entries |
|---|---|
| R | [1, 0, 0, 1] |
| G | [0, 1, 0, 1] |
| B | [0, 0, 1, 1] |
x₁ = [1, 0, 0, 1, 0, 1, 0, 1, 0, 0, 1, 1]
Shape: (1, 12).
We use all R values, then G, then B: 12 values for this patch. Pixel-by-pixel RGB also works if the weight columns follow that order. Our PyTorch code uses the channel-first order.
At the image input: 12 pixel values → nn.Linear(12, 2) → 2 embedding coordinates.
c₁ = x₁W + b. This affine map has no ReLU or GELU afterward.
Later, after attention: the block MLP uses Linear(D, H) → GELU → Linear(H, D).
Feature widths: D → H → H → D. D is the embedding width; H is the hidden width.
Our patch embedding uses one affine layer. The later block MLP puts GELU between two linear layers. D is the embedding width; H is the MLP’s hidden width.
| Quantity | Shape |
|---|---|
| x₁: pixel row | (1,12) |
| W: weights | (12,2) |
| b: bias row | (1,2) |
| c₁: content embedding | (1,2) |
proj = nn.Linear(12, 2)
c₁ = x₁ W + b.
A projection forms weighted sums of the pixel values and adds a bias. This one linear layer turns 12 inputs into a 2-coordinate patch embedding.
A.R is the red value of pixel A. Each pixel contributes three input nodes.
Every input connects to both outputs: 24 weights, plus one bias for each output. c₁ = [y₁, y₂].
Each input node holds one RGB value. Both outputs read all 12 inputs. Together, y₁ and y₂ form the embedding for this one patch.
Highlighted edges carry the nonzero weights for output 1. Its other incoming weights are zero.
A.R + B.R + C.R + D.R
1 + 0 + 0 + 1 + 0.5 (bias) = 2.5.
Multiply each input by its connection weight, add the contributions, then add the bias. The highlighted connections show the nonzero weights for this output.
Highlighted edges carry the nonzero weights for output 2. Its other incoming weights are zero.
A.G + B.G − C.G − D.G
0 + 1 − 0 − 1 − 0.5 (bias) = -0.5.
Multiply each input by its connection weight, add the contributions, then add the bias. The highlighted connections show the nonzero weights for this output.
c1 = proj(x1)
(1,12) → (1,2)
c₁ = [2.5, −0.5]
No activation: −0.5 stays −0.5.
Later block MLP: Linear → GELU → Linear.
c₁ represents the content of patch 1. Its two coordinates are features for the model; the image classifier comes later.
One prepared image: 224 rows × 224 columns × 3 RGB values.
Next, divide this image into patches.
We now carry this one image through the real model’s input pipeline. It has already been resized and cropped to 224 × 224 RGB pixels.
Cut every 16 pixels along x and along y. Count the resulting patches on the next slide.
The x-axis runs right; the y-axis runs down. Both span 224 pixels. Cut every 16 pixels along each axis. Each piece keeps its RGB values; we will count the pieces next.
| Count | Calculation |
|---|---|
| Columns | 224 ÷ 16 = 14 |
| Rows | 224 ÷ 16 = 14 |
| Total patches | 14 × 14 = 196 |
Patch array: 196 × 16 × 16 × 3.
P63 is row 5, column 7; P64 is the next patch to its right.
Number left to right, then continue on the next row. Use the image and patch sizes to calculate the row length and total before revealing the answers.
| Pixel | RGB |
|---|---|
| First | [16, 17, 12] |
| Second | [41, 42, 37] |
16 × 16 = 256 pixels.
256 × 3 = 768 channel values.
P63 contains 256 pixels. Its first pixel is RGB [16, 17, 12]; its next pixel is [41, 42, 37]. Three values per pixel give 768 numbers in this same patch.
| Pixel 1 | Values |
|---|---|
| Raw RGB | [16, 17, 12] |
| Normalized | [−0.875, −0.867, −0.906] |
| Pixel 2 | Values |
|---|---|
| Raw RGB | [41, 42, 37] |
| Normalized | [−0.678, −0.671, −0.710] |
Apply (value / 255 − 0.5) / 0.5 to each channel. The patch still has 16 × 16 × 3 values.
The checkpoint rescales every channel using the same formula. For the first red value, (16 / 255 − 0.5) / 0.5 ≈ −0.875. P63 still contains 768 values.
| Channel | First normalized values |
|---|---|
| R | [−0.875, −0.678, −0.561, …] |
| G | [−0.867, −0.671, −0.553, …] |
| B | [−0.906, −0.710, −0.592, …] |
x₆₃: 1 × 768. Entry order: R₁…R₂₅₆, G₁…G₂₅₆, B₁…B₂₅₆.
Concatenate 256 red values, 256 green values and 256 blue values. This gives one 768-number patch row. F.unfold and the reshaped Conv2d weights use this exact ordering.
x₆₃ (1 × 768) → nn.Linear(768, 192) → c₆₃ (1 × 192).
| Parameter | Shape |
|---|---|
| W_patch | 768 × 192 |
| Bias | 192 entries |
Weighted sums + bias, with no activation. One shared set of 147,648 parameters.
This layer reads 768 values and computes 192 output features. Its learned weights and biases are shared by every patch. One patch row enters; one embedding row leaves.
| Output coordinate | Measured value |
|---|---|
| 1 | −0.852 |
| 2 | 1.339 |
| 3 | 0.504 |
| 192 | −1.199 |
c₆₃: 1 × 192. One learned content representation of P63.
The layer produces c₆₃, the content embedding for P63. These are actual outputs from the pretrained model, rounded here. The first feature is −0.852 and the last is −1.199.
| Patch | Content embedding |
|---|---|
| 1 | [−0.205, 0.082, −0.205, …] |
| 63 | [−0.852, 1.339, 0.504, …] |
| 64 | [0.161, 0.988, −1.148, …] |
| 196 | [−1.309, 1.327, −2.234, …] |
X (196 × 768) → shared layer → C (196 × 192).
196 patch rows, each with 192 learned features.
Repeat the same projection for all 196 patches and keep their order. The input matrix X is 196 × 768. The output matrix C is 196 × 192: one content embedding per image patch.
Inspect all 196 output rows (first three coordinates)
| Patch | First three output features |
|---|---|
| 1 | [−0.205, 0.082, −0.205, …] |
| 2 | [−0.205, 0.082, −0.205, …] |
| 3 | [−0.205, 0.082, −0.205, …] |
| 4 | [−0.205, 0.082, −0.205, …] |
| 5 | [−0.118, 0.263, 0.080, …] |
| 6 | [−0.125, −0.208, −0.452, …] |
| 7 | [−0.121, 0.006, 0.165, …] |
| 8 | [−0.265, 0.158, −0.533, …] |
| 9 | [−0.063, −0.232, −0.245, …] |
| 10 | [0.203, 1.288, −0.448, …] |
| 11 | [1.059, 0.144, 0.550, …] |
| 12 | [−0.514, −0.216, −0.150, …] |
| 13 | [−0.079, 0.055, 0.461, …] |
| 14 | [−1.323, 3.365, −0.333, …] |
| 15 | [−0.193, 0.119, −0.216, …] |
| 16 | [−0.151, 0.199, −0.278, …] |
| 17 | [−0.185, 0.225, −0.261, …] |
| 18 | [−0.215, 0.117, −0.171, …] |
| 19 | [−0.110, 0.348, −0.011, …] |
| 20 | [−0.986, 1.153, −0.731, …] |
| 21 | [−0.423, 1.554, 0.095, …] |
| 22 | [−1.129, −0.134, 0.466, …] |
| 23 | [−1.841, −1.377, 0.110, …] |
| 24 | [−0.226, 0.270, −1.916, …] |
| 25 | [1.073, 1.062, −1.398, …] |
| 26 | [−0.745, 0.609, −1.959, …] |
| 27 | [0.049, 1.647, 0.476, …] |
| 28 | [−1.435, 1.945, −2.812, …] |
| 29 | [−0.010, −0.061, −0.096, …] |
| 30 | [−0.021, −0.400, −0.374, …] |
| 31 | [−0.940, −0.087, 1.085, …] |
| 32 | [−0.075, −0.096, 0.544, …] |
| 33 | [−1.020, 1.022, −0.271, …] |
| 34 | [−0.515, −0.100, −0.238, …] |
| 35 | [−0.537, 0.097, −0.351, …] |
| 36 | [−0.598, −0.020, −0.388, …] |
| 37 | [−0.822, −0.272, −0.154, …] |
| 38 | [−0.614, 0.837, −0.108, …] |
| 39 | [0.234, 2.156, −2.354, …] |
| 40 | [−1.637, 0.503, −1.848, …] |
| 41 | [−0.451, 2.639, −0.193, …] |
| 42 | [−0.991, −0.307, −0.661, …] |
| 43 | [−0.444, 1.215, −1.874, …] |
| 44 | [−0.251, −0.527, 1.191, …] |
| 45 | [−0.532, −0.300, 0.378, …] |
| 46 | [0.968, 1.458, −1.695, …] |
| 47 | [−0.515, −0.806, 0.246, …] |
| 48 | [−0.630, −0.557, 0.359, …] |
| 49 | [−0.753, −0.418, −0.373, …] |
| 50 | [−0.593, 0.160, 0.306, …] |
| 51 | [−0.593, −0.206, −0.517, …] |
| 52 | [−0.619, −0.178, −0.410, …] |
| 53 | [−0.970, −0.009, −0.243, …] |
| 54 | [−0.608, −2.775, 3.317, …] |
| 55 | [−1.582, 0.068, −0.702, …] |
| 56 | [−1.321, 1.534, 1.108, …] |
| 57 | [−0.766, −2.618, 2.394, …] |
| 58 | [−0.057, 2.208, 0.862, …] |
| 59 | [−0.111, 0.609, 0.450, …] |
| 60 | [0.327, −0.281, 0.687, …] |
| 61 | [−0.566, 0.268, −0.527, …] |
| 62 | [−0.695, −0.309, −0.192, …] |
| 63 | [−0.852, 1.339, 0.504, …] |
| 64 | [0.161, 0.988, −1.148, …] |
| 65 | [−0.250, 0.146, −0.390, …] |
| 66 | [−0.578, 0.009, −0.289, …] |
| 67 | [−0.767, −0.137, −0.566, …] |
| 68 | [−0.842, −1.307, −1.350, …] |
| 69 | [−0.572, −0.098, −1.013, …] |
| 70 | [−0.149, 0.373, 0.458, …] |
| 71 | [−0.770, 2.427, −2.082, …] |
| 72 | [−0.433, 0.832, −1.060, …] |
| 73 | [−0.839, 1.788, −0.246, …] |
| 74 | [0.983, 1.150, −1.336, …] |
| 75 | [−0.606, −0.620, 0.026, …] |
| 76 | [−0.615, −0.623, 0.060, …] |
| 77 | [−0.394, 0.341, 0.139, …] |
| 78 | [−0.798, −0.765, −0.379, …] |
| 79 | [−0.591, −0.144, −0.476, …] |
| 80 | [−0.587, −0.047, −0.065, …] |
| 81 | [−0.707, 0.114, −0.298, …] |
| 82 | [−1.319, −0.202, −1.935, …] |
| 83 | [−0.284, 0.372, −0.270, …] |
| 84 | [−0.228, −0.141, −0.244, …] |
| 85 | [−0.969, −3.528, −1.112, …] |
| 86 | [−0.394, 0.557, −1.440, …] |
| 87 | [−0.106, 0.459, 0.099, …] |
| 88 | [0.141, 0.838, 0.430, …] |
| 89 | [−0.720, −0.162, −0.066, …] |
| 90 | [−0.775, −0.106, −0.369, …] |
| 91 | [−1.269, 1.772, −2.034, …] |
| 92 | [−1.148, −1.202, 1.874, …] |
| 93 | [−0.546, −0.205, −0.379, …] |
| 94 | [−0.588, 0.031, −0.364, …] |
| 95 | [−0.663, −0.021, −0.324, …] |
| 96 | [−1.472, −1.028, −2.259, …] |
| 97 | [−0.405, 0.179, 0.167, …] |
| 98 | [−0.564, 0.877, −0.916, …] |
| 99 | [−0.814, 0.330, −0.925, …] |
| 100 | [−0.071, 1.857, −2.427, …] |
| 101 | [0.699, −0.949, 0.395, …] |
| 102 | [−0.373, 0.132, −0.137, …] |
| 103 | [−0.613, 0.046, −0.198, …] |
| 104 | [−0.753, −0.075, −0.720, …] |
| 105 | [−0.730, 0.531, −0.013, …] |
| 106 | [−0.248, 1.284, −0.949, …] |
| 107 | [−0.545, 0.228, −0.367, …] |
| 108 | [−0.649, −0.225, −0.384, …] |
| 109 | [−0.602, −0.171, −0.334, …] |
| 110 | [−2.213, −0.990, −2.786, …] |
| 111 | [0.157, −1.443, 3.069, …] |
| 112 | [−0.591, 0.013, 0.172, …] |
| 113 | [−1.120, 0.547, −0.546, …] |
| 114 | [−0.743, −1.882, 0.202, …] |
| 115 | [0.651, 2.026, −1.209, …] |
| 116 | [−0.605, −0.327, −0.029, …] |
| 117 | [−0.438, 0.140, −0.538, …] |
| 118 | [−0.588, −0.107, −0.247, …] |
| 119 | [−0.497, 0.069, −0.208, …] |
| 120 | [−0.650, −0.027, −0.273, …] |
| 121 | [−0.552, 0.075, −0.384, …] |
| 122 | [−0.703, −0.238, −0.406, …] |
| 123 | [−0.626, −0.004, −0.196, …] |
| 124 | [−1.706, 0.474, −0.672, …] |
| 125 | [−0.576, 4.801, −4.632, …] |
| 126 | [−0.421, 0.690, −1.300, …] |
| 127 | [−0.878, −0.028, −2.594, …] |
| 128 | [0.207, −0.317, 0.132, …] |
| 129 | [−0.218, 0.773, −0.158, …] |
| 130 | [−0.619, −0.187, −0.087, …] |
| 131 | [−0.575, 0.067, −0.450, …] |
| 132 | [−0.645, −0.073, −0.068, …] |
| 133 | [−0.437, −0.007, −0.381, …] |
| 134 | [−0.533, −0.336, −0.141, …] |
| 135 | [−0.729, −0.356, −0.178, …] |
| 136 | [−0.704, −0.005, 0.039, …] |
| 137 | [−0.628, −0.116, 0.036, …] |
| 138 | [−0.568, 0.032, −0.201, …] |
| 139 | [−2.346, −1.107, 0.300, …] |
| 140 | [−0.289, −0.602, 0.342, …] |
| 141 | [−1.132, 1.496, −1.984, …] |
| 142 | [−0.525, 0.362, −0.494, …] |
| 143 | [−0.109, −0.010, 0.552, …] |
| 144 | [−0.573, 0.124, −0.264, …] |
| 145 | [−0.533, 0.169, −0.494, …] |
| 146 | [−0.547, −0.123, −0.221, …] |
| 147 | [−0.509, −0.082, −0.465, …] |
| 148 | [−0.570, 0.102, −0.162, …] |
| 149 | [−0.607, 0.291, −0.704, …] |
| 150 | [−0.670, −0.159, −0.142, …] |
| 151 | [−0.618, −0.169, −0.339, …] |
| 152 | [−0.595, −0.098, −0.339, …] |
| 153 | [−1.848, −1.944, −0.137, …] |
| 154 | [−0.155, −0.913, −1.044, …] |
| 155 | [−0.555, 1.527, −0.953, …] |
| 156 | [−0.842, −0.488, 1.044, …] |
| 157 | [−0.273, −0.047, 0.176, …] |
| 158 | [−0.567, 0.026, −0.313, …] |
| 159 | [−0.628, 0.102, −0.479, …] |
| 160 | [−0.673, 0.112, −0.050, …] |
| 161 | [−0.655, −0.364, −0.359, …] |
| 162 | [−0.701, −0.170, −0.321, …] |
| 163 | [−0.658, −0.174, −0.150, …] |
| 164 | [−0.638, −0.018, −0.161, …] |
| 165 | [−0.405, −0.035, −0.398, …] |
| 166 | [−0.597, −0.016, −0.560, …] |
| 167 | [−1.682, −0.439, −0.968, …] |
| 168 | [−0.677, 1.777, −1.965, …] |
| 169 | [−1.678, −1.940, −0.434, …] |
| 170 | [0.207, −0.497, −2.812, …] |
| 171 | [0.666, −0.523, 1.817, …] |
| 172 | [−0.556, 0.096, −0.490, …] |
| 173 | [−0.737, 0.005, −0.251, …] |
| 174 | [−0.517, −0.060, −0.467, …] |
| 175 | [−0.592, −0.109, −0.263, …] |
| 176 | [−0.629, 0.184, −0.451, …] |
| 177 | [−0.758, −0.115, −0.330, …] |
| 178 | [−0.620, 0.008, −0.350, …] |
| 179 | [−0.549, −0.085, −0.138, …] |
| 180 | [−0.594, −0.294, −1.117, …] |
| 181 | [−1.816, 1.131, 0.547, …] |
| 182 | [−0.954, 1.171, −1.776, …] |
| 183 | [−0.377, −3.448, 0.850, …] |
| 184 | [−0.894, 0.878, −1.173, …] |
| 185 | [0.880, 2.078, −1.083, …] |
| 186 | [−0.499, −0.137, −0.307, …] |
| 187 | [−0.568, −0.541, −0.129, …] |
| 188 | [−0.503, −0.184, −0.525, …] |
| 189 | [−0.594, −0.161, −0.854, …] |
| 190 | [−0.643, 0.254, −0.713, …] |
| 191 | [−0.754, −0.189, −0.078, …] |
| 192 | [−0.621, 0.015, −0.477, …] |
| 193 | [−0.599, −0.061, −0.424, …] |
| 194 | [−0.917, −0.366, −0.959, …] |
| 195 | [−1.206, −2.572, −0.466, …] |
| 196 | [−1.309, 1.327, −2.234, …] |
Optional recap and extra examples
How can a red pixel be three numbers?
Each pixel records a red, green and blue channel.
Which pixel goes first in the row?
Top row: 1 → 0
Bottom row: 0 → 1
Flattened row: [1, 0, 0, 1]
Read left to right, then return to the start of the next row. The 0s and 1s are pixel values, not step numbers.
How does this connect to text embeddings?
| Text | Image |
|---|---|
| Token ID | 12 pixel values |
| nn.Embedding(V, 4) | nn.Linear(12, 2) |
| Look up 4 coordinates | Compute 2 coordinates; no activation |
The patch layer computes x₁W + b. No ReLU or GELU follows this projection.
Here the patch embedding is x₁W + b, with no activation afterward. Text looks up a learned row; the image layer computes one from pixels.
Opening attention again with real ViT numbers
Section 3
Prepare the rows, then classify the image
Our photograph has become 196 rows of patch features.
Where is each patch? How do we get one image summary?
Add location → Introduce CLS → Resume the forward pass
Keep the photograph fixed. Give each patch its location, introduce the summary row, then follow the complete input through the classifier.
The photograph is unchanged. P63 is in row 5, column 7 of the 14 × 14 grid: 4 × 14 + 7 = 63.
c₆₃: 192 content features computed from its pixels.
p₆₃: 192 learned coordinates for this grid slot.
Recognizing a face depends on how its parts are arranged. The shared pixel projection has not received the row or column. Add position so attention can use appearance and layout.
Recognizing a face depends on how its parts are arranged. c₆₃ describes this crop; p₆₃ tells attention where it belongs. Add them so the model can use both appearance and layout.
| Item | Meaning |
|---|---|
| P | 196 patch slots × 192 coordinates |
| P63 | Row 5, column 7 selects p₆₃ |
| Saved p₆₃ preview | [−0.815, −0.083, …] |
| Across images | Reuse the same table |
| Patch-position parameters | 37,632 |
| Patch-projection parameters | 147,648 |
| Including the later CLS position | 197 × 192 = 37,824 |
Each grid slot has a trainable 192-number row, shared across images. Slot 63 always selects p₆₃; its pixel-derived content changes from photo to photo. The patch-position table has 37,632 parameters.
| Row · shape 1 × 192 | First two coordinates |
|---|---|
| c₆₃: projected content | [−0.852, 1.339, …] |
| + p₆₃: position | [−0.815, −0.083, …] |
| = e₆₃: block input | [−1.666, 1.256, …] |
All coordinates are added in matching positions. The width remains 192. Values are rounded.
c₆₃ comes from the patch pixels. p₆₃ is learned for its grid location. Add matching coordinates to get e₆₃. All three rows have 192 coordinates; addition does not double the width.
| Step | What happens |
|---|---|
| Initialize | Small random values, once when training from scratch |
| Forward | c₆₃ + p₆₃ → Transformer → class scores → loss |
| Backward | The loss supplies gradients for the shared position table |
| Illustrative SGD | 0.020 − 0.1 × 0.30 = −0.010 |
| Inference | Reuse the trained table without updates |
| Fixed sinusoidal alternative | Compute sin/cos values; no trainable encoding entries |
The class loss supplies gradients for the position table and other model parameters. Training updates the table; inference reuses it unchanged. Fixed sinusoidal encodings instead compute position vectors from a formula.

Problem: we have 196 patch feature rows, but want one label for the whole photo. Which representation should the classifier read?
- Add CLS: reserve an extra learned row before the Transformer blocks.
- Gather image information: all 197 rows pass through the blocks. Attention lets CLS read patch features; the patch rows are updated too.
- Read final CLS: its updated, image-dependent features go to the classifier, which scores image labels.
CLS starts as shared model parameters. The next slides show where those numbers come from.
CLS is one design choice. Averaging the final patch rows is another readout option.
Patch embeddings give us many rows, while the target is one label for the photograph. We need a way to combine patch information into the representation that the classifier will read.
The same dog photograph still gives 196 real patches.
P63: 16 × 16 × 3 pixels → 768 values → Linear(768,192) → 192 content features.
CLS: a separately stored parameter → 192 coordinates. It has no pixels and no patch projection.
196 patch rows + 1 CLS row = 197 rows.
A patch row comes from pixels. CLS comes from a separate set of model parameters. No extra crop is cut from the photograph. Its 192 coordinates match the width of every patch row.
Same dog image → 196 patch rows. Add the shared starting CLS.
- Block 1: CLS can read CLS and all patches. Every patch can also read CLS and all patches. All attention outputs use the same incoming rows.
- Block 2: its attention reads both the updated CLS and the updated patch rows produced by block 1. The new outputs become the next block’s inputs.
- Continue through block 12: every block has its own learned parameters. All 197 rows remain 192 features wide.
Why a summary? The classifier reads final CLS. The class loss trains the network to make its features useful for predicting the image’s class.
We do not supply a target summary vector or assign a meaning to each coordinate. The diagram is conceptual; the next slide shows saved values.
CLS can read all patch rows; each patch can read CLS and every patch. All updates use the rows entering that block. Block 2 reads the updated CLS and patches produced by block 1.
| Row | Content + position | Shape |
|---|---|---|
| CLS | learned token + p₀ | 1 × 192 |
| P1 | c₁ + p₁ | 1 × 192 |
| … | … | … |
| P196 | c₁₉₆ + p₁₉₆ | 1 × 192 |
Stack all rows: E has shape 197 × 192.
Each patch keeps its content-plus-position row. CLS gets its own learned position too. Stack the summary and patch rows into E: 197 rows, each with 192 features. Batch size one is omitted here.

Prepared rows: 196 patches + one CLS → 197 × 192.
Block 1: attention shares information across rows → MLP transforms each row.
Output: 197 × 192, with updated feature values. First follow this one block.
Our 196 patch rows and one CLS row enter block 1. Attention shares information across rows; the MLP transforms each row. The output still has 197 rows and 192 features per row.
Self-attention uses one input sequence. Cross-attention takes queries from one stream and keys and values from another.
No. CLS is one row inside the image sequence. In this pre-LN model, all three projections read the same normalized current rows.
Self-attention uses one input sequence. Cross-attention takes queries from one stream and keys and values from another.
Pixels → patch projection → c₆₃ → add p₆₃ → e₆₃
e₆₃ (1 × 192) → LayerNorm → z₆₃ (1 × 192).
| Head one · separate learned maps | Output shape |
|---|---|
| q₆₃ = z₆₃W_Q + b_Q | 1 × 64 |
| k₆₃ = z₆₃W_K + b_K | 1 × 64 |
| v₆₃ = z₆₃W_V + b_V | 1 × 64 |
Each W: 192 × 64. Each bias: 64 entries. This model uses three heads.
LayerNorm rescales a row’s features and keeps its width. We name it here and study its details later. Three learned maps then make Q, K and V; the diagram shows one of three heads.

X = LayerNorm(E), shape 197 × 192. Rows: CLS, P1, …, P196.
| Projection from the same X | Result |
|---|---|
| X W_Q + b_Q | Q: 197 × 64 |
| X W_K + b_K | K: 197 × 64 |
| X W_V + b_V | V: 197 × 64 |
Each W: 192 × 64; each bias: 64. One head, three different learned projections. Row identities stay the same.
One normalized matrix X feeds three different learned projections. Q, K and V each contain 197 rows of 64 features, in the same token order. This is one head; dots stand for omitted values.
Q (197 × 64) × Kᵀ (64 × 197) / 8 → S (197 × 197).
Rows are receiving queries; columns are source keys. Both axes follow CLS, P1, …, P196.
| Which part of S? | What stays fixed? | What changes? |
|---|---|---|
| CLS row: S[CLS,:] | CLS query (receiver) | Source key across columns |
| CLS column: S[:,CLS] | CLS key (source) | Receiving query down rows |
For the message that updates CLS, use the CLS row.
S[CLS,P63] = dot(q_CLS, k_P63) / 8. Sum 64 coordinate products.
197 × 197 = 38,809 matching scores per head. They can be negative and are not probabilities.
Fix the receiving query to CLS and move across the columns. These scores determine CLS’s source weights after softmax. Moving down the CLS column changes the query while keeping the CLS key fixed.
Highlight row CLS in S (197 × 197), then enlarge that row.
| Source key | Score for the CLS query | After softmax |
|---|---|---|
| CLS itself | s₀ | a₀ |
| P1 | s₁ | a₁ |
| P2 | s₂ | a₂ |
| … | … | … |
| P63 | s₆₃ | a₆₃ |
| … | … | … |
| P196 | s₁₉₆ | a₁₉₆ |
1 CLS self-score + 196 patch scores = 197 scores. For this fixed CLS query, s_j = dot(q_CLS,k_j) / 8.
Softmax across all 197 scores gives 197 weights: a₀ + a₁ + … + a₁₉₆ = 1. Each a_j weights the corresponding source value row.
Why compute the other rows? Each patch gets its own update. The next block’s CLS reads keys and values made from those updated patches.
Final-block exception: if the classifier uses only final CLS, a specialized implementation could compute only that output. It still needs all 197 incoming keys and values. The standard implementation here computes every row.
No causal mask: the whole image is available before predicting its class. The ellipses only shorten the drawing; no source is excluded from softmax.
Earlier blocks update patches so later CLS queries can read their new features. In the final block, a CLS-only readout could compute just the CLS output, using all 197 incoming keys and values.
S (197 × 197) → softmax across each row → A (197 × 197).
The CLS row contains weights for CLS, P1, …, P196. All 197 weights sum to 1.
A[CLS,P63] = exp(S[CLS,P63]) / sum(exp(S[CLS,k])) over all 197 source rows k.
It is the coefficient on P63’s value row in the message to CLS. The sources are image rows, not animal classes.
Apply softmax across each score row. The CLS row becomes weights over CLS and all 196 patches, summing to one. Each other query gets its own weight row. These weights tell us how to mix the value rows.
Text generation: receiving position t_i may read source positions t_j only when j ≤ i.
This image classifier: every receiving row may read CLS and P1 through P196.
There is no causal mask. The whole photograph is available before predicting its label.
CLS can read P196; P196 can read CLS. Every row can also read itself. Permitted access does not imply equal weights.
Next-token prediction hides later text positions. Here the complete photograph is available before classification. Every query can use every key and value, including itself. The ViT attention matrix has no causal triangle removed.
| Receiver | Sources | Result |
|---|---|---|
| bank | river and other causally allowed tokens | weighted value mixture for bank |
| CLS | P1–P196 and CLS itself | weighted value mixture for CLS |
Same attention operation: use the receiver’s weights to mix source value vectors.
In text, bank receives a weighted mixture of allowed token values. Here, CLS receives a weighted mixture of image-token values, including its own. The query chooses the receiver; the values supply the message.

X = LayerNorm(E): 197 × 192. Apply the same Linear(192,64) to every row.
V = X W_V + b_V: 197 × 64, ordered CLS, P1, …, P196.
P63’s value row begins [−0.174, −1.201, …] and has 64 learned features.
Use the same normalized input X that produced Q and K. The value projection makes 64 features for each of its 197 rows. P63’s crop identifies one source; its value row contains learned features.
- Pick the CLS query row in A (197 × 197).
- Pick column P63: A[CLS,P63] = 0.000504.
- Pick the matching source row P63 in V (197 × 64).
- Multiply feature 1, then repeat for the other 63 features.
| Feature of P63 | Value | Weight × value |
|---|---|---|
| 1 | −0.174 | −0.000088 |
| 2 | −1.201 | −0.000606 |
| 64 | −1.308 | −0.000659 |
a₆₃ × v₆₃ → one weighted row of shape 1 × 64. Every feature gets the same scalar weight. Repeat for all 197 sources.
Receiver: CLS. Source: P63. P63’s own message uses its own query and attention weights.
Use Next to select CLS’s row in A, P63’s column, P63’s row in V, then its feature coordinates. One scalar scales all 64 values. This contribution is addressed to CLS.
| Source | CLS weight | First two value features | First two weighted features |
|---|---|---|---|
| CLS | 0.860556 | 0.418, 0.071 | 0.359857, 0.060762 |
| P1 | 0.002298 | −0.211, −0.027 | −0.000485, −0.000063 |
| P63 | 0.000504 | −0.174, −1.201 | −0.000088, −0.000606 |
| P196 | 0.000233 | −1.607, 1.181 | −0.000374, 0.000275 |
Every product is a 1 × 64 contribution to CLS. These four examples are part of the 197-source sum.
Repeat the P63 calculation for other sources. Each uses its own CLS weight and value row. These are four measured examples from the same head; every contribution is a 64-feature vector addressed to CLS.
| Source group | Weighted features 1 and 2 |
|---|---|
| CLS | 0.359857, 0.060762 |
| P1 | −0.000485, −0.000063 |
| P63 | −0.000088, −0.000606 |
| P196 | −0.000374, 0.000275 |
| Other 193 | −0.005982, −0.003703 |
Add down each feature column. The complete message begins [0.353, 0.057, 0.024, …], shape 1 × 64.
Receiver: CLS. This is one head’s message; the embedding update comes next.
Add corresponding coordinates from all 197 weighted value rows, including CLS itself. The result has 64 features: one message for CLS in this head. The other 193 sources are grouped to keep the addition readable.
One head’s CLS message: 1 × 64.
Combine all three heads’ CLS messages, then apply the output projection: CLS context update, 1 × 192.
Current CLS row + CLS context update → updated CLS row (all 1 × 192). This is the same residual pattern used for text tokens.
The CLS messages become a 192-feature context update after combining heads and projecting. Add this update to the CLS row entering the block. The result is a new representation of the same CLS token, as in text.
| Receiver | Weights over shared V | One-head message |
|---|---|---|
| CLS | A[CLS, all 197 sources] | [0.353, 0.057, …] (1 × 64) |
| P63 | A[P63, all 197 sources] | [0.050, −0.358, …] (1 × 64) |
H = A V contains one 64-feature message for each of the 197 receiving rows.
CLS messages update CLS. P63 messages update P63. All receivers use the same incoming sequence.
P63’s query mixes the same source values using P63’s weights, producing P63’s own message. CLS and every other patch do the same. Each receiver’s messages ultimately update its own input row through the residual path.
Head 1 is complete: Q/K → weights → weighted V → H¹ (197 × 64).
Now follow three parallel heads from the same 197 × 192 input. Combine their messages, project, add E, then continue to the MLP.
We have finished one head’s attention calculation. This model has three heads. Each starts from the same input rows and computes its own messages. Next, follow those parallel paths and combine their outputs before the MLP.
| Possible pattern | Example question |
|---|---|
| Nearby texture | Does this fur continue nearby? |
| Animal parts | Which face clues belong together? |
| Wider context | How does this region fit the scene? |
Illustrative hypotheses. The classification loss learns the head parameters.
Different heads can learn different matching patterns and different messages. Local texture, relationships between parts and wider context are possible intuitions. These are hypotheses, not measured head labels.
X = LayerNorm(E), 197 × 192: CLS plus P1 through P196.
| Head | Independent projections | Each result |
|---|---|---|
| 1 | Q¹, K¹, V¹ | 197 × 64 |
| 2 | Q², K², V² | 197 × 64 |
| 3 | Q³, K³, V³ | 197 × 64 |
Each projection uses its own W (192 × 64) and bias (64). All heads read the same CLS row and all patches.
All three heads read the same normalized matrix X. Each has its own Q, K and V projection weights and biases. Every output matrix has 197 rows and 64 features. The highlighted first row is always CLS.
| Head | Score and weight matrices | Message matrix |
|---|---|---|
| 1 | Q¹(K¹)ᵀ / 8 → softmax → A¹ (197 × 197) | A¹ V¹ → H¹ (197 × 64) |
| 2 | Q²(K²)ᵀ / 8 → softmax → A² (197 × 197) | A² V² → H² (197 × 64) |
| 3 | Q³(K³)ᵀ / 8 → softmax → A³ (197 × 197) | A³ V³ → H³ (197 × 64) |
The three paths use different learned projections and run in parallel from the same X.
Within each head, compare Q with K, scale the scores, apply softmax across each row, then multiply by that head’s V. Each path returns 197 messages of 64 features. The heads operate in parallel.
Shared normalized CLS row: [0.044, −0.011, 0.154, …] (1 × 192).
| Head | Measured CLS message | Shape |
|---|---|---|
| 1 | [0.353, 0.057, 0.024, …] | 1 × 64 |
| 2 | [−0.091, −0.112, −0.086, …] | 1 × 64 |
| 3 | [−0.002, −0.566, 0.020, …] | 1 × 64 |
One CLS token. Different head projections and attention computations produce three different messages.
CLS is the same input row for all heads. Their learned projections produce different queries, keys and values, so their attention computations produce different messages. These measured messages each contain 64 features. There is still one CLS token.
| Joined feature range | Source |
|---|---|
| 1–64 | Head 1 message |
| 65–128 | Head 2 message |
| 129–192 | Head 3 message |
Concatenation appends coordinates: (1 × 64), (1 × 64), (1 × 64) → 1 × 192.
For all queries: concatenate H¹, H² and H³ → joined messages (197 rows × 192 features). No learned parameters in this step.
Keep all three messages by placing their coordinates side by side. The joined CLS row has 192 features: 64 from head 1, then 64 from head 2, then 64 from head 3. Concatenation has no learned parameters.
Concatenate heads → joined messages (197 × 192) → Linear(192,192) → Delta (197 × 192).
The joined matrix has 197 rows: CLS plus 196 patches. Each row has 64 + 64 + 64 = 192 features.
U = E + Delta, still 197 × 192.
The skip path carries the original E, before LayerNorm, directly to the addition.
Next: LayerNorm(U) → MLP → add U. The output projection and MLP are different learned layers.
Follow two paths from E. The heads produce messages; concatenate and project them to make ΔE. The skip path carries E directly to addition. U = E + ΔE is the contextualized representation passed to the MLP.
Original CLS: [−0.704, −0.067, …] (1 × 192).
Projected attention update: [1.167, −0.051, …] (1 × 192).
Add coordinate by coordinate: [0.463, −0.118, …] (1 × 192).
Same dog → MLP → later blocks → final CLS → class scores.
Add the projected attention update coordinate by coordinate to the original CLS embedding. This produces a contextualized CLS row of the same width. It continues through the MLP and later blocks before the classifier reads the final summary.
Input E (197 × 192) → block 1 → output E¹ (197 × 192).
- Normalize → attention → add E: obtain U.
- Normalize U → MLP → add U: obtain E¹.
Both branches update all rows. Each CLS or patch row stays 192 features wide.
One block makes two updates. Attention lets rows exchange information; the MLP transforms each resulting row. Each branch adds its update to its input. The block receives and returns 197 rows, each 192 features wide.
| Operation | Shape for CLS | First values |
|---|---|---|
| LayerNorm(u₀) | 1 × 192 | n₁, n₂, … |
| Linear(192,768) | 1 × 768 | [−0.091, 0.248, …] |
| GELU | 1 × 768 | [−0.042, 0.148, …] |
| Linear(768,192) | 1 × 192 | [−0.100, −0.041, …] |
Both linear layers include a learned bias. The same parameters process each row.
Each linear layer connects every input feature to every output unit. GELU transforms each hidden activation. Only a few neurons are drawn; dots omit the rest. The displayed numbers are measured values for this dog’s CLS row.
CLS after attention: [0.463, −0.118, …].
Add the MLP update: [−0.100, −0.041, …].
Block 1 CLS output: [0.363, −0.159, …] (1 × 192).
For all rows: E¹ = U + MLP(LayerNorm(U)), with shape 197 × 192.
Keep the row produced by attention and add the MLP’s update. The result is this row’s output from block 1. Apply the same operation to all 197 rows. Their feature values change; their identities and width stay the same.
E (CLS, P1, …, P196) → block 1 → E¹ → block 2 → E².
Every matrix is 197 × 192. Each block includes attention + residual, then MLP + residual.
Block 2 has its own learned parameters and receives all the updated rows.
The entire output matrix continues to the next block: CLS and all 196 patch rows. Block 2 has its own attention and MLP weights. It applies the same sequence of operations to the feature values produced by block 1.
| Quantity | What happens |
|---|---|
| Features, Q/K/V, weights from softmax, MLP activations | Recomputed from each block’s input |
| Learned parameters | Separate set per block; fixed during the forward pass |
| Rows and output width | CLS + 196 patches; 192 features per row |
Training changes the learned parameters through backpropagation and an optimizer step.
Each block recomputes queries, keys, values, attention weights and MLP activations from its input. Its learned parameters stay fixed during this forward pass. The next block uses a different parameter set. Matrix dimensions and row identities stay the same.
E → block 1 → E¹ → block 2 → E² → … → block 12.
All 12 blocks have their own weights. Every block receives and returns 197 × 192 features.
After block 12: final normalization → select CLS → class scores → one prediction.
Pass the updated features through blocks 1 to 12 in order. The architecture repeats, with separate learned parameters in each block. After block 12 and final normalization, read CLS and calculate one set of class scores for this image.

Block 12 output (197 × 192) → final LayerNorm (197 × 192) → select CLS (1 × 192).
Final CLS begins [−0.144, −0.340, …]. These are image-summary features.
After block 12, all 197 rows are still present. Final LayerNorm keeps the same shape. Select the highlighted CLS row: one image summary with 192 features. Its values now reflect information gathered from this dog’s patches.
Final CLS: 1 × 192 → nn.Linear(192,1000) → class scores: 1 × 1000.
Weight matrix: 192 × 1000 in row-vector math; 1000 × 192 in PyTorch. Bias: 1000.
| Class | Measured score |
|---|---|
| Newfoundland | 15.476 |
| Tibetan mastiff | 11.401 |
| briard | 10.521 |
Every output reads all 192 features. No hidden layer or GELU in this head.
Every class has one output neuron connected to all 192 CLS features. Each neuron uses its own learned weights and bias to produce a score. Only three of the 1,000 class neurons are drawn. Their scores are measured.
Newfoundland: sum all 192 feature × weight products, then add its bias.
−0.0010 −0.0066 + 15.4822 + 0.0017 ≈ 15.4763
The same CLS features feed every class, with separate learned weights and biases.
Multiply each CLS feature by this class’s learned weight. Sum all 192 products and add its bias. For this dog, the Newfoundland neuron produces 15.4763. Every other class performs the same calculation with its own weights and bias.
| Class | Score | Probability |
|---|---|---|
| Newfoundland | 15.476 | 95.73% |
| Tibetan mastiff | 11.401 | 1.63% |
| briard | 10.521 | 0.67% |
Other 997 classes: 1.97% combined. Softmax uses all 1,000 scores.
p(c) = exp(z_c) / sum_j exp(z_j). Select the largest probability.
Exponentiate each score and divide by the sum over all 1,000 classes. Newfoundland receives 95.73% probability. The displayed classes are only three alternatives; the remaining 997 also enter the denominator. Choose the class with the largest probability.
Attention normalizes across 197 source rows for each query and head. The classifier normalizes across 1,000 labels for each image. Both sum to one along their chosen axis; only the second distribution predicts the image label.
No. It means that one source contributes strongly to that query’s value mixture. The class head later scores labels using the final CLS features.
Attention normalizes across 197 source rows for each query and head. The classifier normalizes across 1,000 labels for each image. Both sum to one along their chosen axis; only the second distribution predicts the image label.

Final CLS → Linear(192,1000) → softmax → select the largest probability.
Newfoundland: 95.73% for this photograph.
The classifier reads the final CLS, scores 1,000 labels, and selects Newfoundland as the most probable class. Its probability for this photograph is 95.73%. We have completed one forward pass from pixels to an image prediction.
Optional recap and extra examples
Load the learned CLS for our dog example
| 192-coordinate row | First three coordinates |
|---|---|
| Stored CLS parameter s | [−0.356, −0.038, −0.072, …] |
| + position p₀ | [−0.349, −0.030, −0.074, …] |
| = input e₀ | [−0.704, −0.067, −0.146, …] |
Saved trained parameters, reused for every image. Rounded previews; 189 coordinates are omitted.
The saved model already contains the learned CLS vector and its position vector. Add them to form its starting input row. Load these same parameters for every photograph; no pixel values are needed for this row.
The dog’s patches turn CLS into this image’s summary
The same dog image → 196 patch input rows. Add the shared CLS input row → 197 rows through 12 blocks.
| CLS activation | First three of 192 coordinates |
|---|---|
| At the input | [−0.704, −0.067, −0.146, …] |
| After the blocks and final normalization | [−0.144, −0.340, −1.348, …] |
Final CLS (192 features) → fully connected Linear(192,1000) → 1,000 class scores → softmax → class probabilities.
| Class neuron | Measured score |
|---|---|
| Newfoundland | 15.48 |
| Tibetan mastiff | 11.40 |
Each class neuron reads all 192 features with its own learned weights and bias. Softmax uses all 1,000 scores. Newfoundland has the highest probability: 95.73% for this photograph.
The activation changes with the image. The stored starting CLS parameter remains fixed during inference.
Here is the learned summary for this dog: 192 features read by the class head. Each class neuron combines all 192 features with its learned weights and bias. Softmax gives Newfoundland the highest probability.
Read the two final summaries to classify the photographs
Dog
Final CLS: [−0.144, −0.340, −1.348, …]
Newfoundland · 95.73%
Cat
Final CLS: [−2.999, 5.643, −0.745, …]
Persian cat · 96.71%
Same starting parameter, same 12 blocks, same class head. Different image-dependent final summaries.
Every summary has 192 coordinates. The class head produces 1,000 scores.
After all 12 blocks and final normalization, the two CLS representations differ. The same classifier reads each 192-number summary. It predicts Newfoundland for this dog and Persian cat for this cat.
With CLS, read the dog’s final summary row

196 dog patch rows + one CLS row → 197 × 192.
12 blocks update every row. Read the final CLS → 1 × 192.
Linear(192,1000) → class probabilities → Newfoundland, 95.73% for this photograph.
The dog supplies 196 patch rows. Add the shared CLS row, update all 197 rows through the blocks, then read the final CLS. Its 192 features produce the saved Newfoundland prediction.
Without CLS, combine the dog’s final patch rows

196 dog patch rows → 196 × 192 through the Transformer blocks.
Mean pooling: average final patch features across the 196 rows → 1 × 192.
Linear(192,1000) → class scores. Train with the breed label Newfoundland for a labelled photo like this.
This is a proposed model design, with no measured prediction shown.
Build a second design with 196 patch rows and no CLS. Attention still lets the patches share information. Average their final features to get one 192-number summary. Train the classifier and blocks with this readout.
Mean pooling: average each feature across the patches

Separate small example: four large patches and two final features. These numbers are chosen for calculation.
| Patch row after the blocks | Final features |
|---|---|
| P1 | [2, 0] |
| P2 | [4, 2] |
| P3 | [2, 4] |
| P4 | [0, 2] |
Feature 1: (2 + 4 + 2 + 0) / 4 = 2.
Feature 2: (0 + 2 + 4 + 2) / 4 = 2.
Image summary: [2, 2]. Shape 4 × 2 → 1 × 2.
For the full design: 196 × 192 → 1 × 192. Average final features, after attention.
Each row is a patch representation after it has read context. Average feature 1 across patches, then feature 2. Four rows become one row. In the full model, 196 rows become one 192-feature summary.
Why is CLS a common choice if pooling also works?
- History: BERT’s classification token → original ViT → our checkpoint.
- Role: CLS is a dedicated summary updated through attention; classification loss trains it to support the image label. Patch rows are updated too.
- Evidence: the original ViT study found similar performance for CLS and mean pooling after tuning their learning rates. CLS is not a universal accuracy winner.
Keep the readout used by a pretrained checkpoint. Train or adapt and evaluate a different choice.
ViT paper, Appendix D.3: class token and average pooling.This diagram follows CLS; patch rows receive updates too. Keep the readout a checkpoint was trained to use, or adapt and evaluate the changed model.
Optional reference and extra examples
Create 192 trainable numbers for CLS
Choose the same width as each patch row: 192 features.
Create s = [s₁, s₂, …, s₁₉₂] once, before training.
cls = nn.Parameter(
torch.randn(1, 1, 192) * 0.02
)This illustrative code initializes a trainable parameter with small random values. It does not read pixels.
Shape: one shared batch entry × one token × 192 features. Reuse the same starting parameter across images.
Before training, initialize one 192-number vector. It is stored in the model, like a learned token embedding in text. The numbers do not come from this photograph. Training will adjust them.
The image label teaches the starting CLS numbers
Illustrative training pair: dog photograph with the label Newfoundland.
Patch rows + CLS → blocks → class scores → compare with the known label → loss.
Loss → backward through the model → CLS gradient → optimizer updates its 192 parameters.
The label is used only in the loss. Repeat across many labelled images; other trainable parameters also learn.
The label supplies a loss on the prediction. Backpropagation reaches the starting CLS vector through the classifier and blocks. The optimizer adjusts its 192 parameters, along with other trainable weights, across many labelled images.
Optional whole-model walkthrough
One example: follow the full model to a label loss, then reverse the path to compute gradients and update the parameters.
Connect the previous result to this new question. Keep the same image classifier as the reference.
One example: follow the full model to a label loss, then reverse the path to compute gradients and update the parameters.
Blocks preserve the shape. They change what each row represents.
Prepending CLS changes 196 rows to 197. The head maps the selected 192-feature row to 1,000 logits. Expand the next diagram to see inside every block.
Blocks preserve the shape. They change what each row represents.
Repeat this structure 12 times, with different learned weights in each block. All 197 rows are updated in parallel.
Self-attention mixes value rows. The MLP is applied independently to each row. LayerNorm precedes each branch; residual additions preserve the current representation.
Repeat this structure 12 times, with different learned weights in each block. All 197 rows are updated in parallel.
Reusable block diagram (SVG). The full model reference below expands the attention heads, classifier and label loss.
Label loss → class head and final CLS → blocks 12 through 1 → learned input parameters.
Gradients reach the head, attention and MLP weights, normalization parameters, patch projection, position table and starting CLS. Input image and label remain fixed.
backward(): compute θ.grad. step(): update parameters. For SGD, θ ← θ − η∇θL. Then run a new forward pass.
Reverse the forward dependencies to compute gradients for the trainable parameters. The optimizer uses these gradients to update them. The next forward pass uses the updated model. The image and its label remain fixed.
Optional reference and extra examples
The whole Vision Transformer in one figure
- RGB image (3 × 224 × 224) → 196 flattened patches (196 × 768) → shared projection (196 × 192).
- Prepend learned CLS and add learned positions: E⁰ is 197 × 192.
- Pass all rows through blocks 1–12. Each has separate learned weights.
- Inside each block: LN₁ → three parallel attention heads. Each forms Q, K, V (197 × 64), QKᵀ/√64 (197 × 197), row softmax, then AV (197 × 64). Join the heads, project to width 192, and add E to get U.
- LN₂(U) → Linear(192,768) → GELU → Linear(768,192) → add U.
- After block 12: final LayerNorm → select CLS (1 × 192) → class head (1,000 logits) → softmax → top label.
The known label y enters only at the loss. Here y = Newfoundland, p(y) = 0.95727, and L = −log p(y) = 0.0437.
Forward computes activations and loss. Model parameters remain fixed during this calculation.
Follow the numbered path. The expanded block shows three parallel heads, both residual additions, and the MLP. All 12 blocks keep every row. Final CLS feeds the classifier; the known label enters only at the loss.
Optional CNN comparison
Compare context, mixing, wider views and readouts, then connect those choices to inductive bias and practical data and compute constraints.
Connect the previous result to this new question. Keep the same image classifier as the reference.
Compare context, mixing, wider views and readouts, then connect those choices to inductive bias and practical data and compute constraints.
| Point | CNN | ViT |
|---|---|---|
| 1. Gather context | One output uses a local neighbourhood. | One query can use near and distant patches. |
| 2. Mix information | The learned filter is reused at every location. | Query–key matches determine the source weights for this input. |
First compare the available connections. Then press Next to compare the mixing weights. CNN filters stay fixed during a forward pass. ViT projection parameters also stay fixed, while attention weights are computed from the current query and source features.
First compare the available connections. Then press Next to compare the mixing weights. CNN filters stay fixed during a forward pass. ViT projection parameters also stay fixed, while attention weights are computed from the current query and source features.
| Point | CNN | ViT |
|---|---|---|
| 3. Build a wider view | Successive local layers combine larger regions. | A global attention layer connects distant patches; later blocks refine the features. |
| 4. Read out one label | Often average the final spatial features, then apply a class head. | This model reads final CLS, then applies a class head. Pooling is another option. |
Trace context through the layers, then press Next to inspect the readout. Both models can use the whole image. A classifier needs one image-level vector: CNNs often use spatial pooling; this ViT reads final CLS.
Trace context through the layers, then press Next to inspect the readout. Both models can use the whole image. A classifier needs one image-level vector: CNNs often use spatial pooling; this ViT reads final CLS.
| Comparison | CNN | ViT |
|---|---|---|
| Spatial structure | Local filters, reused across the image | Patches + position; flexible global mixing |
| The shifted stripe | The same filter can detect it elsewhere | Position can change the attention pattern |
| Learned from data | Which local filters help | Which patch relations help |
A CNN builds in local reuse: one learned detector can look for a pattern at many locations. ViT also has structure—patches, shared projections and position signals—but gives attention more freedom to learn spatial relationships.
A CNN builds in local reuse: one learned detector can look for a pattern at many locations. ViT also has structure—patches, shared projections and position signals—but gives attention more freedom to learn spatial relationships.
| Situation | CNN starting point | ViT starting point |
|---|---|---|
| Few labels; train from scratch | Useful baseline: built-in local reuse | Pay close attention to the training recipe |
| Few labels; pretrained weights | Fine-tune a suitable pretrained CNN | Fine-tune a suitable pretrained ViT |
| Phone or tight latency budget | Try a compact CNN; measure on the device | Try an efficient ViT; measure on the device |
| Distant clues across the image | Use enough depth for broad context | Global attention links distant patches directly |
Treat these as starting points, not a ranking. Pretraining can matter more than the architecture name on a small dataset. Compare suitable models using the same validation set and measure speed and memory on the target hardware.
Treat these as starting points, not a ranking. Pretraining can matter more than the architecture name on a small dataset. Compare suitable models using the same validation set and measure speed and memory on the target hardware.
Implementation lab · optional
All implementation details are retained. Every operation is paired with shapes, diagrams or a numerical equivalence check.
Exactly the same affine map with the same weights and biases, up to floating-point roundoff.
All implementation details are retained. Every operation is paired with shapes, diagrams or a numerical equivalence check.
# rgb: (224, 224, 3), RGB per pixel
x = rgb.permute(2, 0, 1) # 3, 224, 224
x = x.unsqueeze(0) # 1, 3, 224, 224
# Patch order: all R, then G, then BThe leading axis B. One photograph uses B=1; a batch of photographs uses the same model parameters for each.
The prepared photograph is 224×224 with three RGB channels. PyTorch puts channels before height and width, then adds a batch axis. B counts images; it does not count patches.
Each RGB patch supplies 768 numbers. One shared projection turns it into 192 features. Reusing that projection across the 14×14 grid produces 196 rows. We will implement this exact operation in two ways.
No. Each patch supplies different inputs to the same learned projection. The same parameters also serve the cat and every other image.
Each RGB patch supplies 768 numbers. One shared projection turns it into 192 features. Reusing that projection across the 14×14 grid produces 196 rows. We will implement this exact operation in two ways.
| Quantity | Count |
|---|---|
| Inputs per patch | 3×16×16 = 768 |
| Output neurons | 192 |
| Weights | 192×768 = 147,456 |
| Biases | 192 |
| Total | 147,648 |
c₁ = w₁₁x₁ + … + w₁₇₆₈x₇₆₈ + b₁. This projection has no hidden layer or activation.
All 768 inputs connect to every output neuron. Each neuron computes a weighted sum and adds its own bias. The patch projection is one dense layer; the later Transformer MLP has two layers and a nonlinear activation.
linear = nn.Linear(768, 192, bias=True)
def project_with_linear(images, linear):
patches = F.unfold(images, kernel_size=16, stride=16)
# B × 768 × 196; each patch: all R, then G, then B
# Within each channel: left to right, top to bottom
patches = patches.transpose(1, 2)
# B × 196 × 768: one flattened patch per row
return linear(patches) # B × 196 × 192Unfold extracts each non-overlapping patch as a column. Transpose makes patches into rows. Linear changes only the last axis, from 768 inputs to 192 features, reusing its weights for every patch and image.
conv = nn.Conv2d(3, 192, kernel_size=16, stride=16,
padding=0, bias=True)
def project_with_conv(images, conv):
grid = conv(images) # B × 192 × 14 × 14
return grid.flatten(2).transpose(1, 2) # B × 196 × 192Each of 192 filters has 3×16×16 weights and one bias. Output: B×192×14×14; flatten and transpose to B×196×192.
Reshape each neuron’s 768 weights into a 3×16×16 filter. It reads the same pixels and adds the same bias. Conv2d applies all 192 filters at every patch location, giving a 192-channel feature grid.
| Parameter | Linear | Conv2d |
|---|---|---|
| Weight shape | 192 × 768 | 192 × 3 × 16 × 16 |
| Weights | 147,456 | 147,456 |
| Biases | 192 | 192 |
| Total | 147,648 | 147,648 |
with torch.no_grad():
linear.weight.copy_(conv.weight.flatten(1))
linear.bias.copy_(conv.bias)Both implementations have 147,456 weights and 192 biases. Copying the exact weights and biases makes their computations equivalent. More patches or images create more output values, while the learned parameter count stays 147,648.
Flatten RGB channels in order, row-major within each channel. Green [2,3] has flat index 291.
Linear.weight[0,291] = Conv2d.weight[0,1,2,3]. Both multiply the same input value by this same learned weight.
Both paths pair the same 768 pixels with the same 768 weights, sum their products, and add the same bias.
| Image | Linear P63 preview | Conv2d P63 preview |
|---|---|---|
| Dog | [-0.8515409827232361, 1.338600516319275, 0.5036810636520386] | [-0.8515405654907227, 1.3386002779006958, 0.5036811232566833] |
| Cat | [-0.1888497918844223, 0.10422777384519577, -0.20799635350704193] | [-0.18884989619255066, 0.1042276993393898, -0.207997128367424] |
dense_rows = project_with_linear(images, linear)
conv_rows = project_with_conv(images, conv)
torch.testing.assert_close(
dense_rows, conv_rows, rtol=1e-5, atol=2e-5)
def project_with_conv(images, conv):
grid = conv(images) # B × 192 × 14 × 14
return grid.flatten(2).transpose(1, 2) # B × 196 × 192Using identical pretrained weights, both paths produce the same dog and cat features within floating-point rounding. Conv2d packages patch extraction and projection into one image operation. Flatten and transpose its grid before adding CLS and positions.
dense_rows = project_with_linear(images, linear)
conv_rows = project_with_conv(images, conv)
torch.testing.assert_close(
dense_rows, conv_rows, rtol=1e-5, atol=2e-5) Runnable dog/cat equivalence example · Measured outputs and checks. PyTorch Linear · PyTorch Conv2d · PyTorch Unfold. Conv2d accepts the image layout directly and handles the repeated patch computations. It avoids an explicit unfolded patch tensor in our code and uses optimized convolution backends. Actual speed and memory depend on the hardware.
No. The same 147,648 parameters compute the same outputs. The advantage is a convenient image operation and backend support; benchmark actual speed and memory.
Conv2d accepts the image layout directly and handles the repeated patch computations. It avoids an explicit unfolded patch tensor in our code and uses optimized convolution backends. Actual speed and memory depend on the hardware.
# x: (B, 3, 224, 224)
B = x.shape[0] # integer: number of images
grid = self.patch(x) # (B, 192, 14, 14)
rows = grid.flatten(2) # (B, 192, 196)
rows = rows.transpose(1, 2) # (B, 196, 192)No. flatten(2) starts at axis 2. The batch and 192 feature channels remain separate.
Flatten combines the 14×14 spatial grid into 196 locations. Transpose puts the 192 features last. Each row now describes one patch.
def embed(self, x):
# Continue with rows: (B, 196, 192)
cls = self.cls.expand(B, -1, -1) # (B, 1, 192)
rows = torch.cat([cls, rows], dim=1) # (B, 197, 192)
return rows + self.pos # (B, 197, 192)
# self.cls: (1, 1, 192) self.pos: (1, 197, 192)torch.cat adds CLS along dim=1. Adding positions leaves the shape unchanged.
Expand reuses the learned CLS start across the batch. Concatenation adds one row. Position addition changes the numbers while keeping 197 rows and 192 features.
B, N, D = x.shape # B, 197, 192
qkv = self.qkv(x) # B, 197, 576
qkv = qkv.reshape(B, N, 3, 3, 64) # B, N, QKV, heads, features
q, k, v = qkv.permute(2, 0, 3, 1, 4).unbind(0)
# q, k, v: each (B, 3, 197, 64)One selects Q, K or V. The other selects the attention head. B always remains a separate image axis.
Project each row once to produce all queries, keys and values. Split the 576 outputs into three roles, each with three heads of width 64.
scores = (q @ k.transpose(-2, -1)) / 8 # B, 3, 197, 197
weights = scores.softmax(dim=-1) # B, 3, 197, 197
messages = weights @ v # B, 3, 197, 64The last axis: source keys. Every receiving query gets its own distribution.
For each image and head, every query scores all 197 sources. Softmax normalizes each query row. Multiplying by V combines the sources into one 64-feature message per query.
def forward(self, x):
joined = messages.transpose(1, 2) # B, 197, 3, 64
joined = joined.reshape(B, N, D) # B, 197, 192
return self.proj(joined) # B, 197, 192Neither. Each image and query keeps its own three head messages. Only head features are joined.
Move the head axis beside its features, then join 3×64 into 192. A learned output projection mixes these features before the residual addition.
self.norm1 = nn.LayerNorm(192, eps=1e-6)
self.attn = Attention()
self.norm2 = nn.LayerNorm(192, eps=1e-6)
self.mlp = nn.Sequential(
nn.Linear(192, 768), nn.GELU(),
nn.Linear(768, 192))No. Only the feature axis expands. B and N stay unchanged.
Create two normalization layers, one attention module and one MLP. The MLP expands and then restores the feature width for each row.
def forward(self, x):
# x: (B, 197, 192)
x = x + self.attn(self.norm1(x))
x = x + self.mlp(self.norm2(x))
return xThe x already updated by attention. Each branch returns the same shape as the input it is added to.
Normalize, compute the attention update, and add the input. Then normalize the updated rows, compute the MLP update, and add them again.
self.blocks = nn.ModuleList([Block() for _ in range(12)])
self.norm = nn.LayerNorm(192, eps=1e-6)
self.head = nn.Linear(192, num_classes)No. The list comprehension creates twelve separate Block objects.
ModuleList creates twelve distinct blocks. The final normalization and class head read the result of the entire stack.
def forward(self, x):
rows = self.embed(x) # B, 197, 192
for block in self.blocks:
rows = block(rows) # B, 197, 192
summary = self.norm(rows)[:, 0] # B, 192: final CLS
return self.head(summary) # B, 1000 logitsCross-entropy accepts logits directly. During inference, logits.softmax(-1) converts scores into class probabilities.
Feed each block’s output into the next. Normalize the final rows and select CLS for each image. The class head returns 1,000 scores.
model.train()
optimizer.zero_grad(set_to_none=True)
logits = model(images) # B, 1000
loss = F.cross_entropy(logits, labels) # labels: B
loss.backward()
optimizer.step()No. It already includes log-softmax. labels is a length-B integer tensor; the model returns B×1000 logits for this example.
Cross-entropy consumes logits and class-index labels. Backward computes gradients; the optimizer updates the parameters it owns. Clear previous gradients before the next batch. This code shows the procedure without claiming a trained result.
Optional extensions: transfer and evaluation
These experiments extend the core ImageNet classification story.
These experiments extend the core ImageNet classification story.
These experiments extend the core ImageNet classification story.
Choose trainable parameters, follow one batch, then evaluate on held-out images. We will describe the procedure without running a new experiment.
Connect the previous result to this new question. Keep the same image classifier as the reference.
Choose trainable parameters, follow one batch, then evaluate on held-out images. We will describe the procedure without running a new experiment.
The checkpoint was pretrained on ImageNet-21k and fine-tuned on ImageNet-1k. Its current head predicts 1,000 categories. Newfoundland is one of those labels.
No. It ran inference with the existing ImageNet checkpoint. Its learned encoder and 1,000-class head were already available.
The checkpoint was pretrained on ImageNet-21k and fine-tuned on ImageNet-1k. Its current head predicts 1,000 categories. Newfoundland is one of those labels.
Our pet application groups breeds into two categories. We want two scores for every image, trained and evaluated for this dog/cat task. Start by reusing the encoder’s visual features.
Yes. The encoder already represents visual patterns. A new head can learn how those features separate the two target classes.
Our pet application groups breeds into two categories. We want two scores for every image, trained and evaluated for this dog/cat task. Start by reusing the encoder’s visual features.
The labels can stay dog and cat while the image domain changes. A photo classifier may struggle with sketches. Evaluate on the target domain, then compare head training with encoder fine-tuning.
No. The two labels are unchanged. The features may need adaptation because sketches remove colour and texture cues.
The labels can stay dog and cat while the image domain changes. A photo classifier may struggle with sketches. Evaluate on the target domain, then compare head training with encoder fine-tuning.
Pet image → pretrained ViT → CLS features (192) → Linear(192,2) → dog/cat scores.
192×2+2=386 new trainable parameters. This is a proposed workflow, with no claimed benchmark results.
For our dog/cat task, replace the ImageNet head with a new two-output linear layer. The encoder supplies a 192-number image representation. The new 386 head parameters must learn from labelled pet images.
Eye: vision encoder. Snowflake: fixed weights. Flame: trainable weights.
| Mode | Encoder | Head |
|---|---|---|
| Head only | All encoder weights frozen; 192 image-dependent features | 192 inputs → dog and cat scores; 386 trainable parameters |
Snowflake: weights stay fixed. Flame: weights can learn. Train the two-class head on labelled pet images. The frozen encoder still computes different features for different images.
Unfreeze the last block and train it together with the head. Earlier encoder weights stay fixed. This lets some visual features adapt to the target data; validation determines whether it helps.
When might a head alone be insufficient?
Unfreeze the last block and train it together with the head. Earlier encoder weights stay fixed. This lets some visual features adapt to the target data; validation determines whether it helps.
Images B×3×224×224 → logits B×2; labels B → mean loss.
zero_grad → forward → loss → backward → step.
Attention stays within each image.
Pair every image with its dog/cat label. Compute the average loss for the batch, backpropagate once, and update the trainable parameters. Repeat this procedure over training batches.
- Pretrain: many labelled images → learn encoder and classifier weights.
- Adapt: target images and labels → update the selected weights for the new task.
- Infer: a new image → run the trained model → predict a label with weights fixed.
Orange connections and flames mark learning. Blue connections and snowflakes mark fixed parameters. During inference, a new image changes the features while the stored weights stay fixed.
model.eval()
with torch.inference_mode():
probabilities = model(images).softmax(dim=-1) # B, 2The image produces its own CLS features and class scores. Both snowflakes mark fixed weights. Inference computes a prediction without a label, loss or optimizer update.
model.eval()
with torch.inference_mode():
probabilities = model(images).softmax(dim=-1) # B, 2| Split | Use |
|---|---|
| Train | Change weights |
| Validation | Select settings/checkpoints |
| Test | Evaluate the selected procedure |
Inspect a 2×2 confusion table and representative failures. No new measured results are claimed.
Use held-out images to measure accuracy and the two kinds of confusion. Inspect successes and mistakes, including confident errors. A probability on one photograph is not the classifier’s test accuracy.
Optional extensions: new-image inference
Return to the saved dog and cat predictions. These outputs come from the ImageNet checkpoint, not from a newly trained pet classifier.
Connect the previous result to this new question. Keep the same image classifier as the reference.
Return to the saved dog and cat predictions. These outputs come from the ImageNet checkpoint, not from a newly trained pet classifier.
The model receives the square crop on the right, after RGB normalization.
| ImageNet label | Probability |
|---|---|
| Newfoundland | 95.73% |
| Tibetan mastiff | 1.63% |
| briard | 0.67% |
These probabilities come from running this photograph through the pretrained model.
Try the same learned model on another photograph
Persian_98: the top prediction is Persian cat, with probability 0.9671.
Reproduce real-image inference · Inspect every worksheet parameter and calculation · Open the worked notebook.
Keep the earlier lessons beside this one
Part I: representations, scores, probabilities, and loss. Part II: queries/keys choose weights; values supply the messages. Part III: separate head messages, concatenation, output projection, and residual.
Images and teaching references
Photographs: Oxford-IIIT Pet dataset, Parkhi, Vedaldi, Zisserman and Jawahar, via the timm mirror. Original files newfoundland_31 and Persian_98, test partition. Dataset and image license: CC BY-SA 4.0. Original copyright remains with the image owners. The cropped views above retain this attribution.
Explanatory references: Stanford CS231n, Lecture 8, Jay Alammar, D2L, UvA, and the original ViT paper. The opening reproduces the original Transformer and ViT paper figures with attribution. Other teaching diagrams and the four-patch calculation are original to this series.
To estimate accuracy, count correct predictions over an appropriate test set.
Keep the same model and preprocessing. Change only the photograph.
Optional extensions: interpreting the model
Read attention maps, then cover parts of the photograph and measure how the prediction changes.
Connect the previous result to this new question. Keep the same image classifier as the reference.
Read attention maps, then cover parts of the photograph and measure how the prediction changes.
The selected ear-side patch matches patches on the other side of the dog. This measures representation similarity, rather than which values attention mixes.
No. This compares contextual patch features after block 12 using cosine similarity. P82 is a measured high-similarity patch; a correspondence is not a guaranteed segmentation.
The selected ear-side patch matches patches on the other side of the dog. This measures representation similarity, rather than which values attention mixes.
Saved pretrained ViT-Tiny, block 12, query P74, selected source P82. Gold indicates positive cosine similarity; blue indicates negative similarity. Self-similarity is 1. Nine guided examples and free exploration.
Teal shows source weights. Both maps use the same 0–10.35% scale. Within each head, all 197 source weights—including CLS—sum to 100%.
Each head has its own learned query and key projections. These lead to different weights on the source value rows. Head 2’s weaker peak is shown on the same color scale.
Teal shows source weights. Both maps use the same 0–10.35% scale. Within each head, all 197 source weights—including CLS—sum to 100%.
The two maps are calculated from saved Q and K in block 4. Each softmax includes CLS and all 196 patches; the displayed patch values are not renormalized. Change the head or block in the interactive lab.
The query is CLS. Its weights select a mixture of all source value rows. Later operations and the classifier turn the final CLS features into class scores.
No. It is the last block’s first attention head. P64 receives 24.86% of this query’s source weight; CLS itself receives 0.08%. Other heads, residuals, MLPs and the class head also contribute.
The query is CLS. Its weights select a mixture of all source value rows. Later operations and the classifier turn the final CLS features into class scores.
This is measured attention, not a segmentation mask or a causal importance map. Open all nine guided examples.
Original normalized input: (1, 3, 224, 224). Copy it and set the first 112 rows and columns, in all three channels, to zero. This is mid-gray RGB; image size and patch count stay the same.
x = transform(photo).unsqueeze(0) # (1, 3, 224, 224)
covered = x.clone() # a separate copy
covered[:, :, :112, :112] = 0 # all RGB channels, top leftCopy the normalized input, then overwrite one quadrant. The rest of the image stays unchanged; no patch rows are deleted.
The original RGB photo is loaded with Pillow. The checkpoint’s evaluation transform resizes, center-crops to 224×224 and normalizes each channel as (pixel − 0.5) / 0.5. unsqueeze(0) adds the batch axis. clone() makes a separate tensor so the original is preserved.
The four indices are batch, channel, row, column. Both colons select the whole batch and all RGB channels; :112 selects indices 0 through 111 in each spatial axis. This changes 112×112×3 input values. The 14×14 patch grid remains: 49 of its patches now contain constant gray pixels. Their projected features still include the learned projection bias and position information.
This is an input-pixel intervention. It does not remove tokens or apply a mask to the attention matrix. The diagram draws the gray cover on the exact model crop; the experiment changes the tensor before running the model.
x = transform(photo).unsqueeze(0) # (1, 3, 224, 224)
covered = x.clone() # a separate copy
covered[:, :, :112, :112] = 0 # all RGB channels, top leftOriginal → same trained ViT → P(Newfoundland) = 95.73%. Covered copy → same trained ViT → 83.00%. All activations are recomputed; weights remain fixed.
model.eval()
with torch.inference_mode():
p_before = model(x).softmax(-1)[0] # (1000,)
p_after = model(covered).softmax(-1)[0] # (1000,)
before, after = p_before[256], p_after[256] # NewfoundlandRecompute patch features, attention, CLS and class scores. Softmax gives 1,000 probabilities; compare the Newfoundland entry in both runs. The weights stay fixed.
Both inputs use the same pretrained checkpoint in evaluation mode. model.eval() selects evaluation behaviour; torch.inference_mode() avoids recording gradients. Neither call trains the model. Each model call recomputes patch embeddings, all 12 blocks, final CLS and the 1,000 class scores. The CLS start vector, position embeddings and every learned weight remain fixed between runs.
softmax(-1) converts the 1×1,000 logits to class probabilities; [0] selects the only image in the batch. Index 256 is Newfoundland in this checkpoint’s ImageNet label order. Keep that index fixed when comparing probabilities, even if an intervention changes the top prediction.
model.eval()
with torch.inference_mode():
p_before = model(x).softmax(-1)[0] # (1000,)
p_after = model(covered).softmax(-1)[0] # (1000,)
before, after = p_before[256], p_after[256] # NewfoundlandThe results are saved measurements, not a live browser inference. Reproduce the preprocessing and both model calls · Saved probabilities and metadata. This measures sensitivity to this particular gray replacement; it does not assign a unique importance to the removed visual content.
Top-right covering causes the largest drop: 16.11 percentage points. All four images still predict Newfoundland. Attention shows internal mixing; occlusion tests how a changed input changes the prediction.
The top-right cover gives the largest drop among these four tests. Keep the target class, fill value and model fixed. This does not isolate a semantic object part; the next slide uses smaller covers.
Top-right covering causes the largest drop: 16.11 percentage points. All four images still predict Newfoundland. Attention shows internal mixing; occlusion tests how a changed input changes the prediction.
Each trial starts from the original normalized input and independently replaces one 112×112 quadrant with zero (gray RGB). Saved measurements. Probability drops are in percentage points and should not be added.
Large cover: 112×112 pixels, 49 model patches, four tests. Small cover: 16×16 pixels, one model patch, 196 tests. Each test starts with the original image and uses the same model.
Only the cover size changes. The trained model still uses 16×16 patches and a 224×224 input. Every test starts from the original image.
The original experiment used four non-overlapping 112×112 covers. The finer experiment uses 196 non-overlapping 16×16 covers aligned with the existing 14×14 patch grid. The model’s patch projection, weights, input resolution and preprocessing stay fixed.
Each trial starts with a fresh copy. Replace one patch with normalized zero, rerun the whole model and compute 100 × (original probability − covered probability). The covers are never accumulated. A small change does not show that a region is useless: other regions can carry related information, and features can interact. Drops from different trials should not be added together.
Among 196 separate tests, covering P78 gives the largest drop: 95.73% → 91.89%, or 3.84 percentage points. All top labels remain Newfoundland. Red marks a probability decrease; blue marks an increase.
Smaller covers localize sensitivity. P78 causes the largest drop here, but no single-patch cover changes the top label. This map measures probability changes, not attention weights.
This is a new saved experiment on the same trained checkpoint and photograph. The largest drop is at P78, grid row 6, column 8 (counting from one). All 196 interventions retain Newfoundland as the top label. The target probabilities range from 91.89% to 96.05%. Three replacements slightly increase the target probability, so a cover need not always reduce it.
Red means a positive drop, blue a negative drop, with a fixed symmetric −4 to +4 percentage-point color scale. The outlined patch is the same location in the one-test image and in the result map. This is neither a similarity map nor an attention map: it summarizes changes in the final class probability.
The result depends on this image, target, model, cover size and gray fill. It locates sensitivity to these replacements, not an object segmentation or a complete explanation of recognition. All measurements and verification · Reproduce the experiment.
Optional reference and extra examples
Explore a trained ViT, one example at a time
Open the nine-example interactive lab
1 · Locate the purple patch
2 · Where does this query read?
Strongest patch sources
3 · Inspect one connection
Follow nine guided examples on the trained model. Each preset shows what to notice and one takeaway. Then choose Free exploration to select any patch, block, head or view. Each square is a 16×16 patch.
Start here: keep the preset fixed, locate the purple query, read the gold similarity or teal attention map, then discuss the takeaway. Next example changes the settings for you. Free exploration exposes the full controls; Guided examples returns to the last preset.
| Example | Settings | Notice | Takeaway |
|---|---|---|---|
| 1. Background finds background | Feature similarity · block 12 · P182 | The background query finds gold around the dog. | Learned features can separate foreground and background. Purple marks the lower-right background patch. Gold appears around the dog, with less on its body. This is a similarity pattern, not a predicted segmentation mask. |
| 2. Check a true corner | Feature similarity · block 12 · P1 | The top-left corner matches the other three corners. | A strong match need not be an object part. The strongest matches have cosine similarity almost 1. Inspect where the matches occur before assigning them a meaning. |
| 3. One ear finds the other side | Feature similarity · block 12 · P74 | The ear query matches the other side of the head. | Similar features can connect distant patches. Purple marks P74 on the left side of the dog’s head. P82 on the opposite side is the strongest other match. This correspondence is specific to this trained model and photograph; it is not guaranteed for every image. |
| 4. Move the reference to the nose | Feature similarity · block 12 · P63 | The nose query finds nearby face and muzzle patches. | Change the query; change the matches. The query moves to P63 near the nose. The map compares every patch representation with this selected reference. |
| 5. Follow the body’s dark fur | Feature similarity · block 12 · P147 | The chest query finds similar fur lower on the dog. | One object can contain several kinds of features. The reference is now on the chest. Similarity need not highlight the entire dog uniformly: the lower body and nose have different features. |
| 6. Rewind to block 1 | Feature similarity · block 1 · P74 | Same ear pair: 0.322 here → 0.902 at block 12. | Same pixels. Different features after each block. Keep the P74-to-P82 ear pair from example 3. Its cosine similarity is 0.322 after block 1 and 0.902 after block 12. The representation changes while the pixels stay fixed; not every pair’s similarity must increase. |
| 7. Ask what the ear reads | Attention, head 1 · block 4 · P74 | The ear query gives P60’s value 9.09% weight. | Attention mixes values into the query’s message. This is block 4, head 1, query P74. An attention weight scales a source’s value; feature similarity instead compares the patch representations. These are different measurements. |
| 8. Change only the attention head | Attention, head 2 · block 4 · P74 | Same query and block. Head 2 favours P38. | Different heads gather different information. Head 2 puts its largest patch weight on P38, at about 2.29%. Teal is rescaled within each attention map: compare percentages, not brightness, across heads. |
| 9. Let CLS gather an image summary | Attention, head 1 · block 12 · CLS | CLS gives the face patch P64 about 25% weight. | CLS gathers image information for classification. CLS is the query: an extra token, not an image patch. P64 receives about 24.86% of the weight in block 12, head 1. This is one head in one block, not a complete explanation of the final label. |
These are saved activations from the same pretrained ViT-Tiny classifier used throughout this lecture. It predicts Newfoundland for this image with 95.73% probability. No training runs in the browser. One block loads at a time; the browser computes the selected attention row from saved Q and K.
Attention: one head’s query compares with all 197 keys. Softmax produces 196 patch weights plus one CLS weight. The map shows the patch weights without renormalizing them; the CLS weight is reported separately. Teal contrast adapts to each map, so compare numeric weights across blocks. Clicking a source shows the score, weight, and where its value vector enters the weighted sum.
Feature similarity: compare the 192-dimensional patch representations after the selected block’s attention, MLP, and residual additions, before the next normalization. Cosine uses a fixed −1 to 1 scale. The selected patch’s self-match is omitted from the top-three list. This uses the DINO-style interaction idea with our supervised classifier; DINO learns its features with a different, self-supervised objective. Neither view establishes which pixels causally determined the class. The next slides intervene on the image to ask a different question.
Keyboard: tab to either image grid, use arrow keys to move between patches, then Enter or Space to select. Play blocks runs once from block 1 to 12 and stops when you leave the slide. Data provenance and reproduction.
What happens if we cover the top right?
Replace this quadrant with gray pixels, then rerun the same trained model. Compare the Newfoundland probability with the original image.
What happens if we cover the bottom left?
Replace this quadrant with gray pixels, then rerun the same trained model. Compare the Newfoundland probability with the original image.
What happens if we cover the bottom right?
Replace this quadrant with gray pixels, then rerun the same trained model. Compare the Newfoundland probability with the original image.
Optional cost exercises and references
Halve the patch width and height. Predict the change in attention work before calculating it.
Connect the previous result to this new question. Keep the same image classifier as the reference.
Halve the patch width and height. Predict the change in attention work before calculating it.
At fixed image size, half the patch width gives four times as many patch rows.
Even this small ViT compares many pairs of rows.
Work out the token count before you change the controls.
224×224 RGB → shared 768-to-192 patch projection → 196 patch rows → add CLS and positions → 197×192 → twelve blocks → final LayerNorm → CLS (192) → Linear head → 1,000 class scores.
Each block: LayerNorm → three attention heads → join and project → residual addition; LayerNorm → MLP → residual addition. All rows update.
One shared patch layer builds the rows. Twelve blocks refine every row. The classifier reads final CLS to score image labels.
This is the same trained tiny ViT throughout the lecture: 16×16 RGB patches, 196 patch rows, 192 features, three 64-wide attention heads per block, and twelve blocks. LN means LayerNorm. The shared patch projection maps 768 input values to 192 features; Conv2d with kernel and stride 16 implements the same affine map as a shared Linear layer on flattened patches.
Within a head, softmax(QKᵀ/√64)V gives one message per query. Concatenation and the output projection make a 192-wide update for every row. Each plus sign adds that update to the incoming row. The MLP transforms each row separately. Final LayerNorm precedes selection of CLS.
The original detailed reference figure, including Q/K/V, shapes, class probabilities and the label loss, remains in the optional whole-model reference. The top path is the full model; the lower panel is a zoom of one block, not an additional block.
- Content + location: shared patch projection plus learned position.
- Every token gains context through attention and is transformed by the MLP.
- Image-label loss trains useful representations; inference keeps weights fixed.
- At patch size 16, increasing the image from 224 to 448 gives 197 to 785 tokens: about sixteen times as many attention scores.
Pixels supply content. Position supplies location. Attention builds context. Training makes the final representation useful for the task.
The token diagram is schematic: the real model updates CLS and all 196 patches. The loss diagram summarizes backpropagation through the head and all preceding operations. CLS is a learned readout choice; mean pooling is another valid choice when the model is trained for it.
At fixed 16×16 patch size, 224×224 gives 197 tokens with CLS and 448×448 gives 785. The score counts per head and block are 38,809 and 616,225, a factor of 15.88. The matrix icons are schematic. This counts dense attention scores, not total model computation or the memory allocated by an optimized attention kernel. Adapting this checkpoint to a larger grid also requires adapting its positional embeddings.
Optional practice and next steps
The main lecture ends above. Use these questions to check your understanding in the reading view.
Explain one classifier from pixels to learning
- Count patch and CLS rows.
- Find one head’s width.
- Count two-class head parameters.
- Trace the backward path.
- Explain the no-CLS alternative.
Use the architecture diagram to answer each question. Name the operation, its input and output shapes, and how its parameters receive a learning signal. Explain a mean-pooling alternative to CLS.
Your turn: work out the message
Find the weights first. Then multiply each value by its own weight.
Your turn: trace every important shape
Keep patch count, embedding width and per-head width separate.
Layout check: what if we rearranged the patches?
Original: face in row 2
Rearranged: face in row 4
This is a thought experiment, not a step in our classifier. Keep each crop intact but change its location. Which part of its input row should change? The animal label need not change.
Does the patch layer notice the move?
The identical face crop moves from row 2, column 2 to row 4, column 2.
Same pixels → same W and b → same content vector c_face.
Its new location has not entered that calculation.
The same pixels pass through the same linear layer, so they produce the same content embedding. The layer has not been given the patch’s location.
Give the row its location as well as its content
Before: c_face + p₂,₂ = e_before.
After: c_face + p₄,₂ = e_after.
The same content is paired with its current location. Each vector has D coordinates; p₂,₂ names a slot, not two feature values.
Each grid slot has a position vector. Add it to the content embedding before attention, just as we added position to the text embeddings.
Did we move the image, or just reorder its rows?
Follow one patch and its position vector in each experiment.
Next: classify an image by writing the labels
Current lecture: image → representation → trained class head → fixed label choices.
Next: compare image features with features computed from candidate text descriptions.
We now understand a complete image classifier. CLIP will connect its image representation with text representations, so we can compare a photograph with candidate descriptions.
A class score is a dot product with its learned weight vector, plus a bias. The current head has 1,000 fixed label slots. New tasks can train a replacement head.
It is a learned row of the classifier weight matrix. The class name itself is not read by this classifier.
A class score is a dot product with its learned weight vector, plus a bias. The current head has 1,000 fixed label slots. New tasks can train a replacement head.
CLIP trains image and text representations to match. Learned projections and unit normalization make their vectors comparable; prompts can then describe candidate classes.
Not directly. Image and text encoders need compatible projections and joint alignment training. CLIP compares normalized image and text vectors, with a learned score scale.
CLIP trains image and text representations to match. Learned projections and unit normalization make their vectors comparable; prompts can then describe candidate classes.
The next lecture uses the same branches and notation: hᵢ → Wᵢ → normalized u; hₜ → Wₜ → normalized v; then u · v. CLIP also learns a score scale for its training objective.
Optional reference diagrams
Follow the numbered path. The expanded block shows three parallel heads, both residual additions, and the MLP. All 12 blocks keep every row. Final CLS feeds the classifier; the known label enters only at the loss.
Start at pixels, then patch projection, CLS and positions. Trace all 12 blocks. Open block 1: normalize, three parallel Q/K/V → scores → row softmax → AV lanes, concatenate, output projection, add the original E. Normalize U, apply Linear → GELU → Linear, and add U. Continue through the remaining distinct blocks, final normalization, final CLS and the class head. Softmax gives the prediction; logits and the known label give cross-entropy.
Follow the numbered path. The expanded block shows three parallel heads, both residual additions, and the MLP. All 12 blocks keep every row. Final CLS feeds the classifier; the known label enters only at the loss.
Image patches become tokens. The encoder builds a summary for classification.
Read this figure from the image upward. Reveal the patch projection, the encoder, and the classification readout in that order. Point to the expanded encoder on the right. Then begin our dog-photo example.
Image patches become tokens. The encoder builds a summary for classification.
Before training, initialize one 192-number vector. It is stored in the model, like a learned token embedding in text. The numbers do not come from this photograph. Training will adjust them.
The chosen embedding width is 192. The initialization routine supplies the numbers once, before training; no image pixels are needed.
Before training, initialize one 192-number vector. It is stored in the model, like a learned token embedding in text. The numbers do not come from this photograph. Training will adjust them.
The label supplies a loss on the prediction. Backpropagation reaches the starting CLS vector through the classifier and blocks. The optimizer adjusts its 192 parameters, along with other trainable weights, across many labelled images.
No. The image’s class label provides the loss; the chain rule supplies gradients for the starting vector.
The label supplies a loss on the prediction. Backpropagation reaches the starting CLS vector through the classifier and blocks. The optimizer adjusts its 192 parameters, along with other trainable weights, across many labelled images.