Deep learning · Transformer architectures
From Attention to Applications
Read the diagram labels
- From attention mechanisms to three Transformer architectures and their tasks.
- Encoder, decoder-only and encoder-decoder Transformers
- ENCODER
- Read the complete input
- Token and sentence labels
- DECODER-ONLY
- Read the available prefix
- Next-token generation
- ENCODER-DECODER
- Read source and target prefix
- Conditional generation
Authored teaching diagram
How do attention blocks become models for different tasks?
We will use familiar attention blocks to construct decoder-only, encoder and encoder-decoder models, then examine how cross-attention connects a target sequence to its source.
Teaching note
How do attention blocks become models for different tasks? Introduce the lecture, then recall the attention mechanisms studied previously.
Recap · Self-attention, multiple heads and positions
Recap: the Transformer components we have studied
Read the diagram labels
- 1 · Positional encoding. 2 · Self-attention. 3 · Multiple heads. token embedding + position. ei + pi. Token identity and order. Q, K, V from embeddings. scores → softmax → mi. A weighted sum of values. head 1. head 2. head 3. Concatenate messages. Project with WO → Δei. Residual update: ei ← ei + Δei. Then the MLP and its residual update. Result: one contextualised embedding at every token position
- 1 · Positional encoding
- 2 · Self-attention
- 3 · Multiple heads
- token embedding + position
- e
i + pi - Token identity and order
- Q, K, V from embeddings
- scores → softmax → m
i - A weighted sum of values
- head 1
- head 2
- head 3
- Concatenate messages
- Project with W
O → Δei - Residual update: e
i ← ei + Δei - Then the MLP and its residual update
- Result: one contextualised embedding at every token position
Authored teaching diagram · Primary source
What does each attention block change?
Positional encoding supplies order information to token embeddings. Self-attention forms queries, keys and values from the same sequence and combines values into a message for each position. Multiple heads use separate learned projections; their messages are concatenated and projected into an embedding update. Residual addition and the MLP produce contextualised embeddings.
Teaching note
What does each attention block change? Recall the additive positional encoding studied earlier: eᵢ + pᵢ. The eᵢ in the residual equation denotes the current embedding entering that sublayer. Each head uses its own Q/K/V projections. The diagram uses the existing message mᵢ and update Δeᵢ convention; normalization is omitted. Today we vary which positions can attend and which output embeddings the task uses.
00 · What should the Transformer output?
Different tasks for the same input
What should the Transformer output?
Read the diagram labels
- Raghav goes to school in Delhi.
- Raghav goes to school in Delhi.
Authored teaching diagram
What kinds of output could we request for this sentence?
A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
Teaching note
What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
00 · What should the Transformer output?
Different tasks for the same input
What should the Transformer output?
Read the diagram labels
- Raghav goes to school in Delhi.. What comes next?. A continuation
- Raghav goes to school in Delhi.
- What comes next?
- A continuation
Authored teaching diagram
What kinds of output could we request for this sentence?
A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
Teaching note
What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
00 · What should the Transformer output?
Different tasks for the same input
What should the Transformer output?
Read the diagram labels
- Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACE
- Raghav goes to school in Delhi.
- What comes next?
- A continuation
- What is each token?
- PERSON … PLACE
Authored teaching diagram
What kinds of output could we request for this sentence?
A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
Teaching note
What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
00 · What should the Transformer output?
Different tasks for the same input
What should the Transformer output?
Read the diagram labels
- Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACE. Say it in Hindi. राघव दिल्ली में स्कूल जाता है।
- Raghav goes to school in Delhi.
- What comes next?
- A continuation
- What is each token?
- PERSON … PLACE
- Say it in Hindi
- राघव दिल्ली में स्कूल जाता है।
Authored teaching diagram
What kinds of output could we request for this sentence?
A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
Teaching note
What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.
00 · What should the Transformer output?
Review: contextual token embeddings
The subscript tells us which token; the diagram shows its embedding after the blocks.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. Token embeddings + positions → Transformer blocks. . e1. . e2. . e3. . e4. . e5. . e6. ei = final contextualised embedding · d numbers. E = final source embeddings stacked as rows · 6 × d
- Raghav
- goes
- to
- school
- in
- Delhi
- Token embeddings + positions → Transformer blocks
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - e
i = final contextualised embedding · d numbers - E = final source embeddings stacked as rows · 6 × d
Authored teaching diagram · Primary source
Does a Transformer automatically return just one summary vector?
No. Each token retains an embedding as it passes through the blocks. We write its final contextualised embedding as eᵢ, containing d numbers. E stacks the final source embeddings as rows; here its shape is 6 × d.
Teaching note
Does a Transformer automatically return just one summary vector? Trace Raghav to e₁, then goes to e₂. Recall eᵢ → attention message mᵢ → update Δeᵢ → e′ᵢ = eᵢ + Δeᵢ → MLP and residual. We keep the token-position subscript as its embedding is updated; the diagram labels identify the stage. These final embeddings are not the attention messages.
00 · What should the Transformer output?
Attention, output representations and training targets
Read the diagram labels
- 01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?
- 01
- WHICH POSITIONS CAN ATTEND TO EACH OTHER?
- 02
- WHICH EMBEDDINGS DOES THE TASK USE?
- 03
- WHAT IS THE TRAINING TARGET?
Authored teaching diagram
What should we identify before naming a model family?
The allowed information flow, the states we read, and the targets used for training.
Teaching note
What should we identify before naming a model family? Return to these questions whenever the task changes.
01 · Generate the next token
Part I · Generate the next token
Read the diagram labels
- Review of the causal Transformer.. Raghav goes to. school …. Predict the next token from the available prefix.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Review of the causal Transformer.
- Raghav goes to
- school …
- Predict the next token from the available prefix.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part I · Generate the next token
Review of the causal Transformer.
Teaching note
Part I · Generate the next token Pause and establish the requested output before advancing.
01 · Generate the next token
Review: predicting the next token
m₃ updates token 3 (“to”); its final embedding predicts token 4 (“school”).
Read the diagram labels
- Raghav. goes. to. e1. e2. e3. Causal attention · query from position 3. m3 = Σj α3j vj. message received by “to”. Δe3 = m3 WO. e′3 = e3 + Δe3. keep the original e3. MLP + residual. e3 · final embedding. Vocabulary head → school (word 4). One head; final block shown.. Normalisation omitted.
- Raghav
- goes
- to
- e
1 - e
2 - e
3 - Causal attention · query from position 3
- m
3 = Σj α3j vj - message received by “to”
- Δe
3 = m3 WO - e′
3 = e3 + Δe3 - keep the original e
3 - MLP + residual
- e
3 · final embedding - Vocabulary head → school (word 4)
- One head; final block shown.
- Normalisation omitted.
Authored teaching diagram · Primary source
Which state predicts school after Raghav goes to?
The final contextualised embedding e₃, at the position of to, goes through the vocabulary head. m₃ is the attention message received at position 3; it is projected into Δe₃, added to e₃, and followed by the MLP and its residual. It is not a message for the word being predicted. A decoder-only stack has causal self-attention and feed-forward sublayers; it does not have a separate encoder or source cross-attention. The term decoder names its generation role. Scores cover the vocabulary even though the drawing shows one selected word.
Teaching note
Which state predicts school after Raghav goes to? Trace the message, projection, residual addition and MLP. The subscript identifies the receiving input position: m₃ belongs to to; m₄ would belong to school only after school is supplied. For several heads, concatenate their messages before W_O. The figure shows the final block and omits normalisation; earlier blocks repeat the same kind of update. A decoder-only stack has causal self-attention and feed-forward sublayers; it does not have a separate encoder or source cross-attention. The term decoder names its generation role. Scores cover the vocabulary even though the drawing shows one selected word.
01 · Generate the next token
Causal attention: current and earlier positions
Filled cell = connection allowed, not an attention weight.
Read the diagram labels
- 6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. × goes. × to. × school. × in. × Delhi
- 6 × 6 · allowed connections
- Raghav goes to school in Delhi.
- key columns
- R
- R
- g
- g
- t
- t
- s
- s
- i
- i
- D
- D
- Raghav can read…
- Raghav
- × goes
- × to
- × school
- × in
- × Delhi
Authored teaching diagram · Primary source
Can Raghav’s state use the later word Delhi?
Not under this causal mask. A query at position i can use positions up to and including i. The training target is the following token.
Teaching note
Can Raghav’s state use the later word Delhi? Read the first row, then the last row. Filled cells mean access is allowed, not that weights are equal.
01 · Generate the next token
Generating a sequence one token at a time
Each word has an updated embedding. Read the last one to predict the next word.
Read the diagram labels
- Raghav. e1 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. Read e1 → vocabulary head → goes
- Raghav
- e
1 · input - Causal Transformer blocks
- m
i → Δei → ei + Δei → MLP + residual - e
1 · Raghav - updated embedding
- Read e
1 → vocabulary head → goes
Authored teaching diagram · Primary source
What changes after we predict the next token?
The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
Teaching note
What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
01 · Generate the next token
The hidden layer inside a Transformer block
The MLP updates the embedding; the vocabulary layer reads it afterward.
Read the diagram labels
- Zoom into the MLP at Raghav’s position in the last block. After attention + residual. Hidden layer + activation. MLP output. 8 features. 12 hidden units. 8 features. +. . e1. updated. embedding. residual: all 8 features. Example widths: 8 → 12 → 8. The stack repeats attention and MLP updates.
- Zoom into the MLP at Raghav’s position in the last block
- After attention + residual
- Hidden layer + activation
- MLP output
- 8 features
- 12 hidden units
- 8 features
- +
- e
1 - updated
- embedding
- residual: all 8 features
- Example widths: 8 → 12 → 8. The stack repeats attention and MLP updates.
Authored teaching diagram · Primary source
Is the vocabulary head another stack of hidden layers?
In the standard decoder shown here, the hidden processing is already inside the Transformer blocks. Each block has an MLP with a hidden layer. This drawing expands the final block at Raghav’s position: 8 input features, 12 hidden units with a nonlinear activation, and 8 output features, followed by residual addition. Widths 8 and 12 are chosen for drawing, not a specific model. Normalization is omitted, and its exact placement depends on the architecture. The left vector already includes the attention update and its residual. The final e₁ denotes the updated embedding at position 1, not the embedding of the word to be generated. Multiple stacked blocks provide multiple hidden layers. Some model variants use more elaborate prediction heads, but they are not required here.
Teaching note
Is the vocabulary head another stack of hidden layers? Follow the dense connections through the MLP. Trace the residual bypass into the plus sign. The resulting updated e₁ is still eight numbers, ready for the next slide. Widths 8 and 12 are chosen for drawing, not a specific model. Normalization is omitted, and its exact placement depends on the architecture. The left vector already includes the attention update and its residual. The final e₁ denotes the updated embedding at position 1, not the embedding of the word to be generated. Multiple stacked blocks provide multiple hidden layers. Some model variants use more elaborate prediction heads, but they are not required here.
01 · Generate the next token
One output neuron for every vocabulary token
Every vocabulary token gets a score, including tokens unrelated to the input.
Read the diagram labels
- Example: d = 8 · vocabulary size |V| = 50,000. Updated e1. 50,000 output neurons. logit. Raghav. 1.0. bicycle. -2.0. goes. 5.5. ##ing. -1.0. school. 1.5. banana. -3.0. to. 3.0. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. 8 features. learned weights. e1 (1 × 8) × W (8 × 50,000) + b → 50,000 logits. Selected entries shown; the dots stand for thousands of other tokens.
- Example: d = 8 · vocabulary size |V| = 50,000
- Updated e
1 - 50,000 output neurons
- logit
- Raghav
- 1.0
- bicycle
- -2.0
- goes
- 5.5
- ##ing
- -1.0
- school
- 1.5
- banana
- -3.0
- to
- 3.0
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- 8 features
- learned weights
- e
1 (1 × 8) × W (8 × 50,000) + b → 50,000 logits - Selected entries shown; the dots stand for thousands of other tokens.
Authored teaching diagram · Primary source
Are the candidate outputs limited to words in the sentence?
No. The output layer scores the full tokenizer vocabulary. This example has 50,000 output neurons. We show seven entries, including bicycle, banana and a subword fragment, with dots for the omitted entries. A learned W maps the eight embedding features to 50,000 logits; b adds one bias per output. Vocabulary size 50,000 and embedding width eight are illustrative, not specifications of a named model. W has shape 8 × 50,000, b has shape 1 × 50,000 and the output has shape 1 × 50,000. The vocabulary is fixed by the tokenizer; ##ing is an illustrative subword notation. The neuron list is a selection, not vocabulary order or a ranking. All scores are authored examples. The seven visible logits are [1, -2, 5.5, -1, 1.5, -3, 3]; each omitted entry has logit -8 for the numerical softmax example. Some models omit the bias. W here is distinct from the attention projection W_O.
Teaching note
Are the candidate outputs limited to words in the sentence? Trace all eight input features to the output neurons. Point to bicycle and banana: these are valid vocabulary entries even though they are unlikely continuations here. The highlighted goes neuron has the largest illustrative logit. Vocabulary size 50,000 and embedding width eight are illustrative, not specifications of a named model. W has shape 8 × 50,000, b has shape 1 × 50,000 and the output has shape 1 × 50,000. The vocabulary is fixed by the tokenizer; ##ing is an illustrative subword notation. The neuron list is a selection, not vocabulary order or a ranking. All scores are authored examples. The seven visible logits are [1, -2, 5.5, -1, 1.5, -3, 3]; each omitted entry has logit -8 for the numerical softmax example. Some models omit the bias. W here is distinct from the attention projection W_O.
01 · Generate the next token
From vocabulary scores to the next token
Greedy decoding selects the highest probability across the full vocabulary.
Read the diagram labels
- Softmax over all 50,000 vocabulary scores. token. logit. probability. Raghav. 1.0. 0.94%. bicycle. -2.0. 0.05%. goes. 5.5. 84.58%. ##ing. -1.0. 0.13%. school. 1.5. 1.55%. banana. -3.0. 0.02%. to. 3.0. 6.94%. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. Highest: goes. Raghav goes. Append and repeat. Other 49,993 tokens combined:. 5.80%. pj = exp(scorej) / Σ exp(scores) · sum over all 50,000 tokens
- Softmax over all 50,000 vocabulary scores
- token
- logit
- probability
- Raghav
- 1.0
- 0.94%
- bicycle
- -2.0
- 0.05%
- goes
- 5.5
- 84.58%
- ##ing
- -1.0
- 0.13%
- school
- 1.5
- 1.55%
- banana
- -3.0
- 0.02%
- to
- 3.0
- 6.94%
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- ⋮
- Highest: goes
- Raghav goes
- Append and repeat
- Other 49,993 tokens combined:
- 5.80%
- p
j = exp(scorej ) / Σ exp(scores) · sum over all 50,000 tokens
Authored teaching diagram · Primary source
Does softmax consider only the displayed tokens?
No. Its denominator includes all 50,000 scores. The displayed probabilities use the full vocabulary, and the combined probability of the other 49,993 entries appears below the table. Greedy decoding selects goes, appends it to Raghav, and continues with the new prefix. The seven visible logits and 49,993 omitted logits of -8 define this authored distribution at temperature 1. Percentages are rounded to two decimal places; exact probabilities sum to one. The omitted rows are distinct token outputs, not one extra token. Greedy decoding chooses the maximum; sampling can choose another token. Softmax computes probabilities and has no learned parameters; selection is a separate operation. During training, cross-entropy against the known next token updates the vocabulary layer and Transformer. Later generation steps reuse the same projection.
Teaching note
Does softmax consider only the displayed tokens? Compare goes with the unrelated tokens. The omitted entries contribute to normalization even though they are not drawn. Trace the selected token into Raghav goes; its updated e₂ is read at the next step. The seven visible logits and 49,993 omitted logits of -8 define this authored distribution at temperature 1. Percentages are rounded to two decimal places; exact probabilities sum to one. The omitted rows are distinct token outputs, not one extra token. Greedy decoding chooses the maximum; sampling can choose another token. Softmax computes probabilities and has no learned parameters; selection is a separate operation. During training, cross-entropy against the known next token updates the vocabulary layer and Transformer. Later generation steps reuse the same projection.
01 · Generate the next token
Generating a sequence one token at a time
Each word has an updated embedding. Read the last one to predict the next word.
Read the diagram labels
- Raghav. e1 · input. goes. e2 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. Read e2 → vocabulary head → to
- Raghav
- e
1 · input - goes
- e
2 · input - Causal Transformer blocks
- m
i → Δei → ei + Δei → MLP + residual - e
1 · Raghav - updated embedding
- e
2 · goes - updated embedding
- Read e
2 → vocabulary head → to
Authored teaching diagram · Primary source
What changes after we predict the next token?
The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
Teaching note
What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
01 · Generate the next token
Generating a sequence one token at a time
Each word has an updated embedding. Read the last one to predict the next word.
Read the diagram labels
- Raghav. e1 · input. goes. e2 · input. to. e3 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. . e3 · to. updated embedding. Read e3 → vocabulary head → school
- Raghav
- e
1 · input - goes
- e
2 · input - to
- e
3 · input - Causal Transformer blocks
- m
i → Δei → ei + Δei → MLP + residual - e
1 · Raghav - updated embedding
- e
2 · goes - updated embedding
- e
3 · to - updated embedding
- Read e
3 → vocabulary head → school
Authored teaching diagram · Primary source
What changes after we predict the next token?
The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
Teaching note
What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.
01 · Generate the next token
Decoder-only models: next-token prediction
At each step, read the last updated embedding to predict the next token.
Read the diagram labels
- Raghav. goes. to. input e1. input e2. input e3. Causal self-attention + residual. MLP + residual. Each token reads. itself + its past. past → next. . e1. . e2. . e3. Vocabulary head → school. × L. blocks. Read the last embedding. Append school. Repeat.
- Raghav
- goes
- to
- input e
1 - input e
2 - input e
3 - Causal self-attention + residual
- MLP + residual
- Each token reads
- itself + its past
- past → next
- e
1 - e
2 - e
3 - Vocabulary head → school
- × L
- blocks
- Read the last embedding. Append school. Repeat.
Authored teaching diagram · Primary source
What makes this a decoder-only model, and what is self-attention attending to?
A single causal Transformer stack processes the prompt and generated tokens as one sequence. In each self-attention layer, Q, K and V are projected from the current embeddings of that sequence. The causal mask lets each position read itself and earlier positions. The last updated embedding feeds the vocabulary head. This is the standard text-only decoder-only architecture: no separate source encoder and no encoder-to-decoder cross-attention. Each block has causal multi-head self-attention and an MLP, with residual connections; normalization is omitted here. L counts repeated blocks. Training normally predicts the next token at every eligible position using a shifted target, while generation reads the last position. The same vocabulary head is shared across positions. The matrix shows allowed access, not measured attention. Decoding includes a probability distribution and token-selection step, expanded on the preceding slides.
Teaching note
What makes this a decoder-only model, and what is self-attention attending to? Follow the three supplied tokens into the repeated blocks and their three updated embeddings. Only the last embedding feeds the displayed next-token head. Read the causal square in token order: Raghav, goes, to. Append school and repeat. Colored cells mark allowed attention, not its magnitude; input embedding lookup and normalization are omitted in this summary. This is the standard text-only decoder-only architecture: no separate source encoder and no encoder-to-decoder cross-attention. Each block has causal multi-head self-attention and an MLP, with residual connections; normalization is omitted here. L counts repeated blocks. Training normally predicts the next token at every eligible position using a shifted target, while generation reads the last position. The same vocabulary head is shared across positions. The matrix shows allowed access, not measured attention. Decoding includes a probability distribution and token-selection step, expanded on the preceding slides.
01 · Generate the next token
Decoder-only examples: text and code completion
One sequence, one causal Transformer, one next-token prediction at each step.
Read the diagram labels
- Supplied prefix. Continuation, token by token. Text. Raghav goes to. school …. Code. def square(x): return. x * x. At each step: the prefix and tokens generated so far form one sequence.. . Current sequence. embeddings + positions. . Causal Transformer. self-attention + MLP blocks. . Vocabulary head. select the next token. Append the selected token to the same sequence.. Q, K and V come from this sequence; there is no separate source encoder.
- Supplied prefix
- Continuation, token by token
- Text
- Raghav goes to
- school …
- Code
- def square(x): return
- x * x
- At each step: the prefix and tokens generated so far form one sequence.
- Current sequence
- embeddings + positions
- Causal Transformer
- self-attention + MLP blocks
- Vocabulary head
- select the next token
- Append the selected token to the same sequence.
- Q, K and V come from this sequence; there is no separate source encoder.
Authored teaching diagram · Primary source
Does a prompt and an answer imply an encoder-decoder architecture?
No. In a decoder-only model, the prompt and the generated answer occupy successive positions in one sequence. Causal self-attention processes that sequence. A separate source encoder is not present. Text and code completion illustrate this directly. The two rows are separate illustrative completions, not measured model outputs or two inputs processed together. In every self-attention layer, Q, K and V come from the same current sequence. Read the final embedding to produce vocabulary logits, apply softmax and select a token. Decoder-only models can also translate, summarize and chat when suitably trained: the instruction and any source text remain inside the prompt. Those task names do not determine the architecture. A standard Transformer encoder-decoder instead has a separate source encoder whose outputs are available to the decoder through cross-attention. Discuss that distinction after the translation example. Capability depends on training and data; the architecture alone does not guarantee useful completions.
Teaching note
Does a prompt and an answer imply an encoder-decoder architecture? Read each joined strip as one sequence. Purple tokens are supplied; amber tokens are appended one at a time. Only tokens generated so far are inputs at the next step. Follow the shared computation below and the return arrow after token selection. The two rows are separate illustrative completions, not measured model outputs or two inputs processed together. In every self-attention layer, Q, K and V come from the same current sequence. Read the final embedding to produce vocabulary logits, apply softmax and select a token. Decoder-only models can also translate, summarize and chat when suitably trained: the instruction and any source text remain inside the prompt. Those task names do not determine the architecture. A standard Transformer encoder-decoder instead has a separate source encoder whose outputs are available to the decoder through cross-attention. Discuss that distinction after the translation example. Capability depends on training and data; the architecture alone does not guarantee useful completions.
02 · Understand the supplied input
Part II · The whole input is available
Read the diagram labels
- Predict labels for a complete input sentence.. Complete sentence. Token or sentence labels. Use contextual embeddings to classify the input.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Predict labels for a complete input sentence.
- Complete sentence
- Token or sentence labels
- Use contextual embeddings to classify the input.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part II · The whole input is available
Predict labels for a complete input sentence.
Teaching note
Part II · The whole input is available Pause and establish the requested output before advancing.
02 · Understand the supplied input
Which words refer to people or places?
Assign entity labels to words in the complete input.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. PERSON. PLACE
- Raghav
- goes
- to
- school
- in
- Delhi
- PERSON
- PLACE
Authored teaching diagram
Do we need to generate a new sentence?
No. This task assigns a label to each input position. Raghav is a person and Delhi is a place in this teaching example.
Teaching note
Do we need to generate a new sentence? Point to Raghav and Delhi in the same source sentence used on the opening slide.
02 · Understand the supplied input
Predicting a label at each input position
Six input positions → six label decisions.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Do we need to generate a new sentence?
- Raghav
- goes
- to
- school
- in
- Delhi
- Transformer · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - PERSON
- OTHER
- OTHER
- OTHER
- OTHER
- PLACE
- Do we need to generate a new sentence?
Authored teaching diagram · Primary source
What should the unhighlighted words receive?
OTHER, under our simplified label menu. Every position receives a label, including words that are not named entities. The four-class inventory is PERSON, ORGANIZATION, PLACE and OTHER. This sentence does not have an organization, but that class remains available at every position.
Teaching note
What should the unhighlighted words receive? Count all six outputs, not just the two entity names. The four-class inventory is PERSON, ORGANIZATION, PLACE and OTHER. This sentence does not have an organization, but that class remains available at every position.
02 · Understand the supplied input
Self-attention over the complete input
Filled cell = connection allowed, not an attention weight.
Read the diagram labels
- 6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. goes. to. school. in. Delhi
- 6 × 6 · allowed connections
- Raghav goes to school in Delhi.
- key columns
- R
- R
- g
- g
- t
- t
- s
- s
- i
- i
- D
- D
- Raghav can read…
- Raghav
- goes
- to
- school
- in
- Delhi
Authored teaching diagram · Primary source
Why is it legitimate for Raghav’s state to use Delhi?
We supplied the whole input and ask for labels about it. Later source words are available information, not future target answers.
Teaching note
Why is it legitimate for Raghav’s state to use Delhi? Compare the full matrix with the familiar causal triangle. Do not introduce a new attention equation.
02 · Understand the supplied input
Queries, keys and values in an encoder
The attention calculation is unchanged; each position can now use the complete input.
Read the diagram labels
- Raghav’s embedding e1. q1 = e1 × WQ. Current embeddings stacked in E. K = E × WK. V = E × WV. m1 = Attention(q1, K, V) = softmax(q1KT / √dk) V. The projections are shared across positions within a head.
- Raghav’s embedding e
1 - q
1 = e1 × WQ - Current embeddings stacked in E
- K = E × W
K - V = E × W
V - m
1 = Attention(q1 , K, V) = softmax(q1 KT / √dk ) V - The projections are shared across positions within a head.
Authored teaching diagram · Primary source
Do we need a new attention operation to read the whole input?
No. The same query–key scores, softmax and weighted values apply. The mask now permits all supplied positions. We use row-vector notation: q₁ = e₁ × W_Q, K = E × W_K, V = E × W_V. e₁ is Raghav’s current embedding; E stacks all current input embeddings for this layer. The retrieved result m₁ is its attention message. Q, K and V in the projection subscripts identify the learned maps. Scaling uses the query/key width dₖ, not necessarily the model width.
Teaching note
Do we need a new attention operation to read the whole input? Trace the Raghav query and the source-wide keys and values. Within one head/layer, positions share the projection matrices. We use row-vector notation: q₁ = e₁ × W_Q, K = E × W_K, V = E × W_V. e₁ is Raghav’s current embedding; E stacks all current input embeddings for this layer. The retrieved result m₁ is its attention message. Q, K and V in the projection subscripts identify the learned maps. Scaling uses the query/key width dₖ, not necessarily the model width.
02 · Understand the supplied input
One contextual vector per token
E has shape n × d.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Raghav in context. Delhi in context. E ∈ ℝ6 × d. 6 token positions × d features
- Raghav
- goes
- to
- school
- in
- Delhi
- Transformer · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Raghav in context
- Delhi in context
- E ∈ ℝ
6 × d - 6 token positions × d features
Authored teaching diagram · Primary source
Is e₁ just the original embedding for Raghav?
No. It is the final contextual representation at that position after the Transformer blocks, and can depend on the supplied sentence.
Teaching note
Is e₁ just the original embedding for Raghav? Distinguish the input token identity from its final vector. The drawn bars indicate multiple features, not semantic coordinates.
02 · Understand the supplied input
A shared classifier for all token positions
One W and b, reused at all six positions.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. ONE shared classifier: zi = ei W + b → softmax over 4 labels. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Entity labels at every position: PERSON · ORGANIZATION · PLACE · OTHER
- Raghav
- goes
- to
- school
- in
- Delhi
- Transformer · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - ONE shared classifier: z
i = ei W + b → softmax over 4 labels - PERSON
- OTHER
- OTHER
- OTHER
- OTHER
- PLACE
- Entity labels at every position: PERSON · ORGANIZATION · PLACE · OTHER
Authored teaching diagram · Primary source
Are we training six unrelated classifiers?
No. A single shared classifier maps each d-dimensional state to the same four label scores. Softmax operates across the four classes at each position.
Teaching note
Are we training six unrelated classifiers? Trace both Raghav and Delhi through the same wide classifier band. Ask which parameters are shared.
02 · Understand the supplied input
Four entity labels for each token
Training target: one entity label per position.
Read the diagram labels
- . ei. same W, same b. PERSON. ORGANIZATION. PLACE. OTHER. d features → 4 scores → 4 probabilities
- e
i - same W, same b
- PERSON
- ORGANIZATION
- PLACE
- OTHER
- d features → 4 scores → 4 probabilities
Authored teaching diagram · Primary source
Does the classifier for Delhi still score PERSON?
Yes. Every position receives all four class scores. A gold label per token supervises the corresponding probability distribution. A typical supervised loss sums or averages token cross-entropies. The task head must be trained; merely attaching a new head does not teach labels.
Teaching note
Does the classifier for Delhi still score PERSON? Point at all four scores, then say that the same computation repeats for every input position. A typical supervised loss sums or averages token cross-entropies. The task head must be trained; merely attaching a new head does not teach labels.
02 · Understand the supplied input
Encoder: a contextual representation for each token
Use the embeddings for token labels, sentence readouts or similarity.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. input e1. input e2. input e3. input e4. input e5. input e6. Whole-input self-attention + residual. MLP + residual. Every token can read. the complete input. all ↔ all. × L. blocks. . e1. . e2. . e3. . e4. . e5. . e6. Six input tokens → six contextual embeddings.
- Raghav
- goes
- to
- school
- in
- Delhi
- input e
1 - input e
2 - input e
3 - input e
4 - input e
5 - input e
6 - Whole-input self-attention + residual
- MLP + residual
- Every token can read
- the complete input
- all ↔ all
- × L
- blocks
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Six input tokens → six contextual embeddings.
Authored teaching diagram · Primary source
Must an encoder reduce the input to a single vector?
No. Its output is a sequence of contextual states. Pooling, a readout token or a token classifier is a separate downstream choice. BERT is an example of this encoder family. Its pretraining objectives are not needed to understand the arrangement here.
Teaching note
Must an encoder reduce the input to a single vector? Trace each of the six input positions to its updated output embedding. The full attention square means every supplied position can read every other position. The six separate output arrows retain one embedding per input position; this does not pool the sentence into a single vector. BERT is an example of this encoder family. Its pretraining objectives are not needed to understand the arrangement here.
02 · Understand the supplied input
Predicting one label for a sentence
The movie was really awful. → NEGATIVE
Read the diagram labels
- The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. Which state should we read?. One review → one sentiment label
- The
- movie
- was
- really
- awful
- Encoder
- e
1 - e
2 - e
3 - e
4 - e
5 - Which state should we read?
- One review → one sentiment label
Authored teaching diagram
We have five contextual states. How do we get one class prediction?
Choose a whole-sequence readout, such as pooling the states or selecting a learned readout position, then apply a classifier.
Teaching note
We have five contextual states. How do we get one class prediction? Pause with all five states visible before suggesting a pooling operation.
02 · Understand the supplied input
Sentence classification with mean pooling
The pooled vector is the input to the sentence classifier.
Read the diagram labels
- The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. mean: c = (e1 + … + e5) / 5. c → classifier → NEGATIVE
- The
- movie
- was
- really
- awful
- Encoder
- e
1 - e
2 - e
3 - e
4 - e
5 - mean: c = (e
1 + … + e5 ) / 5 - c → classifier → NEGATIVE
Authored teaching diagram · Primary source
What shape does mean pooling produce?
Averaging n vectors of width d produces one vector of width d. A classifier maps that vector to C class scores. Mean pooling is a possible design, not a guarantee of useful features. It must be used with suitable training. This example has no padding.
Teaching note
What shape does mean pooling produce? Follow the five state arrows into the mean and then into the classifier. Remember this pooling drawing for translation. Mean pooling is a possible design, not a guarantee of useful features. It must be used with suitable training. This example has no padding.
02 · Understand the supplied input
Sentence classification with a CLS token
[CLS] participates in ordinary whole-input attention.
Read the diagram labels
- [CLS]. The. movie. was. really. awful. Encoder · every supplied position is available. . eCLS. . e1. . e2. . e3. . e4. . e5. Classifier → NEGATIVE
- [CLS]
- The
- movie
- was
- really
- awful
- Encoder · every supplied position is available
- e
CLS - e
1 - e
2 - e
3 - e
4 - e
5 - Classifier → NEGATIVE
Authored teaching diagram · Primary source
How can a token at the beginning collect information from later words?
The whole-input mask allows its query to read every supplied position, including the later words. We select its final contextual state for the task head. CLS also has keys and values and uses the ordinary per-head/layer projection matrices. The encoder now returns six states for five review words plus CLS; the classifier reads one.
Teaching note
How can a token at the beginning collect information from later words? Trace the added input position through the same stack. Follow only its final state to the sentiment head. CLS also has keys and values and uses the ordinary per-head/layer projection matrices. The encoder now returns six states for five review words plus CLS; the classifier reads one.
02 · Understand the supplied input
Where does the CLS embedding come from?
[CLS] has a token ID and a trainable row in the embedding table.
Read the diagram labels
- Embedding table · trainable. . token ID. d numbers. . [CLS]. . Raghav. . goes. [CLS] has its own row, like a word.. How does that row start?. From scratch: small random values, once.. Pretrained: load the saved row.. Look up eCLS, then add position. Encoder + sentence → updated eCLS. Same table row for every sentence; the encoder output depends on the sentence.
- Embedding table · trainable
- token ID
- d numbers
- [CLS]
- Raghav
- goes
- [CLS] has its own row, like a word.
- How does that row start?
- From scratch: small random values, once.
- Pretrained: load the saved row.
- Look up e
CLS , then add position - Encoder + sentence → updated e
CLS - Same table row for every sentence; the encoder output depends on the sentence.
Authored teaching diagram · Primary source
Do we initialize CLS with a new random vector for every sentence?
No. In this text encoder, [CLS] is a special token with its own ID. That ID selects a trainable row of d numbers from the same embedding table used for ordinary tokens. From scratch, initialise that row with small random values once; a pretrained model supplies its saved row. Add positional information, then send the sequence through the encoder. The table is a model parameter, not a fresh random draw for each example. Backpropagation and optimizer steps learn its CLS row. This is the standard special-token lookup design for text; a separate trainable vector is another implementation choice. The coloured cells are schematic vector coordinates, not measured values.
Teaching note
Do we initialize CLS with a new random vector for every sentence? Point to the ordinary token rows, then the CLS row. Follow the lookup arrow into the encoder input. The initial row is shared across sentences at any fixed set of model parameters; the updated CLS embedding depends on the supplied sentence. The table is a model parameter, not a fresh random draw for each example. Backpropagation and optimizer steps learn its CLS row. This is the standard special-token lookup design for text; a separate trainable vector is another implementation choice. The coloured cells are schematic vector coordinates, not measured values.
02 · Understand the supplied input
CLS learns through the prediction loss
Compare prediction ŷ with label y; use the loss to train the CLS representation.
Read the diagram labels
- Forward pass: the CLS embedding table row and review-word embeddings enter the encoder. Select the updated CLS embedding, pass it into a schematic neural classifier and softmax, and obtain predicted probabilities y hat. The supplied positive label y and probabilities define cross-entropy loss minus log 0.8, approximately 0.223. Dashed backward arrows show gradients through softmax, classifier, encoder and the input table lookup. The optimizer updates trainable parameters. The displayed probabilities are an illustrative example, not model measurements.
- FORWARD PASS → compute the prediction and loss
- input
- [CLS] row
- e
CLS - Encoder
- a good movie
- review words
- updated
- e
CLS - Neural classifier
- 2 class scores
- softmax
- prediction
- ŷ
- [0.2, 0.8]
- NEG, POS
- L(ŷ, y)
- given label
- y = POSITIVE
- −log(0.8)
- ≈ 0.223
- ← BACKWARD PASS: loss gradients flow through the same operations
- Optimizer updates: CLS table row + encoder + classifier parameters.
Authored teaching diagram · Primary source
Do we provide a target embedding for CLS, or a true label y?
We supply the true class label y. The CLS table row and review words pass through the encoder; the updated CLS embedding enters a neural classifier. Softmax gives the predicted class probabilities ŷ. The loss compares ŷ with y. Backpropagation computes gradients through the classifier and encoder to the original CLS embedding-table row. Here ŷ denotes the full predicted probability vector, while y is the supplied class label, not another prediction. The loss uses the probability assigned to the true class; no argmax is inserted in the differentiable training path. The classifier drawing is a schematic MLP with one hidden layer; a linear classifier is also possible. Neuron counts do not specify the encoder width. Input embeddings receive positional information; all encoder positions participate even though only the CLS output is drawn. The dashed lane is an offset view of gradients through the same computational graph, not a direct shortcut from loss to CLS. This shows full fine-tuning of unfrozen parameters. The optimizer changes model parameters, not the stored value of a per-example updated CLS activation. Frozen parameters would not update. Natural logarithms are used, and the probabilities are illustrative.
Teaching note
Do we provide a target embedding for CLS, or a true label y? Follow the solid forward arrows left to right. For the positive review, the illustrative probabilities are NEG 0.2 and POS 0.8; cross-entropy is −log(0.8), about 0.223. Then follow the dashed lower lane from right to left. Finally distinguish gradient computation from the optimizer step. Here ŷ denotes the full predicted probability vector, while y is the supplied class label, not another prediction. The loss uses the probability assigned to the true class; no argmax is inserted in the differentiable training path. The classifier drawing is a schematic MLP with one hidden layer; a linear classifier is also possible. Neuron counts do not specify the encoder width. Input embeddings receive positional information; all encoder positions participate even though only the CLS output is drawn. The dashed lane is an offset view of gradients through the same computational graph, not a direct shortcut from loss to CLS. This shows full fine-tuning of unfrozen parameters. The optimizer changes model parameters, not the stored value of a per-example updated CLS activation. Frozen parameters would not update. Natural logarithms are used, and the probabilities are illustrative.
02 · Understand the supplied input
Token classification and sentence classification
Read every state for NER; read CLS for one sentence label.
Read the diagram labels
- A LABEL PER TOKEN. Raghav goes to school in Delhi. Encoder. . e1. PERSON. . e6. PLACE. …. shared head. ONE SENTENCE LABEL. [CLS] Raghav goes to school in Delhi. Encoder. . eCLS. one classifier → one label
- A LABEL PER TOKEN
- Raghav goes to school in Delhi
- Encoder
- e
1 - PERSON
- e
6 - PLACE
- …
- shared head
- ONE SENTENCE LABEL
- [CLS] Raghav goes to school in Delhi
- Encoder
- e
CLS - one classifier → one label
Authored teaching diagram · Primary source
What changes when the whole sentence needs one label?
The task determines the readout and head. NER reads all ordinary token states through shared parameters. Sentence classification selects the final CLS state. For example, one could define a topic classification dataset; this drawing does not invent a measured topic prediction.
Teaching note
What changes when the whole sentence needs one label? Compare the identical Raghav sentence on both sides. The sentence-classification label is deliberately unspecified until a dataset defines the task. For example, one could define a topic classification dataset; this drawing does not invent a measured topic prediction.
02 · Understand the supplied input
Six words, eight features per word
6 counts token positions; 8 counts features inside each embedding.
Read the diagram labels
- Raghav goes to school in Delhi. Encoder · choose d = 8 for this example. 1. 2. 3. 4. 5. 6. 7. 8. Raghav. goes. to. school. in. Delhi. 6 rows = 6 token positions. 8 columns = 8 features. E has shape 6 × 8. Each row is one updated embedding. The cells stand for numbers.
- Raghav goes to school in Delhi
- Encoder · choose d = 8 for this example
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- Raghav
- goes
- to
- school
- in
- Delhi
- 6 rows = 6 token positions
- 8 columns = 8 features
- E has shape 6 × 8
- Each row is one updated embedding. The cells stand for numbers.
Authored teaching diagram · Primary source
What does a single row of this matrix represent?
One word’s updated embedding. We choose a toy model width of eight, so six input words produce six rows of eight numbers: E has shape 6 × 8. A feature column is not a word or a class label. This is one sentence with no batch dimension. The width eight is chosen only to make the shape easy to draw. It does not come from the six-word sentence or the four entity classes.
Teaching note
What does a single row of this matrix represent? Read the Raghav row across all eight cells, then count all six rows. The cells represent vector coordinates, not actual values. This is one sentence with no batch dimension. The width eight is chosen only to make the shape easy to draw. It does not come from the six-word sentence or the four entity classes.
02 · Understand the supplied input
NER: follow each word from input to scores
6 × 8 input features → 6 × 8 updated features → 6 × 4 class scores.
Read the diagram labels
- Six words are aligned with six input embeddings. Each embedding has eight features in this toy example. Whole-input self-attention and MLP blocks with residual updates return six updated eight-feature embeddings. One shared classifier multiplies every updated embedding by the same 8 by 4 weight matrix and adds the same four-component bias. It returns four scores per word, in the order PERSON, ORGANIZATION, PLACE, OTHER. The twenty-four output boxes are score slots, not measured values. Input embeddings include position information.
- Raghav
- input e
1 - goes
- input e
2 - to
- input e
3 - school
- input e
4 - in
- input e
5 - Delhi
- input e
6 - Whole-input self-attention → MLP
- residual updates · repeated encoder blocks
- updated e
1 - updated e
2 - updated e
3 - updated e
4 - updated e
5 - updated e
6 - Same classifier at every position: × W (8 × 4), then + b (4)
- Each group of four boxes holds one word’s PERSON, ORG, PLACE and OTHER scores.
Authored teaching diagram · Primary source
What happens before we multiply by the classifier weights?
Each word has an input embedding, with positional information. Whole-input self-attention and MLP blocks update those embeddings using the sentence. The encoder still returns one eight-feature vector per word. The shared classifier then multiplies each updated vector by W and adds b, producing four class scores. We retain eᵢ at position i and label its input and updated stages explicitly. The model width eight is illustrative. W has shape 8 × 4; b has four components and is added, not multiplied. In matrix form, scores = E W + b, with b broadcast across the six rows. The four output slots per word use PERSON, ORGANIZATION, PLACE, OTHER order; they are not numerical predictions. The next slide rearranges those six sets of scores into rows and applies softmax independently within each row. Positional information and normalization are omitted from the drawing.
Teaching note
What happens before we multiply by the classifier weights? Follow Raghav straight down, then Delhi. At every stage there are six token positions. The encoder changes their features while preserving width eight. The task head changes the feature width from eight to four class scores; its parameters are shared across all six positions. We retain eᵢ at position i and label its input and updated stages explicitly. The model width eight is illustrative. W has shape 8 × 4; b has four components and is added, not multiplied. In matrix form, scores = E W + b, with b broadcast across the six rows. The four output slots per word use PERSON, ORGANIZATION, PLACE, OTHER order; they are not numerical predictions. The next slide rearranges those six sets of scores into rows and applies softmax independently within each row. Positional information and normalization are omitted from the drawing.
02 · Understand the supplied input
NER: turn each set of scores into a label
The classifier changes 8 features into 4 scores at every position.
Read the diagram labels
- One shared classifier reads each of the six embeddings. E: 6 × 8. W: 8 × 4; b: 4. Scores: 6 × 4. PERSON. ORG. PLACE. OTHER. Raghav. PERSON. goes. OTHER. to. OTHER. school. OTHER. in. OTHER. Delhi. PLACE. chosen label. Softmax within each row → 4 probabilities → one label per word.
- One shared classifier reads each of the six embeddings
- E: 6 × 8
- W: 8 × 4; b: 4
- Scores: 6 × 4
- PERSON
- ORG
- PLACE
- OTHER
- Raghav
- PERSON
- goes
- OTHER
- to
- OTHER
- school
- OTHER
- in
- OTHER
- Delhi
- PLACE
- chosen label
- Softmax within each row → 4 probabilities → one label per word.
Authored teaching diagram · Primary source
Why is the result 6 × 4 rather than one vector of four scores?
We classify every input position. The same W of shape 8 × 4 and bias of length four are used for all six rows: scores = E W + b has shape 6 × 4. Softmax across the four columns gives one distribution per word. Dark cells indicate the illustrative highest-scoring class; their shading is not a numerical score or probability. Raghav is labelled PERSON, Delhi PLACE, and the other words OTHER in this teaching example. The table depicts score slots and chosen labels, not a trained model run.
Teaching note
Why is the result 6 × 4 rather than one vector of four scores? The previous slide showed four score slots beneath every word. Here each word is a row and each class is a column. Point across the Raghav row, then the Delhi row. Both use the same classifier parameters. ORG abbreviates ORGANIZATION. Dark cells indicate the illustrative highest-scoring class; their shading is not a numerical score or probability. Raghav is labelled PERSON, Delhi PLACE, and the other words OTHER in this teaching example. The table depicts score slots and chosen labels, not a trained model run.
02 · Understand the supplied input
CLS: select one row from seven
7 × 8 becomes 1 × 8 when we select the updated CLS embedding.
Read the diagram labels
- Add [CLS] to the same six-word sentence. 7 input positions → encoder → 7 updated embeddings. 1. 2. 3. 4. 5. 6. 7. 8. [CLS]. Raghav. goes. to. school. in. Delhi. 7 rows × 8 features = 7 × 8. Select the updated eCLS row. 1 row × 8 features = 1 × 8. Selection keeps all eight features.
- Add [CLS] to the same six-word sentence
- 7 input positions → encoder → 7 updated embeddings
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- [CLS]
- Raghav
- goes
- to
- school
- in
- Delhi
- 7 rows × 8 features = 7 × 8
- Select the updated e
CLS row - 1 row × 8 features = 1 × 8
- Selection keeps all eight features.
Authored teaching diagram · Primary source
Does selecting CLS shrink every embedding to one feature?
No. It selects one position and keeps all eight features in that row. Six ordinary words plus CLS produce seven updated embeddings: 7 × 8. Reading just the CLS row gives 1 × 8. CLS can contain information from the other words because its embedding has already passed through whole-input attention. This is row selection, not averaging the rows or compressing the feature dimension.
Teaching note
Does selecting CLS shrink every embedding to one feature? Hold the eight columns fixed while highlighting the CLS row. The other rows still exist; this task head simply does not read them directly. CLS can contain information from the other words because its embedding has already passed through whole-input attention. This is row selection, not averaging the rows or compressing the feature dimension.
02 · Understand the supplied input
A neural classifier reads the updated CLS embedding
The one-hot vector identifies CLS; its updated embedding is used for classification.
Read the diagram labels
- CLS token identity can be represented by a vocabulary-length one-hot vector. Embedding lookup selects its learned dense input embedding; the encoder combines it with the sentence and returns an updated dense CLS embedding. That eight-feature vector enters a six-unit hidden layer with a nonlinear activation, then three output neurons. Illustrative logits 2, 1, 0 become probabilities 0.665, 0.245, 0.090 for EDUCATION, TRAVEL, OTHER after softmax. The one-hot identity is not the classifier input. Vocabulary size, embedding width, hidden width and number of classes are distinct.
- [CLS] token identity
- 0
- 0
- 1
- 0
- …
- 0
- one-hot · |V| entries
- Embedding lookup
- input e
CLS - Encoder + sentence
- e
CLS - updated
- Updated e
CLS - Hidden layer
- 3 logits
- Probabilities
- softmax
- 2.0
- 0.665
- EDUCATION
- 1.0
- 0.245
- TRAVEL
- 0.0
- 0.090
- OTHER
- 8 dense features → 6 hidden units → 3 logits → 3 probabilities
- Example scores; the hidden layer has a nonlinear activation.
Authored teaching diagram · Primary source
Does the classifier receive a one-hot CLS vector?
No. A vocabulary-length one-hot vector represents the identity of [CLS] for embedding lookup. That lookup returns its dense input embedding. After the encoder processes CLS with the sentence, the updated CLS vector enters the classifier. Our example has eight input features, six hidden units with a nonlinear activation, and three output logits. The one-hot vector has |V| entries, one for each vocabulary token; it is not a three-class target vector. Implementations normally index the embedding table directly rather than materialize one-hot vectors. Positions are included before the encoder. The schematic classifier is an MLP: h = activation(e_CLS W₁ + b₁), with W₁ of shape 8 × 6 and b₁ of length 6; logits = h W₂ + b₂, with W₂ of shape 6 × 3 and b₂ of length 3. This is a valid alternative to a linear 8 × 3 head. The logit example [2, 1, 0] gives softmax probabilities approximately [0.665, 0.245, 0.090], summing to one before rounding. These are authored numbers, not model measurements. A labelled dataset and an annotation rule define the topic classes. The model width, hidden width, vocabulary size and class count are independent choices.
Teaching note
Does the classifier receive a one-hot CLS vector? Trace the top strip from token identity through lookup and encoding, then follow the arrow to the eight input neurons. Count six hidden units and three output neurons. Read each logit and its probability beside EDUCATION, TRAVEL and OTHER. The softmax operates across all three logits. The one-hot vector has |V| entries, one for each vocabulary token; it is not a three-class target vector. Implementations normally index the embedding table directly rather than materialize one-hot vectors. Positions are included before the encoder. The schematic classifier is an MLP: h = activation(e_CLS W₁ + b₁), with W₁ of shape 8 × 6 and b₁ of length 6; logits = h W₂ + b₂, with W₂ of shape 6 × 3 and b₂ of length 3. This is a valid alternative to a linear 8 × 3 head. The logit example [2, 1, 0] gives softmax probabilities approximately [0.665, 0.245, 0.090], summing to one before rounding. These are authored numbers, not model measurements. A labelled dataset and an annotation rule define the topic classes. The model width, hidden width, vocabulary size and class count are independent choices.
02 · Understand the supplied input
One encoder, several ways to use its embeddings
Selecting updated CLS and pooling token embeddings are two readout choices.
Read the diagram labels
- Encoder. Select updated eCLS. one label. Encoder. Use updated e1 … en. one label / token. Encoder. . Pool token embeddings. or select updated eCLS. representation. Readout = choosing or combining the encoder’s updated embeddings.
- Encoder
- Select updated e
CLS - one label
- Encoder
- Use updated e
1 … en - one label / token
- Encoder
- Pool token embeddings
- or select updated e
CLS - representation
- Readout = choosing or combining the encoder’s updated embeddings.
Authored teaching diagram · Primary source
Does readout mean another token or another operation after CLS?
Readout is the general choice of which encoder outputs to use and how to combine them. Selecting the updated e_CLS is one readout. Pooling token embeddings is another. Token classification uses all ordinary token embeddings. There is no extra readout token or mandatory processing step implied by the word here. The diagram abbreviates the task heads in the first two rows. A sentence or token classifier is still needed to produce labels. Retrieval can use a suitably trained representation, potentially with a learned projection and normalization; neither raw pooling nor raw CLS is automatically useful for similarity. The training objective and evaluation must match the use.
Teaching note
Does readout mean another token or another operation after CLS? Read each concrete choice aloud. The top row selects updated CLS for a sentence classifier. The middle row sends all token embeddings through a shared token classifier. The bottom row forms a whole-input representation by pooling or by selecting updated CLS. The diagram abbreviates the task heads in the first two rows. A sentence or token classifier is still needed to produce labels. Retrieval can use a suitably trained representation, potentially with a learned projection and normalization; neither raw pooling nor raw CLS is automatically useful for similarity. The training objective and evaluation must match the use.
03 · Represent one sequence, generate another
Part III · Sequence-to-sequence prediction
Read the diagram labels
- Encode the English input and generate a Hindi translation.. English source. Hindi target. Represent the source and condition target generation.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Encode the English input and generate a Hindi translation.
- English source
- Hindi target
- Represent the source and condition target generation.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part III · Sequence-to-sequence prediction
Encode the English input and generate a Hindi translation.
Teaching note
Part III · Sequence-to-sequence prediction Pause and establish the requested output before advancing.
03 · Represent one sequence, generate another
Example: English-to-Hindi translation
राघव दिल्ली में स्कूल जाता है।
Read the diagram labels
- Raghav goes to school in Delhi.. राघव दिल्ली में स्कूल जाता है।. Encode the source sentence and generate the target sequence.
- Raghav goes to school in Delhi.
- राघव दिल्ली में स्कूल जाता है।
- Encode the source sentence and generate the target sequence.
Authored teaching diagram · Primary source
Do contextual source vectors alone produce the Hindi sentence?
They represent the input. To generate a variable-length target sentence, we also need a mechanism that predicts target tokens and a stopping token. The translation is authored. Displayed Hindi words are teaching units, not a claim about real model tokenization.
Teaching note
Do contextual source vectors alone produce the Hindi sentence? Notice Delhi moves earlier in the Hindi sentence. Avoid suggesting one fixed source position maps to each output position. The translation is authored. Displayed Hindi words are teaching units, not a claim about real model tokenization.
03 · Represent one sequence, generate another
How does Hindi generation begin?
[शुरू] starts the process; [समाप्त] tells us to stop.
Read the diagram labels
- The Hindi labels [शुरू] and [समाप्त] stand for special beginning-of-sequence and end-of-sequence token IDs. They are not ordinary Hindi words in the translated sentence. Supply the start ID, look up its trainable target embedding, add its position and begin decoder computation. Generate six displayed Hindi word tokens and stop on the end ID. These are teaching labels, not a claim about the exact spellings or special token conventions of a particular tokenizer.
- [शुरू] means “begin generating” · a special token ID
- [शुरू]
- Target embedding table: special row
- e
1 - Look up e
1 , add p1 , then run the decoder with the English source. - We supply [शुरू]. The first predicted Hindi word is राघव.
- [शुरू] राघव दिल्ली में स्कूल जाता है [समाप्त]
- The markers control generation; the translation contains only the six Hindi words.
Authored teaching diagram
Is शुरू the first word of the translation?
No. [शुरू] is our Hindi display label for a special beginning-of-sequence token, often called START or BOS. We supply its token ID and look up its embedding e₁ in the target embedding table. Its decoder output predicts the first Hindi word राघव. [समाप्त] labels the end token; it is not printed as part of the translated sentence. The marker labels are a teaching convention, not a tokenizer configuration change. A real model uses its configured decoder start ID, which can be a BOS, EOS, pad or other designated ID. Here we use a simple BOS/EOS teaching design. The start row is a learned parameter loaded from a checkpoint or initialized before training. Each displayed Hindi word is treated as one token for clarity; actual tokenizers may split it. The target embedding is updated through the decoder while its table row is fixed during inference.
Teaching note
Is शुरू the first word of the translation? Point to the brackets: these mark special tokens. Trace the start ID to its embedding-table row, then read the complete example output. We do not supply the complete target sentence during generation. The marker labels are a teaching convention, not a tokenizer configuration change. A real model uses its configured decoder start ID, which can be a BOS, EOS, pad or other designated ID. Here we use a simple BOS/EOS teaching design. The start row is a learned parameter loaded from a checkpoint or initialized before training. Each displayed Hindi word is treated as one token for clarity; actual tokenizers may split it. The target embedding is updated through the decoder while its table row is fixed during inference.
03 · Represent one sequence, generate another
Step 1: encode the source sentence
The same encoder states we used for NER.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Keep one contextual vector per source position.
Authored teaching diagram · Primary source
What new operation have we introduced so far?
None. English tokens receive embeddings and positions, then the encoder returns one contextual state per source position.
Teaching note
What new operation have we introduced so far? Keep the source strip, encoder band and vector positions fixed for all the following builds.
03 · Represent one sequence, generate another
First attempt: compress the source into one vector
c = Pool(e₁, …, e₆)
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. One fixed source vector, c ∈ ℝd
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Pool → c
- One fixed source vector, c ∈ ℝ
d
Authored teaching diagram · Primary source
Where have we already used this operation?
In whole-sentence classification. Here the pooled vector will condition generation instead of feeding a class head. This first design gives the decoder one pooled source vector. We will choose how the decoder accesses the source in the following section.
Teaching note
Where have we already used this operation? Reveal the pooling arrows below the unchanged encoder states. This first design gives the decoder one pooled source vector. We will choose how the decoder accesses the source in the following section.
03 · Represent one sequence, generate another
Conditioning the decoder on the source vector
p(yₜ | y<ₜ, c)
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi token
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Pool → c
- Causal decoder
- target prefix
- [शुरू] राघव …
- next Hindi token
Authored teaching diagram · Primary source
What information does the decoder need at this step?
It needs the target tokens so far and the source context c. For now, imagine the decoder has access to c at every step; the next diagrams show three ways to provide it. The diagram states a conditional dependency. It does not specify an extra token or a particular source interface. We first show the dependency, then compare concrete implementations.
Teaching note
What information does the decoder need at this step? Add the familiar causal decoder beneath c. Show the target prefix entering from the left. The diagram states a conditional dependency. It does not specify an extra token or a particular source interface. We first show the dependency, then compare concrete implementations.
03 · Represent one sequence, generate another
A fixed source vector at each generation step
The decoder has access to c at every step.
Read the diagram labels
- source. target prefix. next token. same c. +. [शुरू]. राघव. same c. +. [शुरू] राघव. दिल्ली. same c. +. [शुरू] राघव दिल्ली. में. p(yt | y<t, c)
- source
- target prefix
- next token
- same c
- +
- [शुरू]
- राघव
- same c
- +
- [शुरू] राघव
- दिल्ली
- same c
- +
- [शुरू] राघव दिल्ली
- में
- p(y
t | y<t , c)
Authored teaching diagram · Primary source
What changes between these three predictions?
The source context c stays fixed. The target prefix grows, so the decoder state and next-token prediction can change. This frame describes p(yₜ | y<ₜ, c). The following frames implement that dependency in three different ways; the plus sign on this frame means both inputs are available, not necessarily numerical addition. Continue through स्कूल, जाता and है, then predict [समाप्त]. The outputs are teaching examples.
Teaching note
What changes between these three predictions? Read the three rows in order. Keep pointing back to the same c. This frame describes p(yₜ | y<ₜ, c). The following frames implement that dependency in three different ways; the plus sign on this frame means both inputs are available, not necessarily numerical addition. Continue through स्कूल, जाता and है, then predict [समाप्त]. The outputs are teaching examples.
03 · Represent one sequence, generate another
Option 1: add context to every target embedding
Project c to the decoder width, then broadcast the same vector across target positions.
Read the diagram labels
- English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Project: u = c WP. WP: dc × d. Add to every position: et + u. Add pt → target self-attention → MLP. Vocabulary head → softmax → choose में
- English → Encoder → e
1 … e6 → Pool - c
- d
c numbers - [शुरू] राघव दिल्ली
- e
1 , e2 , e3 · each has d numbers - Project: u = c W
P - W
P : dc × d - Add to every position: e
t + u - Add p
t → target self-attention → MLP - Vocabulary head → softmax → choose में
Authored teaching diagram
Can we add c directly to a target embedding?
Only if their dimensions match. Here c has d_c features and each target embedding has d features. Learn W_P with shape d_c × d, compute u = c W_P, then add u to each eₜ before positional information and the causal decoder. This is a concrete teaching architecture, not a required definition of an encoder-decoder. The decoder input has the same number of positions as the target prefix. The source encoder, projection and decoder are trained together on next-target-token loss. The source vector is recomputed for each source sentence and held fixed while that sentence is decoded. Normalisation and biases are omitted for clarity.
Teaching note
Can we add c directly to a target embedding? Trace the same source encoder and pool into c, then follow the projection and the repeated addition at the target positions. This is a concrete teaching architecture, not a required definition of an encoder-decoder. The decoder input has the same number of positions as the target prefix. The source encoder, projection and decoder are trained together on next-target-token loss. The source vector is recomputed for each source sentence and held fixed while that sentence is decoded. Normalisation and biases are omitted for clarity.
03 · Represent one sequence, generate another
Add a position vector at each target position
eᵢ identifies the token; u carries source context; pᵢ marks the position.
Read the diagram labels
- [शुरू]. e1 + u. p1. +. e1 + u + p1. राघव. e2 + u. p2. +. e2 + u + p2. दिल्ली. e3 + u. p3. +. e3 + u + p3. pi is a position vector with d numbers, just like ei and u.. Same source context u; a different pi for each position.. Add matching coordinates. The result still has d features.
- [शुरू]
- e
1 + u - p
1 - +
- e
1 + u + p1 - राघव
- e
2 + u - p
2 - +
- e
2 + u + p2 - दिल्ली
- e
3 + u - p
3 - +
- e
3 + u + p3 - p
i is a position vector with d numbers, just like ei and u. - Same source context u; a different p
i for each position. - Add matching coordinates. The result still has d features.
Authored teaching diagram
What does “add positions” actually add?
In this example, pᵢ is a d-dimensional position vector. We add it coordinate by coordinate to the context-conditioned embedding eᵢ + u. शुरू receives p₁, राघव receives p₂ and दिल्ली receives p₃. The decoder therefore receives e₁ + u + p₁, e₂ + u + p₂ and e₃ + u + p₃. This is an additive positional-encoding design. Position vectors may be learned or fixed sinusoidal vectors; their construction is not needed here. Other Transformer designs incorporate positions differently, for example through rotations of queries and keys. We index illustrated positions from one for teaching. Add the positional information once, after the displayed source-context fusion. For concatenation, add pᵢ after W_F has returned the fused vector to width d. For prepending, the source prefix occupies position 1, so the three target positions receive p₂, p₃ and p₄.
Teaching note
What does “add positions” actually add? Follow each vertical column. The source context u is shared; the token embedding and position vector differ. Point at each plus node and count d features before and after addition. This is an additive positional-encoding design. Position vectors may be learned or fixed sinusoidal vectors; their construction is not needed here. Other Transformer designs incorporate positions differently, for example through rotations of queries and keys. We index illustrated positions from one for teaching. Add the positional information once, after the displayed source-context fusion. For concatenation, add pᵢ after W_F has returned the fused vector to width d. For prepending, the source prefix occupies position 1, so the three target positions receive p₂, p₃ and p₄.
03 · Represent one sequence, generate another
Causal self-attention over the Hindi prefix
Target queries read target keys and values. The causal mask keeps future tokens hidden.
Read the diagram labels
- Causal target self-attention. Query rows शुरू, Raghav and Delhi can read target key/value columns up to their own position. The future token mein is blocked and not supplied. Queries, keys and values are projections of the target-side current embeddings, which already include the added source context and positional information.
- The same attention operation, now among Hindi target positions.
- Keys / values: positions we can read
- [शुरू]
- राघव
- दिल्ली
- में
- [शुरू]
- ✓
- ×
- ×
- ×
- राघव
- ✓
- ✓
- ×
- ×
- दिल्ली
- ✓
- ✓
- ✓
- ×
- Rows = queries · filled cells = allowed
- Q, K, V all come from target e
i - [शुरू]
- reads itself
- राघव
- reads शुरू and itself
- दिल्ली
- reads all three supplied positions
- में is the next prediction. Its embedding is not an input yet.
Authored teaching diagram
Is the decoder using self-attention again?
Yes. शुरू, राघव and दिल्ली exchange information through causal self-attention. Each target position supplies a query, key and value from its current embedding. शुरू can read itself; राघव can also read शुरू; दिल्ली can read all three. The next token में is not supplied yet. These embeddings already contain u and positional information. Their queries, keys and values are all target-side projections: this is self-attention. There is no separate query-to-English-token lookup in this fixed-context design. During parallel training, future supplied target positions are masked; during this generation step, they are absent. The same rule applies after concatenation and projection. With a prepended source vector, that extra earlier position also provides a key and value.
Teaching note
Is the decoder using self-attention again? Read each matrix row as one receiving target position. The gray में column is a teaching placeholder for a future token, not an actual input at this generation step. These embeddings already contain u and positional information. Their queries, keys and values are all target-side projections: this is self-attention. There is no separate query-to-English-token lookup in this fixed-context design. During parallel training, future supplied target positions are masked; during this generation step, they are absent. The same rule applies after concatenation and projection. With a prepended source vector, that extra earlier position also provides a key and value.
03 · Represent one sequence, generate another
Updating the embedding at दिल्ली
The message m₃ updates दिल्ली; its final embedding will predict में.
Read the diagram labels
- One self-attention head at target position 3. Use the current Delhi embedding for q3 and all three supplied target embeddings for keys and values. Scaled query-key scores with a causal mask become softmax weights; the weighted values form message m3. Project with W_O, add the residual, then use the MLP and its residual. With multiple heads, concatenate their messages before W_O. The final updated Delhi embedding predicts the next token mein.
- Here e
i means the current target embedding, after adding u and pi . - दिल्ली: current e
3 - [शुरू] राघव दिल्ली: current e
1 , e2 , e3 - q
3 = e3 WQ - k
i = ei WK ; vi = ei WV - Match q
3 with k1 , k2 , k3 - Mask future positions
- Softmax → α
1 , α2 , α3 - Message m
3 = α1 v1 + α2 v2 + α3 v3 - values
- Δe
3 = m3 WO → add to e3 → MLP + residual → updated e3
Authored teaching diagram
Where does the attention message go?
Project the current Delhi embedding into q₃. Compare it with the keys of the supplied target positions, scale the scores by the square root of the key width, apply the causal mask and softmax, then combine their values into m₃. Project the message with W_O to get Δe₃ and add it to the current embedding. The MLP and its residual follow. The displayed equations describe one head. Multiple heads compute their own messages; concatenate them before W_O. Layer normalization is omitted from the drawing. Here eᵢ denotes the current embedding entering the attention sublayer: in the first block this is the prepared eᵢ + u + pᵢ from the previous slide. Scores use q₃ kᵢ transpose divided by sqrt(d_k). Repeat the block updates before reading the last position through the vocabulary head. The third embedding belongs to दिल्ली, not to the token being predicted.
Teaching note
Where does the attention message go? Point from q₃ to the three keys, then from the weighted values to m₃. Trace the same message, delta and residual convention from the earlier attention lecture. The displayed equations describe one head. Multiple heads compute their own messages; concatenate them before W_O. Layer normalization is omitted from the drawing. Here eᵢ denotes the current embedding entering the attention sublayer: in the first block this is the prepared eᵢ + u + pᵢ from the previous slide. Scores use q₃ kᵢ transpose divided by sqrt(d_k). Repeat the block updates before reading the last position through the vocabulary head. The third embedding belongs to दिल्ली, not to the token being predicted.
03 · Represent one sequence, generate another
From target embeddings to the next Hindi token
The same decoder and vocabulary head are reused at the next generation step.
Read the diagram labels
- [शुरू]. e1 + u + p1. राघव. e2 + u + p2. दिल्ली. e3 + u + p3. . Causal self-attention + residual → MLP + residual. Repeat the blocks. Each position can read itself and earlier positions.. updated e1. updated e2. updated e3. Vocabulary head → |V| logits. Softmax. Choose: में. Read the last target embedding; append में and repeat.
- [शुरू]
- e
1 + u + p1 - राघव
- e
2 + u + p2 - दिल्ली
- e
3 + u + p3 - Causal self-attention + residual → MLP + residual
- Repeat the blocks. Each position can read itself and earlier positions.
- updated e
1 - updated e
2 - updated e
3 - Vocabulary head → |V| logits
- Softmax
- Choose: में
- Read the last target embedding; append में and repeat.
Authored teaching diagram
What is the “last embedding,” and why read that one?
The decoder updates every supplied position through causal self-attention and MLP blocks. Here the last target token is दिल्ली, so we read its updated e₃. The vocabulary head turns it into |V| scores; softmax gives a vocabulary distribution, from which we choose the example next token में. The subscript of e₃ identifies the third target token, not the token being predicted. The earlier updated embeddings still exist; generation uses the last target position to choose the next token. |V| is the target vocabulary size. Softmax does not itself choose the token. In the addition design, every new target input also receives the same source vector u for this source sentence. Concatenation and prepending change input preparation; the causal blocks and last-target-position readout follow the same pattern. Normalization and the individual attention-head projections are omitted from this summary.
Teaching note
What is the “last embedding,” and why read that one? Trace all three prepared inputs to all three updated outputs. Highlight the last one, then follow it through logits, softmax and token selection. Append में, obtain its token embedding and position vector, and run the next step. The subscript of e₃ identifies the third target token, not the token being predicted. The earlier updated embeddings still exist; generation uses the last target position to choose the next token. |V| is the target vocabulary size. Softmax does not itself choose the token. In the addition design, every new target input also receives the same source vector u for this source sentence. Concatenation and prepending change input preparation; the causal blocks and last-target-position readout follow the same pattern. Normalization and the individual attention-head projections are omitted from this summary.
03 · Represent one sequence, generate another
Option 2: concatenate context as extra features
Concatenation increases the feature width; it does not add a token position.
Read the diagram labels
- English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Concatenate features: [et ∥ c]. same c beside each et · d + dc features. Project: [et ∥ c] WF → d features. WF: (d + dc) × d. Add pt → causal blocks → updated e3 → logits → softmax → में
- English → Encoder → e
1 … e6 → Pool - c
- d
c numbers - [शुरू] राघव दिल्ली
- e
1 , e2 , e3 · each has d numbers - Concatenate features: [e
t ∥ c] - same c beside each e
t · d + dc features - Project: [e
t ∥ c] WF → d features - W
F : (d + dc ) × d - Add p
t → causal blocks → updated e3 → logits → softmax → में
Authored teaching diagram
Does concatenating c create another token?
Not here: [eₜ ∥ c] concatenates features within each target position. Its width is d + d_c. A learned W_F of shape (d + d_c) × d maps it back to the decoder width. Then add the d-dimensional position vector pₜ and use the causal decoder, last-target-position readout, vocabulary logits and softmax just shown. Apply the same fusion map at every target position. Train it with the source encoder and decoder through the translation loss. With a linear fusion, [eₜ ∥ c]W_F can be split into a learned transform of eₜ plus a learned transform of c; these are related conditioning designs, not claims of strictly different expressive power. This diagram is an authored teaching construction.
Teaching note
Does concatenating c create another token? Point to the two feature groups, then the projection back to d. Count the target positions: there are still three. Apply the same fusion map at every target position. Train it with the source encoder and decoder through the translation loss. With a linear fusion, [eₜ ∥ c]W_F can be split into a learned transform of eₜ plus a learned transform of c; these are related conditioning designs, not claims of strictly different expressive power. This diagram is an authored teaching construction.
03 · Represent one sequence, generate another
Option 3: prepend context as an extra input position
The target reads the source-prefix vector through ordinary causal self-attention.
Read the diagram labels
- English → Encoder → e1 … e6 → Pool. c. dc numbers. u = c WP · d numbers. Target embeddings: [शुरू] राघव दिल्ली. . u · source. + p1. . e1 · शुरू. + p2. . e2 · राघव. + p3. . e3 · दिल्ली. + p4. Causal decoder → updated target embeddings → select last e3. Vocabulary head → softmax → choose में. p1 for u; p2, p3, p4 for targets
- English → Encoder → e
1 … e6 → Pool - c
- d
c numbers - u = c W
P · d numbers - Target embeddings: [शुरू] राघव दिल्ली
- u · source
- + p
1 - e
1 · शुरू - + p
2 - e
2 · राघव - + p
3 - e
3 · दिल्ली - + p
4 - Causal decoder → updated target embeddings → select last e
3 - Vocabulary head → softmax → choose में
- p
1 for u; p2 , p3 , p4 for targets
Authored teaching diagram
Why place u before शुरू rather than after the target prefix?
Project c to a d-dimensional vector u and put it before the target embeddings. Every target query can then attend to that earlier position under the causal mask. If u were placed after the current prediction position, the causal mask would prevent that position from reading it. The continuous prefix is a vector, not a vocabulary word to predict. Assign positions consistently, including the extra slot, and omit a next-token loss on the source slot in this teaching construction. Train the projection, source encoder and decoder on target next-token loss. Unlike the later cross-attention architecture, this uses one pooled source vector and causal self-attention. The appendix enlarges the same prefix construction.
Teaching note
Why place u before शुरू rather than after the target prefix? Count four input positions. Add p₁ to u, p₂ to the शुरू embedding, p₃ to राघव and p₄ to दिल्ली. Keep e₁, e₂ and e₃ as target-token indices. Read updated e₃, at the fourth total input position, through the vocabulary head and softmax to predict में. The continuous prefix is a vector, not a vocabulary word to predict. Assign positions consistently, including the extra slot, and omit a next-token loss on the source slot in this teaching construction. Train the projection, source encoder and decoder on target next-token loss. Unlike the later cross-attention architecture, this uses one pooled source vector and causal self-attention. The appendix enlarges the same prefix construction.
03 · Represent one sequence, generate another
An encoder-decoder generates a target from a source
Represent the source + generate the target.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi token. ENCODER-DECODER · source context + target prefix
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Pool → c
- Causal decoder
- target prefix
- [शुरू] राघव …
- next Hindi token
- ENCODER-DECODER · source context + target prefix
Authored teaching diagram · Primary source
What are the two sequences in this architecture?
The complete supplied English source and the growing Hindi target. Their positions, lengths and vocabularies need not be identical.
Teaching note
What are the two sequences in this architecture? Name the family now, after following the computation through both stacks.
03 · Represent one sequence, generate another
The decoder cannot revisit the source
The decoder relies on c for all source information.
Read the diagram labels
- Eight source embeddings are pooled into one fixed-width vector c. The decoder receives c and a growing target prefix, with no direct path back to the source embeddings. Different target steps need who, where and when details. If pooling fails to preserve Delhi, the source detail cannot be retrieved later. The Hindi phrases are illustrative fragments, not a full token-by-token translation.
- Raghav goes to school in Delhi every morning.
- Encoder
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - e
7 - e
8 - 8 updated embeddings
- pool
- c
- fixed width
- Decoder + output head
- Growing target prefix
- Different target steps need different source details
- Who?
- Raghav → राघव
- Where?
- Delhi → दिल्ली
- When?
- every morning → हर सुबह
- If c loses “Delhi”, the decoder cannot look back to recover it.
Authored teaching diagram · Primary source
So what is the problem with using one vector?
A fixed vector can work. The restriction is that every source detail needed later must survive pooling into c. If c does not preserve Delhi, a later decoder step cannot query the original Delhi embedding to recover it. Each displayed word is one token for this teaching example, so the source has eight positions. The Hindi items are translated fragments illustrating different information needs, not consecutive next-token predictions. Pooling produces a fixed-width vector even when the source grows longer. This creates a representation and learning burden; it does not prove that a fixed vector must forget information. The decoder can use c differently as its target prefix changes. It may guess a missing detail from learned patterns, but cannot retrieve it from the source. The next architecture keeps all source embeddings available for a fresh lookup.
Teaching note
So what is the problem with using one vector? Trace encoder outputs through pooling into c. Show the separate growing target prefix. Then ask what source information is needed for who, where and when. There is no direct edge from the source embeddings to the decoder. Each displayed word is one token for this teaching example, so the source has eight positions. The Hindi items are translated fragments illustrating different information needs, not consecutive next-token predictions. Pooling produces a fixed-width vector even when the source grows longer. This creates a representation and learning burden; it does not prove that a fixed vector must forget information. The decoder can use c differently as its target prefix changes. It may guess a missing detail from learned patterns, but cannot retrieve it from the source. The next architecture keeps all source embeddings available for a fresh lookup.
04 · Cross-attention to the source
Part IV · Cross-attention to the source
Read the diagram labels
- Retain one encoder output per source position.. Updated target embedding. queries. Source keys and values. Let each target position query the source directly.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Retain one encoder output per source position.
- Updated target embedding
- queries
- Source keys and values
- Let each target position query the source directly.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part IV · Cross-attention to the source
Retain one encoder output per source position.
Teaching note
Part IV · Cross-attention to the source Pause and establish the requested output before advancing.
04 · Cross-attention to the source
Retaining the source embeddings for the decoder
Keep e₁ … e₆ available to the decoder.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Keep one contextual vector per source position.
Authored teaching diagram · Primary source
What could we retain instead of just c?
The full sequence of encoder output states. The decoder can make a source lookup suited to each current target position.
Teaching note
What could we retain instead of just c? Return to the identical source drawing. Remove the pooling constraint while leaving the encoder unchanged.
04 · Cross-attention to the source
Source attention when predicting राघव
All six English embeddings are available; the target query can assign more weight to Raghav.
Read the diagram labels
- To predict the first Hindi token Raghav from शुरू, retain all six English encoder outputs. Illustrative hand-chosen weights 0.72, 0.04, 0.03, 0.09, 0.04, 0.08 put most weight on the source name Raghav. These are explanatory numbers, not measured model attention or guaranteed alignments.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- An illustrative lookup: put most weight on the English name Raghav.
- Target prefix: [शुरू] → example next Hindi token: राघव
- Illustrative attention weights; their sum is 1.
Authored teaching diagram
When predicting राघव, which source position might receive more attention?
For an intuitive example, give the English Raghav position weight 0.72 and distribute the remaining 0.28 over the other five positions. This illustrates a plausible source lookup for the first Hindi token. These numbers are hand-chosen, not measured attention. The weights sum to one. Real attention heads can use broader or different patterns; a name need not receive the largest weight in every head or layer. Crucially, the query for predicting the first Hindi token comes from the शुरू position. Feeding राघव to predict itself would leak the target. The following slides explain the source of the query, keys, values and weights.
Teaching note
When predicting राघव, which source position might receive more attention? Keep the encoder drawing fixed from the previous slide. Read all six weights, then point to शुरू as the supplied target prefix and राघव as the upcoming prediction. The weights sum to one. Real attention heads can use broader or different patterns; a name need not receive the largest weight in every head or layer. Crucially, the query for predicting the first Hindi token comes from the शुरू position. Feeding राघव to predict itself would leak the target. The following slides explain the source of the query, keys, values and weights.
04 · Cross-attention to the source
Step 1: encode the English sentence
Retain all six encoder outputs before starting the target decoder.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Six encoder outputs, one for each English position.. Keep these contextual embeddings available to the decoder.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - Six encoder outputs, one for each English position.
- Keep these contextual embeddings available to the decoder.
Authored teaching diagram
What information is available from the source?
The English encoder produces six contextual embeddings, one per supplied position. Keep them separately so the decoder can attend to different source positions. This sequence expands one head in one decoder block. English source indices and Hindi target indices are separate. Source e₁ belongs to Raghav; the target e₁ introduced next belongs to [शुरू].
Teaching note
What information is available from the source? Stop at the blue encoder outputs. Do not introduce target queries or source keys and values yet. This sequence expands one head in one decoder block. English source indices and Hindi target indices are separate. Source e₁ belongs to Raghav; the target e₁ introduced next belongs to [शुरू].
04 · Cross-attention to the source
Step 2: initialise the target at [शुरू]
Look up the start-token embedding and add its position vector.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. Initial target embedding. राघव has not been generated; the supplied target is only [शुरू].
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - TARGET: begin with the special token [शुरू]
- [शुरू]
- input e
1 + p1 - lookup + position
- Initial target embedding
- राघव has not been generated; the supplied target is only [शुरू].
Authored teaching diagram
Where does the first target embedding come from?
The special start token has an entry in the target embedding table. Look up its input e₁ and add p₁. No Hindi word has been generated yet. [शुरू] is the teaching label for the special beginning-of-sequence token. It is not an ordinary word in the translation. During inference, lookup reads the learned table row; it does not train that row.
Teaching note
Where does the first target embedding come from? Keep the encoder outputs fixed. Reveal only the start token, embedding lookup and position addition. [शुरू] is the teaching label for the special beginning-of-sequence token. It is not an ordinary word in the translation. During inference, lookup reads the learned table row; it does not train that row.
04 · Cross-attention to the source
Step 3: update [शुरू] with target self-attention
Causal self-attention updates the target embedding before cross-attention.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. Target Q, K, V → m1 → Δe1 → e1 + Δe1. Only [शुरू] is available: self-attention reads this one target position.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - TARGET: begin with the special token [शुरू]
- [शुरू]
- input e
1 + p1 - lookup + position
- Causal self-attention
- + residual
- updated e
1 - Target Q, K, V → m
1 → Δe1 → e1 + Δe1 - Only [शुरू] is available: self-attention reads this one target position.
Authored teaching diagram
With only [शुरू] supplied, what can target self-attention read?
Only the start position itself. Its current embedding supplies the self-attention query, key and value. The attention message is projected and added through the residual connection to update e₁. With a single unmasked key, the attention weight is one, but the value and output projections can still produce a nonzero update. Once more Hindi tokens are supplied, causal self-attention can use their preceding target context. Layer normalization is omitted. The target MLP follows cross-attention in this decoder block.
Teaching note
With only [शुरू] supplied, what can target self-attention read? Trace input e₁ plus position through self-attention and the residual to updated e₁. There is no token prediction at this stage. With a single unmasked key, the attention weight is one, but the value and output projections can still produce a nonzero update. Once more Hindi tokens are supplied, causal self-attention can use their preceding target context. Layer normalization is omitted. The target MLP follows cross-attention in this decoder block.
04 · Cross-attention to the source
Step 4: form a new query for cross-attention
Self-attention and cross-attention have separate learned query projections.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Cross-attention uses its own query projection WQ.. Its input is e1 after the target self-attention update.
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - TARGET: begin with the special token [शुरू]
- [शुरू]
- input e
1 + p1 - lookup + position
- Causal self-attention
- + residual
- updated e
1 - q
1 = e1 WQ - Cross-attention uses its own query projection W
Q . - Its input is e
1 after the target self-attention update.
Authored teaching diagram
Is this the query used in the preceding target self-attention?
No. Cross-attention applies its own W_Q to the updated target embedding. The self-attention query was formed within the preceding sublayer from its input embedding. The two sublayers have distinct parameters and read different sequences. W_Q here means the cross-attention query projection; the self-attention sublayer has a different W_Q. The local symbol is reused for the same mathematical role, not to imply shared weights. In implementations with pre-normalization, the projection takes a normalized version of this updated target stream. Normalization is omitted from the drawing.
Teaching note
Is this the query used in the preceding target self-attention? Reveal the new arrow from updated e₁ to q₁. Keep the e₁ subscript because this is still the start position. The cross-attention query will be matched against English source keys. W_Q here means the cross-attention query projection; the self-attention sublayer has a different W_Q. The local symbol is reused for the same mathematical role, not to imply shared weights. In implementations with pre-normalization, the projection takes a normalized version of this updated target stream. Normalization is omitted from the drawing.
04 · Cross-attention to the source
Step 5: match the target query to English keys
Cross-attention: Q from the updated target; K and V from encoder outputs.
Read the diagram labels
- Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. k1, v1. k2, v2. k3, v3. k4, v4. k5, v5. k6, v6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Source: ki = ei WK; vi = ei WV. Compare q1 with all six source keys: q1 kiT / √dk
- Raghav
- goes
- to
- school
- in
- Delhi
- ENCODER · whole-input self-attention
- e
1 - e
2 - e
3 - e
4 - e
5 - e
6 - k
1 , v1 - k
2 , v2 - k
3 , v3 - k
4 , v4 - k
5 , v5 - k
6 , v6 - TARGET: begin with the special token [शुरू]
- [शुरू]
- input e
1 + p1 - lookup + position
- Causal self-attention
- + residual
- updated e
1 - q
1 = e1 WQ - Source: k
i = ei WK ; vi = ei WV - Compare q
1 with all six source keys: q1 ki T / √dk
Authored teaching diagram
How are the English keys and values constructed?
Apply the cross-attention W_K and W_V to each encoder output to obtain six source keys and six values. Compare q₁ with each key using a scaled dot product. The next slide normalizes the scores and combines the values. Cross-attention has its own learned W_Q, W_K and W_V, separate from target and source self-attention parameters. Source keys and values can be computed in advance; this reveal order is for teaching and does not impose an execution order. Keys determine source-position weights; values supply the message. The query at [शुरू] supports prediction of राघव after the remaining decoder computation.
Teaching note
How are the English keys and values constructed? Reveal the six key/value pairs and the query-to-key connections. Follow the English paths down into K and V, then the updated target path into Q. Cross-attention has its own learned W_Q, W_K and W_V, separate from target and source self-attention parameters. Source keys and values can be computed in advance; this reveal order is for teaching and does not impose an execution order. Keys determine source-position weights; values supply the message. The query at [शुरू] supports prediction of राघव after the remaining decoder computation.
04 · Cross-attention to the source
Computing the weighted source message
Keys determine the weights; values supply the information sent to the target.
Read the diagram labels
- Illustrative scaled dot-product scores are the natural logarithms of weights 0.72, 0.04, 0.03, 0.09, 0.04 and 0.08. Scores are displayed rounded to two decimals. Softmax produces those weights, which multiply source values; sum them to obtain cross-attention message c1. Project the message, add it to the target stream, apply MLP and residual processing, and use the final vocabulary head to predict the illustrative Hindi token Raghav. The attention weights are source-position weights, not Hindi vocabulary probabilities.
- Match the target q
1 with each English key: q1 ki T / √dk - Raghav
- key k
1 - -0.33
- goes
- key k
2 - -3.22
- to
- key k
3 - -3.51
- school
- key k
4 - -2.41
- in
- key k
5 - -3.22
- Delhi
- key k
6 - -2.53
- Softmax across the six source scores → weights that sum to 1
- 0.72 × v
1 - 0.04 × v
2 - 0.03 × v
3 - 0.09 × v
4 - 0.04 × v
5 - 0.08 × v
6 - Add weighted source values → message c
1 - c
1 WO + target e1 → MLP + residual → vocabulary head → राघव
Authored teaching diagram
How do the illustrative weights become information the decoder can use?
Score q₁ against all six English keys, scale by sqrt(d_k), and apply softmax across source positions. Multiply each source value by its weight and add the six weighted vectors to get c₁. Project that message and add it to the target embedding; after the remaining decoder computation, the vocabulary head can predict राघव. To make the arithmetic internally consistent, the authored scaled scores are log(0.72), log(0.04), log(0.03), log(0.09), log(0.04), log(0.08); displayed scores are rounded to two decimals. Softmax of the unrounded scores reproduces the weights exactly. These are not trained model scores. c₁ is a cross-attention message, analogous to mᵢ in self-attention. For multiple heads, concatenate their messages before W_O. The decoder retains the target stream through the residual; c₁ alone is neither the final embedding nor a vocabulary distribution.
Teaching note
How do the illustrative weights become information the decoder can use? Trace source scores downward through softmax to the weighted values. The larger 0.72 weights the Raghav value more heavily in this illustration. Keep source-attention weights distinct from output-vocabulary probabilities. To make the arithmetic internally consistent, the authored scaled scores are log(0.72), log(0.04), log(0.03), log(0.09), log(0.04), log(0.08); displayed scores are rounded to two decimals. Softmax of the unrounded scores reproduces the weights exactly. These are not trained model scores. c₁ is a cross-attention message, analogous to mᵢ in self-attention. For multiple heads, concatenate their messages before W_O. The decoder retains the target stream through the residual; c₁ alone is neither the final embedding nor a vocabulary distribution.
04 · Cross-attention to the source
The source message updates शुरू’s embedding
Keep the current target embedding, then add the update retrieved from the source.
Read the diagram labels
- Track शुरू through embedding lookup, position addition and target causal self-attention. Its current target embedding e1 branches: W_Q makes a query, while the residual keeps e1. English encoder outputs supply keys and values. Cross-attention returns illustrative message c1, emphasizing the source Raghav value. W_O projects c1 into delta e1; the plus node adds that delta to the retained current target embedding. This updates the शुरू position, not the English source embeddings. The MLP and next-token prediction follow on the next frame. One head is shown; normalization is omitted.
- [शुरू] → lookup input e
1 → add p1 → causal self-attention + residual - TARGET · शुरू position
- current target e
1 - q
1 = e1 WQ - SOURCE · English encoder outputs
- e
1 … e6 → K and V - One key and value per English position
- Cross-attention(q
1 , K, V) - Message c
1 = 0.72 v1 + … + 0.08 v6 - Embedding update: Δe
1 = c1 WO - Keep the current e
1 - +
- updated e
1 = e1 + Δe1 - The source message changes शुरू’s embedding.
Authored teaching diagram
What exactly changes when शुरू attends to the English sentence?
शुरू begins with an embedding-table lookup and positional information. Target self-attention and its residual produce the current e₁. That embedding makes q₁ while a residual path keeps e₁. Cross-attention retrieves the source message c₁; W_O maps it to Δe₁. Add Δe₁ to the retained e₁, giving the शुरू position an embedding informed by the English source. This is one head in one decoder block; with multiple heads, concatenate their source messages before W_O. Layer normalization is omitted. The residual adds to the current target representation entering cross-attention, which has already passed through target self-attention; it does not jump back to the raw embedding-table row. The message c₁ plays the role of the earlier attention message mᵢ. The shown 0.72 weight is illustrative, not measured attention. A forward pass updates this contextual representation; it does not by itself train or overwrite the embedding table. Here the source is read through K/V, not through an added fixed pooled u.
Teaching note
What exactly changes when शुरू attends to the English sentence? Follow the शुरू stream, then pause at the fork: one copy makes the query and the other follows the long residual arrow. Trace the source message through W_O into the plus node. The result is still the embedding at शुरू. This is one head in one decoder block; with multiple heads, concatenate their source messages before W_O. Layer normalization is omitted. The residual adds to the current target representation entering cross-attention, which has already passed through target self-attention; it does not jump back to the raw embedding-table row. The message c₁ plays the role of the earlier attention message mᵢ. The shown 0.72 weight is illustrative, not measured attention. A forward pass updates this contextual representation; it does not by itself train or overwrite the embedding table. Here the source is read through K/V, not through an added fixed pooled u.
04 · Cross-attention to the source
Read the updated शुरू embedding to predict राघव
शुरू keeps its position; its final embedding is used to predict the next token.
Read the diagram labels
- The source-updated शुरू embedding continues through the MLP and its residual. After the repeated decoder blocks, the vocabulary head reads the final embedding at शुरू to produce logits, softmax probabilities, and the illustrative token Raghav. Append Raghav at target position 2 and obtain its input embedding from the target embedding table; the next step can predict Delhi. शुरू is not replaced by Raghav. Source encoder outputs stay available across steps.
- The शुरू embedding after cross-attention.
- e
1 after source update - MLP + residual
- updated शुरू e
1 - Repeat the decoder blocks; then read the final embedding at शुरू.
- Vocabulary head → logits
- Softmax → probabilities
- Choose: राघव
- Append राघव: [शुरू] राघव → next prediction: दिल्ली
- e
1 still belongs to शुरू. The new token राघव gets input embedding e2 .
Authored teaching diagram
Does the updated शुरू embedding become the embedding of राघव?
No. It remains the contextual embedding at शुरू. After the remaining decoder blocks, a vocabulary head and softmax turn it into next-token probabilities. Choosing राघव appends a new token at position 2, with its own embedding-table lookup e₂. The next generation step can then predict दिल्ली. The decoder repeats causal self-attention, cross-attention, and MLP sublayers with their residuals across its blocks. The top row shows the continuation of the block enlarged on the preceding slide. Read the final target embedding after all blocks. Each new token receives its input embedding and positional information. English encoder outputs, and each layer’s source keys and values, can be reused while the target prefix grows. Predictions are illustrative.
Teaching note
Does the updated शुरू embedding become the embedding of राघव? Trace the updated शुरू embedding through the MLP, final vocabulary scores, probabilities and chosen token. Point to both positions in the new prefix: शुरू remains first, राघव is second. The decoder repeats causal self-attention, cross-attention, and MLP sublayers with their residuals across its blocks. The top row shows the continuation of the block enlarged on the preceding slide. Read the final target embedding after all blocks. Each new token receives its input embedding and positional information. English encoder outputs, and each layer’s source keys and values, can be reused while the target prefix grows. Predictions are illustrative.
04 · Cross-attention to the source
The roles of keys and values
Query-key scores determine the weights used to combine the values.
Read the diagram labels
- Source E = [e1 … e6]. Target embedding et. V = E WV. K = E WK. qt = et WQ. scores: qt KT / √dk. softmax → αt. αt V = ct. values
- Source E = [e
1 … e6 ] - Target embedding e
t - V = E W
V - K = E W
K - q
t = et WQ - scores: q
t KT / √dk - softmax → α
t - α
t V = ct - values
Authored teaching diagram · Primary source
Why do we need both keys and values?
The query is compared with keys to produce scores. Softmax turns those scores into weights. Those weights mix the value vectors to produce cₜ. E contains final encoder embeddings. eₜ is the current target embedding after causal self-attention and before the cross-attention update. cₜ is the retrieved message, playing the same role as mᵢ earlier. Project the message to an embedding update, add it to the target stream, then apply the MLP and its residual. Different layers and heads generally use different projection parameters.
Teaching note
Why do we need both keys and values? Follow the K path into the scores, then the separate V path into the weighted sum. The query affects the sum through its weights. E contains final encoder embeddings. eₜ is the current target embedding after causal self-attention and before the cross-attention update. cₜ is the retrieved message, playing the same role as mᵢ earlier. Project the message to an embedding update, add it to the target stream, then apply the MLP and its residual. Different layers and heads generally use different projection parameters.
04 · Cross-attention to the source
The cross-attention calculation
cₜ is the cross-attention message, just as mᵢ was the self-attention message.
Read the diagram labels
- One target query, all source keys and values. qt = et WQ. ki = ei WK. vi = ei WV. αti = softmax over i (qt kiT / √dk). ct = Σi αti vi
- One target query, all source keys and values
- q
t = et WQ - k
i = ei WK - v
i = ei WV - α
ti = softmax over i (qt ki T / √dk ) - c
t = Σi αti vi
Authored teaching diagram · Primary source
Which dimension does this softmax normalize?
The source-position index i. For each target query, the weights over all permitted source positions sum to one. One head is shown. Multiple head messages are concatenated and projected back to model width before combining with the target stream. All equations use row-vector notation.
Teaching note
Which dimension does this softmax normalize? Connect each equation to its corresponding arrow in the preceding diagram. Separate scores, weights and the weighted message. One head is shown. Multiple head messages are concatenated and projected back to model width before combining with the target stream. All equations use row-vector notation.
04 · Cross-attention to the source
Multi-head cross-attention uses the same mechanism
Each head uses target queries and source keys and values.
Read the diagram labels
- Source embeddings E. Target embeddings D. head 1. E → K,V. D → Q. Attention → C1. head 2. E → K,V. D → Q. Attention → C2. head 3. E → K,V. D → Q. Attention → C3. Concatenate(C1, C2, C3) → WO. Qj = D WQ(j) · Kj = E WK(j) · Vj = E WV(j). Cj = softmax(Qj KjT / √dk) Vj
- Source embeddings E
- Target embeddings D
- head 1
- E → K,V
- D → Q
- Attention → C
1 - head 2
- E → K,V
- D → Q
- Attention → C
2 - head 3
- E → K,V
- D → Q
- Attention → C
3 - Concatenate(C
1 , C2 , C3 ) → WO - Q
j = D WQ (j) · Kj = E WK (j) · Vj = E WV (j) - C
j = softmax(Qj Kj T / √dk ) Vj
Authored teaching diagram · Primary source
What changes when we use several cross-attention heads?
Each head has its own learned query, key and value projections. Every head still takes queries from D (current target embeddings) and keys and values from E (final source embeddings). Concatenate the head outputs, then apply W_O. For head j: Qⱼ = D W_Q⁽ʲ⁾, Kⱼ = E W_K⁽ʲ⁾, Vⱼ = E W_V⁽ʲ⁾. Cⱼ = softmax(QⱼKⱼᵀ / √dₖ)Vⱼ. W_O maps the concatenated messages to model width before the residual update.
Teaching note
What changes when we use several cross-attention heads? Trace the same E and D into all three heads, then join their outputs. Students already know the multi-head calculation; stop after identifying the two origins. For head j: Qⱼ = D W_Q⁽ʲ⁾, Kⱼ = E W_K⁽ʲ⁾, Vⱼ = E W_V⁽ʲ⁾. Cⱼ = softmax(QⱼKⱼᵀ / √dₖ)Vⱼ. W_O maps the concatenated messages to model width before the residual update.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Hindi generation: self-attention, then cross-attention
Hindi context grows; English context stays fixed. Both sets of weights are illustrative.
Read the diagram labels
- Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
1 at [शुरू] - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 1.00
- m
1 = Σ αi vi · Hindi values - Δe
1 = m1 WO ; add to current e1 - Updated e
1 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
1 = updated e1 WQ (cross-attention) - q
1 KT / √dk → softmax over English - 0.72
- 0.04
- 0.03
- 0.09
- 0.04
- 0.08
- c
1 = Σ αi vi · English values - Cross update: Δe
1 = c1 WO → add to e1 → MLP + residual; finish blocks - Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
2 at राघव - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.35
- राघव
- 0.65
- m
2 = Σ αi vi · Hindi values - Δe
2 = m2 WO ; add to current e2 - Updated e
2 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
2 = updated e2 WQ (cross-attention) - q
2 KT / √dk → softmax over English - 0.05
- 0.03
- 0.04
- 0.10
- 0.08
- 0.70
- c
2 = Σ αi vi · English values - Cross update: Δe
2 = c2 WO → add to e2 → MLP + residual; finish blocks - Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
3 at दिल्ली - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.15
- राघव
- 0.25
- दिल्ली
- 0.60
- m
3 = Σ αi vi · Hindi values - Δe
3 = m3 WO ; add to current e3 - Updated e
3 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
3 = updated e3 WQ (cross-attention) - q
3 KT / √dk → softmax over English - 0.04
- 0.04
- 0.03
- 0.09
- 0.70
- 0.10
- c
3 = Σ αi vi · English values - Cross update: Δe
3 = c3 WO → add to e3 → MLP + residual; finish blocks - Step 3: vocabulary head → softmax → में · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
4 at में - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.10
- राघव
- 0.15
- दिल्ली
- 0.45
- में
- 0.30
- m
4 = Σ αi vi · Hindi values - Δe
4 = m4 WO ; add to current e4 - Updated e
4 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
4 = updated e4 WQ (cross-attention) - q
4 KT / √dk → softmax over English - 0.08
- 0.06
- 0.04
- 0.70
- 0.05
- 0.07
- c
4 = Σ αi vi · English values - Cross update: Δe
4 = c4 WO → add to e4 → MLP + residual; finish blocks - Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
5 at स्कूल - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.08
- राघव
- 0.12
- दिल्ली
- 0.20
- में
- 0.25
- स्कूल
- 0.35
- m
5 = Σ αi vi · Hindi values - Δe
5 = m5 WO ; add to current e5 - Updated e
5 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
5 = updated e5 WQ (cross-attention) - q
5 KT / √dk → softmax over English - 0.04
- 0.75
- 0.04
- 0.08
- 0.04
- 0.05
- c
5 = Σ αi vi · English values - Cross update: Δe
5 = c5 WO → add to e5 → MLP + residual; finish blocks - Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
6 at जाता - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.15
- में
- 0.15
- स्कूल
- 0.25
- जाता
- 0.30
- m
6 = Σ αi vi · Hindi values - Δe
6 = m6 WO ; add to current e6 - Updated e
6 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
6 = updated e6 WQ (cross-attention) - q
6 KT / √dk → softmax over English - 0.10
- 0.35
- 0.05
- 0.20
- 0.15
- 0.15
- c
6 = Σ αi vi · English values - Cross update: Δe
6 = c6 WO → add to e6 → MLP + residual; finish blocks - Step 6: vocabulary head → softmax → है · append to Hindi prefix
- 1 · Hindi self-attention
- 2 · English cross-attention
- Self query from e
7 at है - Keys and values: supplied Hindi positions
- Same six English encoder outputs
- Source keys K and values V
- [शुरू]
- 0.05
- राघव
- 0.10
- दिल्ली
- 0.10
- में
- 0.15
- स्कूल
- 0.15
- जाता
- 0.20
- है
- 0.25
- m
7 = Σ αi vi · Hindi values - Δe
7 = m7 WO ; add to current e7 - Updated e
7 after self-attention - Raghav
- goes
- to
- school
- in
- Delhi
- q
7 = updated e7 WQ (cross-attention) - q
7 KT / √dk → softmax over English - 0.16
- 0.18
- 0.16
- 0.17
- 0.16
- 0.17
- c
7 = Σ αi vi · English values - Cross update: Δe
7 = c7 WO → add to e7 → MLP + residual; finish blocks - Step 7: vocabulary head → softmax → [समाप्त] · stop
Authored teaching diagram
What can each attention operation read at this generation step?
Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
Teaching note
What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.
04 · Cross-attention to the source
Fixed context and query-dependent context
Source states stay fixed; the query-dependent message changes.
Read the diagram labels
- ONE FIXED SOURCE VECTOR. SOURCE STATES RETAINED. e1 … e6 → Pool → c. e1 … e6 → K, V. c. predict y1. q1. lookup. c1. c. predict y2. q2. lookup. c2. c. predict y3. q3. lookup. c3
- ONE FIXED SOURCE VECTOR
- SOURCE STATES RETAINED
- e
1 … e6 → Pool → c - e
1 … e6 → K, V - c
- predict y
1 - q
1 - lookup
- c
1 - c
- predict y
2 - q
2 - lookup
- c
2 - c
- predict y
3 - q
3 - lookup
- c
3
Authored teaching diagram · Primary source
Do we re-encode the English sentence for each Hindi token?
No. We can reuse its encoder states. The changing target query produces a potentially different weighted source message at each target step.
Teaching note
Do we re-encode the English sentence for each Hindi token? Contrast the one-vector source interface with retaining all source states. Do not imply that fixed-context decoding has an unchanging target state.
04 · Cross-attention to the source
One decoder block, two attention operations
Residuals + normalization omitted for clarity.
Read the diagram labels
- Target embeddings + positions. Causal self-attention. Cross-attention. Feed-forward network. Q. Encoder states. K,V. target state → vocabulary head
- Target embeddings + positions
- Causal self-attention
- Cross-attention
- Feed-forward network
- Q
- Encoder states
- K,V
- target state → vocabulary head
Authored teaching diagram · Primary source
Why does the decoder need both attention operations?
Causal self-attention supplies target-prefix context. Cross-attention retrieves from the complete source. The feed-forward sublayer then transforms each target position.
Teaching note
Why does the decoder need both attention operations? Trace the target stream down the block, stopping at the side connection from the source.
05 · Attention patterns and applications
Part V · Attention patterns and applications
Read the diagram labels
- Compare what each attention operation can read.. Full, causal and cross-attention. Task predictions. Connect the attention pattern to the required output.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Compare what each attention operation can read.
- Full, causal and cross-attention
- Task predictions
- Connect the attention pattern to the required output.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part V · Attention patterns and applications
Compare what each attention operation can read.
Teaching note
Part V · Attention patterns and applications Pause and establish the requested output before advancing.
05 · Attention patterns and applications
Attention patterns and their applications
Filled cells show allowed connections. The task probability comes from the final head.
Read the diagram labels
- Three allowed-access matrices with concrete tasks. Encoder: six English query rows by six English key/value columns, full access, followed by a topic summary and classification head. Decoder-only continuation: three supplied English prefix positions with a causal triangle, then vocabulary prediction of school. Cross-attention: two Hindi target positions, start and Raghav, querying all six English positions to help predict Delhi. Filled cells are permission, not numerical attention weights. Detailed word-labeled matrices follow.
- English example: Raghav goes to school in Delhi
- Encoder self-attention
- Classify the topic: EDUCATION
- English → English
- Q, K, V: English
- p(label | complete input)
- Decoder self-attention
- Continue: Raghav goes to …
- Prefix → prefix
- Q, K, V: prefix
- p(school | Raghav goes to)
- Cross-attention
- Translate: [शुरू] राघव …
- Hindi → English
- Q: Hindi · K,V: English
- p(दिल्ली | English, Hindi prefix)
Authored teaching diagram
How does the attention pattern relate to the task?
An encoder can read the complete English input to build representations for a sentence classifier. A causal decoder reads the supplied prefix to support next-token prediction. Cross-attention lets Hindi target positions consult the English source while translating. The following slides label every row and column with the actual words. These are three illustrative task settings, not exclusive uses of each architecture. The encoder example has six positions, decoder continuation has three supplied positions, and cross-attention has two Hindi positions by six English positions. Filled cells are permitted access, not measured attention weights or probabilities. The decoder can use previous words and the current supplied token; it cannot see future target words. The encoder returns one vector per input position, with pooling or a selected token used to obtain a summary when the task needs one.
Teaching note
How does the attention pattern relate to the task? Compare the full English square, the causal prefix triangle and the Hindi-to-English rectangle. Read the task probability below each one. These are three illustrative task settings, not exclusive uses of each architecture. The encoder example has six positions, decoder continuation has three supplied positions, and cross-attention has two Hindi positions by six English positions. Filled cells are permitted access, not measured attention weights or probabilities. The decoder can use previous words and the current supplied token; it cannot see future target words. The encoder returns one vector per input position, with pooling or a selected token used to obtain a summary when the task needs one.
05 · Attention patterns and applications
Encoder attention for sentence classification
Example task: classify the sentence’s topic using all supplied words.
Read the diagram labels
- English encoder self-attention with actual word-labeled query rows and key/value columns. Every source position can read all six input positions. The highlighted school row can attend to Raghav, goes, to, school, in and Delhi. The encoder returns six contextual embeddings. For this illustrative sentence-topic classification task, mean pool the outputs, then use a classifier and softmax for p(y equals EDUCATION given the complete sentence x). Pooling is a task choice; an encoder does not inherently output just one summary.
- Input x: Raghav goes to school in Delhi
- Columns: English keys / values
- Raghav
- goes
- to
- school
- in
- Delhi
- Raghav
- goes
- to
- school
- in
- Delhi
- Rows: English queries · all six words are visible
- Encoder → 6 updated embeddings
- Mean pool → one sentence vector
- Topic head → logits → softmax
- p(y = EDUCATION | x)
- The encoder keeps all positions. Pooling gives the summary used by this task.
Authored teaching diagram
Does the encoder itself reduce six words to one summary?
No. It returns six updated token embeddings. In this example we mean-pool them into one vector, then apply a topic classifier and softmax to obtain p(y | x), where x is the complete sentence and y is a topic label. EDUCATION is an illustrative label. Rows are positions whose queries receive messages; columns are positions supplying keys and values. Both axes come from the same English input. CLS selection is another summary choice, and token labeling could instead use every output embedding. This slide uses pooling to keep the six-token attention matrix literal. No trained prediction or numerical attention values are claimed.
Teaching note
Does the encoder itself reduce six words to one summary? Read the highlighted school row against all six column words. Follow all encoder outputs into mean pooling and the topic head. Rows are positions whose queries receive messages; columns are positions supplying keys and values. Both axes come from the same English input. CLS selection is another summary choice, and token labeling could instead use every output embedding. This slide uses pooling to keep the six-token attention matrix literal. No trained prediction or numerical attention values are claimed.
05 · Attention patterns and applications
Causal attention for next-token prediction
The current and previous supplied words are visible; future words are not.
Read the diagram labels
- Decoder causal self-attention over the supplied prefix Raghav goes to. Rows and columns are Raghav, goes and to. Raghav reads itself; goes reads Raghav and itself; to reads all three supplied positions. The future school token is absent. Finish the decoder blocks and use the final to embedding e3 for vocabulary logits and softmax, modeling p(school given Raghav goes to). Self-attention updates embeddings; the final vocabulary head produces the next-token distribution.
- Supplied prefix: Raghav goes to → next token: school
- Columns: supplied prefix keys / values
- Raghav
- goes
- to
- Raghav
- goes
- to
- Causal attention updates embeddings
- Finish blocks → read last e
3 - Vocabulary logits → softmax
- The “to” row can read Raghav, goes and to.
- p(school | Raghav goes to)
- Previous words are available; future words cannot be read.
Authored teaching diagram
Can the “to” position attend to the previous word “goes”?
Yes. The last supplied position to can read Raghav, goes and itself. It cannot read school because school has not been supplied. After the decoder blocks, the vocabulary head reads the final e₃ and models p(school | Raghav goes to). Causal attention alone updates embeddings rather than emitting a word. More generally, next-token language modeling learns p(xₜ₊₁ | x₁,…,xₜ). During training, later tokens may be present in the batch tensor but the causal mask hides them from earlier queries. At generation time those future tokens are absent. The diagonal is allowed because the token at the receiving position is an input used to predict the following token.
Teaching note
Can the “to” position attend to the previous word “goes”? Read the three matrix rows in order. Trace the highlighted bottom row to all three input columns, then the last embedding to the vocabulary distribution. More generally, next-token language modeling learns p(xₜ₊₁ | x₁,…,xₜ). During training, later tokens may be present in the batch tensor but the causal mask hides them from earlier queries. At generation time those future tokens are absent. The diagonal is allowed because the token at the receiving position is an input used to predict the following token.
05 · Attention patterns and applications
Cross-attention between Hindi and English
Rows are Hindi target positions; columns are English source positions.
Read the diagram labels
- Cross-attention has Hindi query rows start and Raghav, and English key/value columns Raghav, goes, to, school, in, Delhi. Both target positions can read the complete English source. Queries use target embeddings after causal self-attention; keys and values use encoder outputs. The retrieved source message updates the target stream before the remaining decoder computation and vocabulary prediction. The model predicts p(Delhi given the English input x and supplied Hindi prefix start Raghav). This full rectangular access pattern is not a future-target leak: English is supplied, whereas future Hindi tokens are unavailable.
- English x: Raghav goes to school in Delhi
- Columns: English keys / values
- Raghav
- goes
- to
- school
- in
- Delhi
- [शुरू]
- राघव
- Rows: Hindi queries after causal self-attention
- Hindi prefix: [शुरू] राघव
- Q: Hindi · K and V: English
- Source message → target update
- Both Hindi positions may read all six English positions.
- Translation task: p(दिल्ली | English x, Hindi prefix [शुरू] राघव)
Authored teaching diagram
Why is the cross-attention matrix full even though the decoder is causal?
The complete English source is already supplied, so each Hindi query may read all six source positions. Causality still applies within the Hindi stream. To predict दिल्ली, use the current embedding at राघव after Hindi self-attention, then query English keys and combine English values. The translation model represents p(yₜ | x, y₁,…,yₜ₋₁), with x the complete English sentence and the supplied Hindi prefix beginning with [शुरू]. The probability of a Hindi output token is produced by the final vocabulary head after target updates, not directly by the attention matrix. Attention weights normalize over English positions; vocabulary probabilities normalize over possible Hindi tokens. Future Hindi tokens are not available to these queries.
Teaching note
Why is the cross-attention matrix full even though the decoder is causal? Read the highlighted राघव row across the six English column words. Contrast the two languages on the axes and read the conditional probability aloud. The translation model represents p(yₜ | x, y₁,…,yₜ₋₁), with x the complete English sentence and the supplied Hindi prefix beginning with [शुरू]. The probability of a Hindi output token is produced by the final vocabulary head after target updates, not directly by the attention matrix. Attention weights normalize over English positions; vocabulary probabilities normalize over possible Hindi tokens. Future Hindi tokens are not available to these queries.
06 · Comparing model architectures
Part VI · Comparing model architectures
Read the diagram labels
- Compare the architectures used for the example tasks.. Inputs and outputs. Architecture choice. Compare attention patterns and task heads.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
- Compare the architectures used for the example tasks.
- Inputs and outputs
- Architecture choice
- Compare attention patterns and task heads.
- 01 · Generation
- 02 · Input labels
- 03 · Translation
- 04 · Cross-attention
- 05 · Patterns
- 06 · Model families
Authored teaching diagram
Part VI · Comparing model architectures
Compare the architectures used for the example tasks.
Teaching note
Part VI · Comparing model architectures Pause and establish the requested output before advancing.
06 · Comparing model architectures
Encoder, decoder-only and encoder-decoder models
Compare the input, attention connections and output of each architecture.
Read the diagram labels
- ENCODER. Encode the input. DECODER. Predict the next token. ENCODER-DECODER. Generate from a source. . . . . . . . Encoder. . . . . . . one vector per token. labels · retrieval · readouts. . . . . Decoder. next. append → repeat. chat · code · continuation. Encoder. . . . target prefix. source K,V. Decoder. next. source + target prefix. translation · summarization
- ENCODER
- Encode the input
- DECODER
- Predict the next token
- ENCODER-DECODER
- Generate from a source
- Encoder
- one vector per token
- labels · retrieval · readouts
- Decoder
- next
- append → repeat
- chat · code · continuation
- Encoder
- target prefix
- source K,V
- Decoder
- next
- source + target prefix
- translation · summarization
Authored teaching diagram · Primary source
What did we add to connect understanding a source with generating a target?
A source encoder and a pathway from its representations into the target decoder. Our final arrangement uses cross-attention for that source pathway.
Teaching note
What did we add to connect understanding a source with generating a target? Compare all three architectures only now, after deriving them from the tasks.
06 · Comparing model architectures
Choosing an architecture for a task
These are common choices; other architectures can perform the same tasks.
Read the diagram labels
- One label for the input. Encoder + readout. sentiment. One label per token. Encoder + shared head. NER. Continue a prefix. Decoder-only. text generation. Turn source X into target Y. Encoder-decoder. English → Hindi
- One label for the input
- Encoder + readout
- sentiment
- One label per token
- Encoder + shared head
- NER
- Continue a prefix
- Decoder-only
- text generation
- Turn source X into target Y
- Encoder-decoder
- English → Hindi
Authored teaching diagram · Primary source
Can a decoder-only model also produce a class label or translation?
Yes. It can generate labels or translations from a suitable prompt and training setup. These rows illustrate useful arrangements rather than exclusive capabilities.
Teaching note
Can a decoder-only model also produce a class label or translation? Ask for a task and have students trace its desired output shape.
06 · Comparing model architectures
Choosing an architecture
The output type and source pathway guide the choice of architecture.
Read the diagram labels
- Do I need to generate new tokens?. NO. YES. Encoder. representations. labels / embeddings. Separate source pathway?. NO. YES. Decoder-only. Encoder-decoder
- Do I need to generate new tokens?
- NO
- YES
- Encoder
- representations
- labels / embeddings
- Separate source pathway?
- NO
- YES
- Decoder-only
- Encoder-decoder
Authored teaching diagram · Primary source
Does having source text force us to use an encoder-decoder?
No. This branch asks whether the chosen design has a separate source pathway. A decoder-only model can instead consume source text within its prefix.
Teaching note
Does having source text force us to use an encoder-decoder? Ask whether the task needs new tokens. For a generated output, ask whether the design uses a separate source pathway. Representations can support labels, embeddings or token outputs.
06 · Comparing model architectures
Review: attention, representations and training
Read the diagram labels
- 01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?. Next lecture: what if our tokens are not words?
- 01
- WHICH POSITIONS CAN ATTEND TO EACH OTHER?
- 02
- WHICH EMBEDDINGS DOES THE TASK USE?
- 03
- WHAT IS THE TRAINING TARGET?
- Next lecture: what if our tokens are not words?
Authored teaching diagram
For NER, classification and translation: who reads whom, what do we read, and what trains it?
NER uses whole-input attention, all token states and token labels. Classification uses a whole-input readout and one sentence label. Translation uses source encoding, causal target states, cross-attention and next-target-token supervision.
Teaching note
For NER, classification and translation: who reads whom, what do we read, and what trains it? Let students reconstruct each example, then reveal the teaser about tokens that are not words.
Appendix · Optional worked examples
Appendix: worked examples
Read the diagram labels
- Tensor shapes, training and a numerical attention example.. Shapes and training. Numerical examples. Optional details for discussion and practice.
- Tensor shapes, training and a numerical attention example.
- Shapes and training
- Numerical examples
- Optional details for discussion and practice.
Authored teaching diagram
Appendix: worked examples
Tensor shapes, training and a numerical attention example.
Teaching note
Appendix: worked examples Pause and establish the requested output before advancing.
Appendix · Optional worked examples
Tensor shapes in cross-attention
Rows are target queries; columns are source positions.
Read the diagram labels
- E ∈ ℝn × d · source states. D ∈ ℝm × d · target states. d = model width. . K ∈ ℝn × dk. . V ∈ ℝn × dv. . Q ∈ ℝm × dk. . A = softmax(QKT / √dk) ∈ ℝm × n. . C = AV ∈ ℝm × dv. Example: m = 4, n = 6 → A is 4 × 6
- E ∈ ℝ
n × d · source states - D ∈ ℝ
m × d · target states - d = model width
- K ∈ ℝ
n × d k - V ∈ ℝ
n × d v - Q ∈ ℝ
m × d k - A = softmax(QK
T / √dk ) ∈ ℝm × n - C = AV ∈ ℝ
m × d v - Example: m = 4, n = 6 → A is 4 × 6
Authored teaching diagram · Primary source
Must source and target sequence lengths match?
No. For m target queries and n source positions, A has shape m × n. Each row sums to one. C = AV has shape m × dᵥ per head. With m = 4 and n = 6, the weight matrix is 4 × 6. This diagram uses a common model width for simplicity; learned projections can also map different source and target widths into a shared query/key width.
Teaching note
Must source and target sequence lengths match? Multiply the inner dimensions aloud. At generation time highlight one query row. This diagram uses a common model width for simplicity; learned projections can also map different source and target widths into a shared query/key width.
Appendix · Optional worked examples
Conditioning with a continuous source prefix
The target can read the source prefix through causal self-attention.
Read the diagram labels
- c = Pool(source states). u = c WP · decoder width. u. [शुरू]. राघव. दिल्ली. continuous prefix. ordinary target embeddings + positions. Causal Transformer → next Hindi token. [CONTEXT] denotes an injected vector, not a new vocabulary word.
- c = Pool(source states)
- u = c W
P · decoder width - u
- [शुरू]
- राघव
- दिल्ली
- continuous prefix
- ordinary target embeddings + positions
- Causal Transformer → next Hindi token
- [CONTEXT] denotes an injected vector, not a new vocabulary word.
Authored teaching diagram · Primary source
Is [CONTEXT] a word the decoder needs to predict?
No. Here it labels an injected continuous input vector u = c W_P. Prepend it to the positioned target embeddings, start with [शुरू], and read the last target position to predict the next word. Give the prefix slot a consistent position, and train this construction end to end on translation targets. The initial u is source-dependent and fixed across decoding steps; target hidden states continue to depend on the growing target prefix. The prefix slot itself has no language-model target in this construction.
Teaching note
Is [CONTEXT] a word the decoder needs to predict? Trace c through the width projection to the prefix slot, then through the ordinary causal stack. Give the prefix slot a consistent position, and train this construction end to end on translation targets. The initial u is source-dependent and fixed across decoding steps; target hidden states continue to depend on the growing target prefix. The prefix slot itself has no language-model target in this construction.
Appendix · Optional worked examples
Next-token prediction during training
Shift the target sequence by one position.
Read the diagram labels
- Source: Raghav goes to school in Delhi.. Decoder input. [शुरू]. राघव. दिल्ली. में. स्कूल. जाता. है. राघव. दिल्ली. में. स्कूल. जाता. है. [समाप्त]. Training target. L = −Σt log p(yt | y<t, x)
- Source: Raghav goes to school in Delhi.
- Decoder input
- [शुरू]
- राघव
- दिल्ली
- में
- स्कूल
- जाता
- है
- राघव
- दिल्ली
- में
- स्कूल
- जाता
- है
- [समाप्त]
- Training target
- L = −Σ
t log p(yt | y<t , x)
Authored teaching diagram · Primary source
Can the decoder see the target it is currently supposed to predict?
Not at the prediction position: its input is the previous target token and its self-attention is causal. The source remains available through cross-attention. Teacher forcing supplies the known target prefix during training. With causal masking, all these next-token losses can be computed in one parallel pass.
Teaching note
Can the decoder see the target it is currently supposed to predict? Trace each input column to the different target below it, including [शुरू] and [समाप्त]. Teacher forcing supplies the known target prefix during training. With causal masking, all these next-token losses can be computed in one parallel pass.
Appendix · Optional worked examples
Training versus generation
The next-token objective is the same; prefix construction differs.
Read the diagram labels
- TRAINING. GENERATION. correct previous target token. previous model prediction. predict next token. predict next token. Known target prefixes supplied. Append prediction; repeat
- TRAINING
- GENERATION
- correct previous target token
- previous model prediction
- predict next token
- predict next token
- Known target prefixes supplied
- Append prediction; repeat
Authored teaching diagram · Primary source
What is fed back at inference time?
The model’s selected prediction. In teacher-forced training we supply the observed previous target token instead.
Teaching note
What is fed back at inference time? Contrast how each prefix is built. Avoid decoding algorithms here.
Appendix · Optional worked examples
Attention head ≠ classification head
Attention mixes states; a task head predicts labels.
Read the diagram labels
- Attention heads. Classification head. mix information. map states to label scores. Their counts are independent.
- Attention heads
- Classification head
- mix information
- map states to label scores
- Their counts are independent.
Authored teaching diagram · Primary source
Do four entity labels imply four attention heads?
No. The number of attention heads and the number of classes are independent. They name different computations.
Teaching note
Do four entity labels imply four attention heads? Locate the attention heads inside the Transformer and the classifier after it.
Appendix · Optional worked examples
A numerical cross-attention example
Chosen scaled scores; probabilities recomputed with softmax.
Read the diagram labels
- One query · four source keys · chosen scaled scores. Raghav. 0.2. 0.110. goes. 0.1. 0.100. school. 0.3. 0.122. Delhi. 2.0. 0.668. softmax across source positions → ct = Σi αti vi
- One query · four source keys · chosen scaled scores
- Raghav
- 0.2
- 0.110
- goes
- 0.1
- 0.100
- school
- 0.3
- 0.122
- Delhi
- 2.0
- 0.668
- softmax across source positions → c
t = Σi αti vi
Authored teaching diagram · Primary source
Which source position gets the most weight?
Delhi. Softmax of scaled scores [0.2, 0.1, 0.3, 2.0] is approximately [0.110, 0.100, 0.122, 0.668]. These weights mix the four source value vectors. This deliberately reduced four-key example is separate from the six-word translation diagram. Scores already include division by √dₖ. No learned model or actual alignment is being claimed.
Teaching note
Which source position gets the most weight? Normalize the four exponentials, then point back to the familiar weighted sum. This deliberately reduced four-key example is separate from the six-word translation diagram. Scores already include division by √dₖ. No learned model or actual alignment is being claimed.
Appendix · Optional worked examples
Computing the weighted sum of values
The weights mix values into cₜ ≈ [1.57, 0.89].
Read the diagram labels
- source. weight α. value vector V. Raghav. 0.110. [1, 0]. goes. 0.100. [0, 1]. school. 0.122. [1, 1]. Delhi. 0.668. [2, 1]. ct = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1]. ct ≈ [1.57, 0.89]
- source
- weight α
- value vector V
- Raghav
- 0.110
- [1, 0]
- goes
- 0.100
- [0, 1]
- school
- 0.122
- [1, 1]
- Delhi
- 0.668
- [2, 1]
- c
t = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1] - c
t ≈ [1.57, 0.89]
Authored teaching diagram · Primary source
What does cross-attention return after softmax?
A weighted sum of value vectors. With the four values shown, cₜ is approximately [1.57, 0.89]. It is a feature vector, not a class label or a word probability distribution. The exact softmax weights give cₜ = [1.567881402964, 0.889620530623]. The slide rounds the weights to three decimals and the result to two. Both agree at the displayed result precision. Values and scores are chosen for this calculation.
Teaching note
What does cross-attention return after softmax? Multiply each value by its weight, then add coordinate by coordinate. The Delhi value contributes most in this chosen example. The exact softmax weights give cₜ = [1.567881402964, 0.889620530623]. The slide rounds the weights to three decimals and the result to two. Both agree at the displayed result precision. Values and scores are chosen for this calculation.