Deep learning · Transformer architectures

From Attention to Applications

From attention mechanisms to three Transformer architectures and their tasks.Encoder, decoder-only and encoder-decoder TransformersENCODERRead the complete inputToken and sentence labelsDECODER-ONLYRead the available prefixNext-token generationENCODER-DECODERRead source and target prefixConditional generation
Read the diagram labels
  • From attention mechanisms to three Transformer architectures and their tasks.
  • Encoder, decoder-only and encoder-decoder Transformers
  • ENCODER
  • Read the complete input
  • Token and sentence labels
  • DECODER-ONLY
  • Read the available prefix
  • Next-token generation
  • ENCODER-DECODER
  • Read source and target prefix
  • Conditional generation

Authored teaching diagram

How do attention blocks become models for different tasks?

We will use familiar attention blocks to construct decoder-only, encoder and encoder-decoder models, then examine how cross-attention connects a target sequence to its source.

Teaching note

How do attention blocks become models for different tasks? Introduce the lecture, then recall the attention mechanisms studied previously.

Frame overflows: split the content

Recap · Self-attention, multiple heads and positions

Recap: the Transformer components we have studied

1 · Positional encoding. 2 · Self-attention. 3 · Multiple heads. token embedding + position. ei + pi. Token identity and order. Q, K, V from embeddings. scores → softmax → mi. A weighted sum of values. head 1. head 2. head 3. Concatenate messages. Project with WO → Δei. Residual update: ei ← ei + Δei. Then the MLP and its residual update. Result: one contextualised embedding at every token position1 · Positional encoding2 · Self-attention3 · Multiple headstoken embedding + positionei + piToken identity and orderQ, K, V from embeddingsscores → softmax → miA weighted sum of valueshead 1head 2head 3Concatenate messagesProject with WO → ΔeiResidual update: ei ← ei + ΔeiThen the MLP and its residual updateResult: one contextualised embedding at every token position
Read the diagram labels
  • 1 · Positional encoding. 2 · Self-attention. 3 · Multiple heads. token embedding + position. ei + pi. Token identity and order. Q, K, V from embeddings. scores → softmax → mi. A weighted sum of values. head 1. head 2. head 3. Concatenate messages. Project with WO → Δei. Residual update: ei ← ei + Δei. Then the MLP and its residual update. Result: one contextualised embedding at every token position
  • 1 · Positional encoding
  • 2 · Self-attention
  • 3 · Multiple heads
  • token embedding + position
  • ei + pi
  • Token identity and order
  • Q, K, V from embeddings
  • scores → softmax → mi
  • A weighted sum of values
  • head 1
  • head 2
  • head 3
  • Concatenate messages
  • Project with WO → Δei
  • Residual update: ei ← ei + Δei
  • Then the MLP and its residual update
  • Result: one contextualised embedding at every token position

Authored teaching diagram · Primary source

What does each attention block change?

Positional encoding supplies order information to token embeddings. Self-attention forms queries, keys and values from the same sequence and combines values into a message for each position. Multiple heads use separate learned projections; their messages are concatenated and projected into an embedding update. Residual addition and the MLP produce contextualised embeddings.

Teaching note

What does each attention block change? Recall the additive positional encoding studied earlier: eᵢ + pᵢ. The eᵢ in the residual equation denotes the current embedding entering that sublayer. Each head uses its own Q/K/V projections. The diagram uses the existing message mᵢ and update Δeᵢ convention; normalization is omitted. Today we vary which positions can attend and which output embeddings the task uses.

Frame overflows: split the content

00 · What should the Transformer output?

Different tasks for the same input

Raghav goes to school in Delhi.Raghav goes to school in Delhi.

What should the Transformer output?

Read the diagram labels
  • Raghav goes to school in Delhi.
  • Raghav goes to school in Delhi.

Authored teaching diagram

What kinds of output could we request for this sentence?

A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Teaching note

What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Frame overflows: split the content

00 · What should the Transformer output?

Different tasks for the same input

Raghav goes to school in Delhi.. What comes next?. A continuationRaghav goes to school in Delhi.What comes next?A continuation

What should the Transformer output?

Read the diagram labels
  • Raghav goes to school in Delhi.. What comes next?. A continuation
  • Raghav goes to school in Delhi.
  • What comes next?
  • A continuation

Authored teaching diagram

What kinds of output could we request for this sentence?

A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Teaching note

What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Frame overflows: split the content

00 · What should the Transformer output?

Different tasks for the same input

Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACERaghav goes to school in Delhi.What comes next?A continuationWhat is each token?PERSON … PLACE

What should the Transformer output?

Read the diagram labels
  • Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACE
  • Raghav goes to school in Delhi.
  • What comes next?
  • A continuation
  • What is each token?
  • PERSON … PLACE

Authored teaching diagram

What kinds of output could we request for this sentence?

A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Teaching note

What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Frame overflows: split the content

00 · What should the Transformer output?

Different tasks for the same input

Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACE. Say it in Hindi. राघव दिल्ली में स्कूल जाता है।Raghav goes to school in Delhi.What comes next?A continuationWhat is each token?PERSON … PLACESay it in Hindiराघव दिल्ली में स्कूल जाता है।

What should the Transformer output?

Read the diagram labels
  • Raghav goes to school in Delhi.. What comes next?. A continuation. What is each token?. PERSON … PLACE. Say it in Hindi. राघव दिल्ली में स्कूल जाता है।
  • Raghav goes to school in Delhi.
  • What comes next?
  • A continuation
  • What is each token?
  • PERSON … PLACE
  • Say it in Hindi
  • राघव दिल्ली में स्कूल जाता है।

Authored teaching diagram

What kinds of output could we request for this sentence?

A continuation, a label at each input position, or a Hindi translation. The supplied information and desired output determine the computation. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Teaching note

What kinds of output could we request for this sentence? Keep the sentence fixed. Reveal one branch per advance; collect suggested answers before naming any architecture. Throughout the lecture, each displayed word is treated as one token for clarity; actual tokenization may split words. Punctuation is omitted from the token strips. Labels and continuations are authored teaching examples, not model predictions.

Frame overflows: split the content

00 · What should the Transformer output?

Review: contextual token embeddings

Raghav. goes. to. school. in. Delhi. Token embeddings + positions → Transformer blocks. . e1. . e2. . e3. . e4. . e5. . e6. ei = final contextualised embedding · d numbers. E = final source embeddings stacked as rows · 6 × dRaghavgoestoschoolinDelhiToken embeddings + positions → Transformer blockse1e2e3e4e5e6ei = final contextualised embedding · d numbersE = final source embeddings stacked as rows · 6 × d

The subscript tells us which token; the diagram shows its embedding after the blocks.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. Token embeddings + positions → Transformer blocks. . e1. . e2. . e3. . e4. . e5. . e6. ei = final contextualised embedding · d numbers. E = final source embeddings stacked as rows · 6 × d
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Token embeddings + positions → Transformer blocks
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • ei = final contextualised embedding · d numbers
  • E = final source embeddings stacked as rows · 6 × d

Authored teaching diagram · Primary source

Does a Transformer automatically return just one summary vector?

No. Each token retains an embedding as it passes through the blocks. We write its final contextualised embedding as eᵢ, containing d numbers. E stacks the final source embeddings as rows; here its shape is 6 × d.

Teaching note

Does a Transformer automatically return just one summary vector? Trace Raghav to e₁, then goes to e₂. Recall eᵢ → attention message mᵢ → update Δeᵢ → e′ᵢ = eᵢ + Δeᵢ → MLP and residual. We keep the token-position subscript as its embedding is updated; the diagram labels identify the stage. These final embeddings are not the attention messages.

Frame overflows: split the content

00 · What should the Transformer output?

Attention, output representations and training targets

01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?01WHICH POSITIONS CAN ATTEND TO EACH OTHER?02WHICH EMBEDDINGS DOES THE TASK USE?03WHAT IS THE TRAINING TARGET?
Read the diagram labels
  • 01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?
  • 01
  • WHICH POSITIONS CAN ATTEND TO EACH OTHER?
  • 02
  • WHICH EMBEDDINGS DOES THE TASK USE?
  • 03
  • WHAT IS THE TRAINING TARGET?

Authored teaching diagram

What should we identify before naming a model family?

The allowed information flow, the states we read, and the targets used for training.

Teaching note

What should we identify before naming a model family? Return to these questions whenever the task changes.

Frame overflows: split the content

01 · Generate the next token

Part I · Generate the next token

Review of the causal Transformer.. Raghav goes to. school …. Predict the next token from the available prefix.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesReview of the causal Transformer.Raghav goes toschool …Predict the next token from the available prefix.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Review of the causal Transformer.. Raghav goes to. school …. Predict the next token from the available prefix.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Review of the causal Transformer.
  • Raghav goes to
  • school …
  • Predict the next token from the available prefix.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part I · Generate the next token

Review of the causal Transformer.

Teaching note

Part I · Generate the next token Pause and establish the requested output before advancing.

Frame overflows: split the content

01 · Generate the next token

Review: predicting the next token

Raghav. goes. to. e1. e2. e3. Causal attention · query from position 3. m3 = Σj α3j vj. message received by “to”. Δe3 = m3 WO. e′3 = e3 + Δe3. keep the original e3. MLP + residual. e3 · final embedding. Vocabulary head → school (word 4). One head; final block shown.. Normalisation omitted.Raghavgoestoe1e2e3Causal attention · query from position 3m3 = Σj α3j vjmessage received by “to”Δe3 = m3 WOe′3 = e3 + Δe3keep the original e3MLP + residuale3 · final embeddingVocabulary head → school (word 4)One head; final block shown.Normalisation omitted.

m₃ updates token 3 (“to”); its final embedding predicts token 4 (“school”).

Read the diagram labels
  • Raghav. goes. to. e1. e2. e3. Causal attention · query from position 3. m3 = Σj α3j vj. message received by “to”. Δe3 = m3 WO. e′3 = e3 + Δe3. keep the original e3. MLP + residual. e3 · final embedding. Vocabulary head → school (word 4). One head; final block shown.. Normalisation omitted.
  • Raghav
  • goes
  • to
  • e1
  • e2
  • e3
  • Causal attention · query from position 3
  • m3 = Σj α3j vj
  • message received by “to”
  • Δe3 = m3 WO
  • e′3 = e3 + Δe3
  • keep the original e3
  • MLP + residual
  • e3 · final embedding
  • Vocabulary head → school (word 4)
  • One head; final block shown.
  • Normalisation omitted.

Authored teaching diagram · Primary source

Which state predicts school after Raghav goes to?

The final contextualised embedding e₃, at the position of to, goes through the vocabulary head. m₃ is the attention message received at position 3; it is projected into Δe₃, added to e₃, and followed by the MLP and its residual. It is not a message for the word being predicted. A decoder-only stack has causal self-attention and feed-forward sublayers; it does not have a separate encoder or source cross-attention. The term decoder names its generation role. Scores cover the vocabulary even though the drawing shows one selected word.

Teaching note

Which state predicts school after Raghav goes to? Trace the message, projection, residual addition and MLP. The subscript identifies the receiving input position: m₃ belongs to to; m₄ would belong to school only after school is supplied. For several heads, concatenate their messages before W_O. The figure shows the final block and omits normalisation; earlier blocks repeat the same kind of update. A decoder-only stack has causal self-attention and feed-forward sublayers; it does not have a separate encoder or source cross-attention. The term decoder names its generation role. Scores cover the vocabulary even though the drawing shows one selected word.

Frame overflows: split the content

01 · Generate the next token

Causal attention: current and earlier positions

6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. × goes. × to. × school. × in. × Delhi6 × 6 · allowed connectionsRaghav goes to school in Delhi.key columnsRRggttssiiDDRaghav can read…Raghav× goes× to× school× in× Delhi

Filled cell = connection allowed, not an attention weight.

Read the diagram labels
  • 6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. × goes. × to. × school. × in. × Delhi
  • 6 × 6 · allowed connections
  • Raghav goes to school in Delhi.
  • key columns
  • R
  • R
  • g
  • g
  • t
  • t
  • s
  • s
  • i
  • i
  • D
  • D
  • Raghav can read…
  • Raghav
  • × goes
  • × to
  • × school
  • × in
  • × Delhi

Authored teaching diagram · Primary source

Can Raghav’s state use the later word Delhi?

Not under this causal mask. A query at position i can use positions up to and including i. The training target is the following token.

Teaching note

Can Raghav’s state use the later word Delhi? Read the first row, then the last row. Filled cells mean access is allowed, not that weights are equal.

Frame overflows: split the content

01 · Generate the next token

Generating a sequence one token at a time

Raghav. e1 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. Read e1 → vocabulary head → goesRaghave1 · inputCausal Transformer blocksmi → Δei → ei + Δei → MLP + residuale1 · Raghavupdated embeddingRead e1 → vocabulary head → goes

Each word has an updated embedding. Read the last one to predict the next word.

Read the diagram labels
  • Raghav. e1 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. Read e1 → vocabulary head → goes
  • Raghav
  • e1 · input
  • Causal Transformer blocks
  • mi → Δei → ei + Δei → MLP + residual
  • e1 · Raghav
  • updated embedding
  • Read e1 → vocabulary head → goes

Authored teaching diagram · Primary source

What changes after we predict the next token?

The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Teaching note

What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Frame overflows: split the content

01 · Generate the next token

The hidden layer inside a Transformer block

Zoom into the MLP at Raghav’s position in the last block. After attention + residual. Hidden layer + activation. MLP output. 8 features. 12 hidden units. 8 features. +. . e1. updated. embedding. residual: all 8 features. Example widths: 8 → 12 → 8. The stack repeats attention and MLP updates.Zoom into the MLP at Raghav’s position in the last blockAfter attention + residualHidden layer + activationMLP output8 features12 hidden units8 features+e1updatedembeddingresidual: all 8 featuresExample widths: 8 → 12 → 8. The stack repeats attention and MLP updates.

The MLP updates the embedding; the vocabulary layer reads it afterward.

Read the diagram labels
  • Zoom into the MLP at Raghav’s position in the last block. After attention + residual. Hidden layer + activation. MLP output. 8 features. 12 hidden units. 8 features. +. . e1. updated. embedding. residual: all 8 features. Example widths: 8 → 12 → 8. The stack repeats attention and MLP updates.
  • Zoom into the MLP at Raghav’s position in the last block
  • After attention + residual
  • Hidden layer + activation
  • MLP output
  • 8 features
  • 12 hidden units
  • 8 features
  • +
  • e1
  • updated
  • embedding
  • residual: all 8 features
  • Example widths: 8 → 12 → 8. The stack repeats attention and MLP updates.

Authored teaching diagram · Primary source

Is the vocabulary head another stack of hidden layers?

In the standard decoder shown here, the hidden processing is already inside the Transformer blocks. Each block has an MLP with a hidden layer. This drawing expands the final block at Raghav’s position: 8 input features, 12 hidden units with a nonlinear activation, and 8 output features, followed by residual addition. Widths 8 and 12 are chosen for drawing, not a specific model. Normalization is omitted, and its exact placement depends on the architecture. The left vector already includes the attention update and its residual. The final e₁ denotes the updated embedding at position 1, not the embedding of the word to be generated. Multiple stacked blocks provide multiple hidden layers. Some model variants use more elaborate prediction heads, but they are not required here.

Teaching note

Is the vocabulary head another stack of hidden layers? Follow the dense connections through the MLP. Trace the residual bypass into the plus sign. The resulting updated e₁ is still eight numbers, ready for the next slide. Widths 8 and 12 are chosen for drawing, not a specific model. Normalization is omitted, and its exact placement depends on the architecture. The left vector already includes the attention update and its residual. The final e₁ denotes the updated embedding at position 1, not the embedding of the word to be generated. Multiple stacked blocks provide multiple hidden layers. Some model variants use more elaborate prediction heads, but they are not required here.

Frame overflows: split the content

01 · Generate the next token

One output neuron for every vocabulary token

Example: d = 8 · vocabulary size |V| = 50,000. Updated e1. 50,000 output neurons. logit. Raghav. 1.0. bicycle. -2.0. goes. 5.5. ##ing. -1.0. school. 1.5. banana. -3.0. to. 3.0. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. 8 features. learned weights. e1 (1 × 8) × W (8 × 50,000) + b → 50,000 logits. Selected entries shown; the dots stand for thousands of other tokens.Example: d = 8 · vocabulary size |V| = 50,000Updated e150,000 output neuronslogitRaghav1.0bicycle-2.0goes5.5##ing-1.0school1.5banana-3.0to3.0⋮⋮⋮⋮⋮⋮⋮⋮⋮8 featureslearned weightse1 (1 × 8) × W (8 × 50,000) + b → 50,000 logitsSelected entries shown; the dots stand for thousands of other tokens.

Every vocabulary token gets a score, including tokens unrelated to the input.

Read the diagram labels
  • Example: d = 8 · vocabulary size |V| = 50,000. Updated e1. 50,000 output neurons. logit. Raghav. 1.0. bicycle. -2.0. goes. 5.5. ##ing. -1.0. school. 1.5. banana. -3.0. to. 3.0. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. 8 features. learned weights. e1 (1 × 8) × W (8 × 50,000) + b → 50,000 logits. Selected entries shown; the dots stand for thousands of other tokens.
  • Example: d = 8 · vocabulary size |V| = 50,000
  • Updated e1
  • 50,000 output neurons
  • logit
  • Raghav
  • 1.0
  • bicycle
  • -2.0
  • goes
  • 5.5
  • ##ing
  • -1.0
  • school
  • 1.5
  • banana
  • -3.0
  • to
  • 3.0
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • 8 features
  • learned weights
  • e1 (1 × 8) × W (8 × 50,000) + b → 50,000 logits
  • Selected entries shown; the dots stand for thousands of other tokens.

Authored teaching diagram · Primary source

Are the candidate outputs limited to words in the sentence?

No. The output layer scores the full tokenizer vocabulary. This example has 50,000 output neurons. We show seven entries, including bicycle, banana and a subword fragment, with dots for the omitted entries. A learned W maps the eight embedding features to 50,000 logits; b adds one bias per output. Vocabulary size 50,000 and embedding width eight are illustrative, not specifications of a named model. W has shape 8 × 50,000, b has shape 1 × 50,000 and the output has shape 1 × 50,000. The vocabulary is fixed by the tokenizer; ##ing is an illustrative subword notation. The neuron list is a selection, not vocabulary order or a ranking. All scores are authored examples. The seven visible logits are [1, -2, 5.5, -1, 1.5, -3, 3]; each omitted entry has logit -8 for the numerical softmax example. Some models omit the bias. W here is distinct from the attention projection W_O.

Teaching note

Are the candidate outputs limited to words in the sentence? Trace all eight input features to the output neurons. Point to bicycle and banana: these are valid vocabulary entries even though they are unlikely continuations here. The highlighted goes neuron has the largest illustrative logit. Vocabulary size 50,000 and embedding width eight are illustrative, not specifications of a named model. W has shape 8 × 50,000, b has shape 1 × 50,000 and the output has shape 1 × 50,000. The vocabulary is fixed by the tokenizer; ##ing is an illustrative subword notation. The neuron list is a selection, not vocabulary order or a ranking. All scores are authored examples. The seven visible logits are [1, -2, 5.5, -1, 1.5, -3, 3]; each omitted entry has logit -8 for the numerical softmax example. Some models omit the bias. W here is distinct from the attention projection W_O.

Frame overflows: split the content

01 · Generate the next token

From vocabulary scores to the next token

Softmax over all 50,000 vocabulary scores. token. logit. probability. Raghav. 1.0. 0.94%. bicycle. -2.0. 0.05%. goes. 5.5. 84.58%. ##ing. -1.0. 0.13%. school. 1.5. 1.55%. banana. -3.0. 0.02%. to. 3.0. 6.94%. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. Highest: goes. Raghav goes. Append and repeat. Other 49,993 tokens combined:. 5.80%. pj = exp(scorej) / Σ exp(scores) · sum over all 50,000 tokensSoftmax over all 50,000 vocabulary scorestokenlogitprobabilityRaghav1.00.94%bicycle-2.00.05%goes5.584.58%##ing-1.00.13%school1.51.55%banana-3.00.02%to3.06.94%⋮⋮⋮⋮⋮⋮Highest: goesRaghav goesAppend and repeatOther 49,993 tokens combined:5.80%pj = exp(scorej) / Σ exp(scores) · sum over all 50,000 tokens

Greedy decoding selects the highest probability across the full vocabulary.

Read the diagram labels
  • Softmax over all 50,000 vocabulary scores. token. logit. probability. Raghav. 1.0. 0.94%. bicycle. -2.0. 0.05%. goes. 5.5. 84.58%. ##ing. -1.0. 0.13%. school. 1.5. 1.55%. banana. -3.0. 0.02%. to. 3.0. 6.94%. ⋮. ⋮. ⋮. ⋮. ⋮. ⋮. Highest: goes. Raghav goes. Append and repeat. Other 49,993 tokens combined:. 5.80%. pj = exp(scorej) / Σ exp(scores) · sum over all 50,000 tokens
  • Softmax over all 50,000 vocabulary scores
  • token
  • logit
  • probability
  • Raghav
  • 1.0
  • 0.94%
  • bicycle
  • -2.0
  • 0.05%
  • goes
  • 5.5
  • 84.58%
  • ##ing
  • -1.0
  • 0.13%
  • school
  • 1.5
  • 1.55%
  • banana
  • -3.0
  • 0.02%
  • to
  • 3.0
  • 6.94%
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • ⋮
  • Highest: goes
  • Raghav goes
  • Append and repeat
  • Other 49,993 tokens combined:
  • 5.80%
  • pj = exp(scorej) / Σ exp(scores) · sum over all 50,000 tokens

Authored teaching diagram · Primary source

Does softmax consider only the displayed tokens?

No. Its denominator includes all 50,000 scores. The displayed probabilities use the full vocabulary, and the combined probability of the other 49,993 entries appears below the table. Greedy decoding selects goes, appends it to Raghav, and continues with the new prefix. The seven visible logits and 49,993 omitted logits of -8 define this authored distribution at temperature 1. Percentages are rounded to two decimal places; exact probabilities sum to one. The omitted rows are distinct token outputs, not one extra token. Greedy decoding chooses the maximum; sampling can choose another token. Softmax computes probabilities and has no learned parameters; selection is a separate operation. During training, cross-entropy against the known next token updates the vocabulary layer and Transformer. Later generation steps reuse the same projection.

Teaching note

Does softmax consider only the displayed tokens? Compare goes with the unrelated tokens. The omitted entries contribute to normalization even though they are not drawn. Trace the selected token into Raghav goes; its updated e₂ is read at the next step. The seven visible logits and 49,993 omitted logits of -8 define this authored distribution at temperature 1. Percentages are rounded to two decimal places; exact probabilities sum to one. The omitted rows are distinct token outputs, not one extra token. Greedy decoding chooses the maximum; sampling can choose another token. Softmax computes probabilities and has no learned parameters; selection is a separate operation. During training, cross-entropy against the known next token updates the vocabulary layer and Transformer. Later generation steps reuse the same projection.

Frame overflows: split the content

01 · Generate the next token

Generating a sequence one token at a time

Raghav. e1 · input. goes. e2 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. Read e2 → vocabulary head → toRaghave1 · inputgoese2 · inputCausal Transformer blocksmi → Δei → ei + Δei → MLP + residuale1 · Raghavupdated embeddinge2 · goesupdated embeddingRead e2 → vocabulary head → to

Each word has an updated embedding. Read the last one to predict the next word.

Read the diagram labels
  • Raghav. e1 · input. goes. e2 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. Read e2 → vocabulary head → to
  • Raghav
  • e1 · input
  • goes
  • e2 · input
  • Causal Transformer blocks
  • mi → Δei → ei + Δei → MLP + residual
  • e1 · Raghav
  • updated embedding
  • e2 · goes
  • updated embedding
  • Read e2 → vocabulary head → to

Authored teaching diagram · Primary source

What changes after we predict the next token?

The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Teaching note

What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Frame overflows: split the content

01 · Generate the next token

Generating a sequence one token at a time

Raghav. e1 · input. goes. e2 · input. to. e3 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. . e3 · to. updated embedding. Read e3 → vocabulary head → schoolRaghave1 · inputgoese2 · inputtoe3 · inputCausal Transformer blocksmi → Δei → ei + Δei → MLP + residuale1 · Raghavupdated embeddinge2 · goesupdated embeddinge3 · toupdated embeddingRead e3 → vocabulary head → school

Each word has an updated embedding. Read the last one to predict the next word.

Read the diagram labels
  • Raghav. e1 · input. goes. e2 · input. to. e3 · input. . Causal Transformer blocks. mi → Δei → ei + Δei → MLP + residual. . e1 · Raghav. updated embedding. . e2 · goes. updated embedding. . e3 · to. updated embedding. Read e3 → vocabulary head → school
  • Raghav
  • e1 · input
  • goes
  • e2 · input
  • to
  • e3 · input
  • Causal Transformer blocks
  • mi → Δei → ei + Δei → MLP + residual
  • e1 · Raghav
  • updated embedding
  • e2 · goes
  • updated embedding
  • e3 · to
  • updated embedding
  • Read e3 → vocabulary head → school

Authored teaching diagram · Primary source

What changes after we predict the next token?

The chosen token is appended to the prefix with its own input embedding and position. Each displayed output is a final contextualised embedding. Read Raghav’s e₁ to predict “goes”, then the embedding e₂ at “goes” to predict “to”, then e₃ at “to” to predict “school”. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Teaching note

What changes after we predict the next token? Trace each word to its input embedding eᵢ and final contextualised embedding eᵢ. Within a block, attention retrieves mᵢ; the output projection gives Δeᵢ; add it to eᵢ, then apply the MLP and its residual. The amber output is the last supplied position, which feeds the vocabulary head. Input embeddings include positions. The input and updated-embedding labels distinguish the stages without introducing a superscript. The displayed continuation is an example sequence, not a model run. Under causal attention, appending a word does not change earlier positions’ final embeddings for a fixed model in evaluation. The new position reads the available prefix; caching can reuse earlier computations. Multiple head messages are concatenated before the output projection. Normalisation is omitted from the diagram.

Frame overflows: split the content

01 · Generate the next token

Decoder-only models: next-token prediction

Raghav. goes. to. input e1. input e2. input e3. Causal self-attention + residual. MLP + residual. Each token reads. itself + its past. past → next. . e1. . e2. . e3. Vocabulary head → school. × L. blocks. Read the last embedding. Append school. Repeat.Raghavgoestoinput e1input e2input e3Causal self-attention + residualMLP + residualEach token readsitself + its pastpast → nexte1e2e3Vocabulary head → school× LblocksRead the last embedding. Append school. Repeat.

At each step, read the last updated embedding to predict the next token.

Read the diagram labels
  • Raghav. goes. to. input e1. input e2. input e3. Causal self-attention + residual. MLP + residual. Each token reads. itself + its past. past → next. . e1. . e2. . e3. Vocabulary head → school. × L. blocks. Read the last embedding. Append school. Repeat.
  • Raghav
  • goes
  • to
  • input e1
  • input e2
  • input e3
  • Causal self-attention + residual
  • MLP + residual
  • Each token reads
  • itself + its past
  • past → next
  • e1
  • e2
  • e3
  • Vocabulary head → school
  • × L
  • blocks
  • Read the last embedding. Append school. Repeat.

Authored teaching diagram · Primary source

What makes this a decoder-only model, and what is self-attention attending to?

A single causal Transformer stack processes the prompt and generated tokens as one sequence. In each self-attention layer, Q, K and V are projected from the current embeddings of that sequence. The causal mask lets each position read itself and earlier positions. The last updated embedding feeds the vocabulary head. This is the standard text-only decoder-only architecture: no separate source encoder and no encoder-to-decoder cross-attention. Each block has causal multi-head self-attention and an MLP, with residual connections; normalization is omitted here. L counts repeated blocks. Training normally predicts the next token at every eligible position using a shifted target, while generation reads the last position. The same vocabulary head is shared across positions. The matrix shows allowed access, not measured attention. Decoding includes a probability distribution and token-selection step, expanded on the preceding slides.

Teaching note

What makes this a decoder-only model, and what is self-attention attending to? Follow the three supplied tokens into the repeated blocks and their three updated embeddings. Only the last embedding feeds the displayed next-token head. Read the causal square in token order: Raghav, goes, to. Append school and repeat. Colored cells mark allowed attention, not its magnitude; input embedding lookup and normalization are omitted in this summary. This is the standard text-only decoder-only architecture: no separate source encoder and no encoder-to-decoder cross-attention. Each block has causal multi-head self-attention and an MLP, with residual connections; normalization is omitted here. L counts repeated blocks. Training normally predicts the next token at every eligible position using a shifted target, while generation reads the last position. The same vocabulary head is shared across positions. The matrix shows allowed access, not measured attention. Decoding includes a probability distribution and token-selection step, expanded on the preceding slides.

Frame overflows: split the content

01 · Generate the next token

Decoder-only examples: text and code completion

Supplied prefix. Continuation, token by token. Text. Raghav goes to. school …. Code. def square(x): return. x * x. At each step: the prefix and tokens generated so far form one sequence.. . Current sequence. embeddings + positions. . Causal Transformer. self-attention + MLP blocks. . Vocabulary head. select the next token. Append the selected token to the same sequence.. Q, K and V come from this sequence; there is no separate source encoder.Supplied prefixContinuation, token by tokenTextRaghav goes toschool …Codedef square(x): returnx * xAt each step: the prefix and tokens generated so far form one sequence.Current sequenceembeddings + positionsCausal Transformerself-attention + MLP blocksVocabulary headselect the next tokenAppend the selected token to the same sequence.Q, K and V come from this sequence; there is no separate source encoder.

One sequence, one causal Transformer, one next-token prediction at each step.

Read the diagram labels
  • Supplied prefix. Continuation, token by token. Text. Raghav goes to. school …. Code. def square(x): return. x * x. At each step: the prefix and tokens generated so far form one sequence.. . Current sequence. embeddings + positions. . Causal Transformer. self-attention + MLP blocks. . Vocabulary head. select the next token. Append the selected token to the same sequence.. Q, K and V come from this sequence; there is no separate source encoder.
  • Supplied prefix
  • Continuation, token by token
  • Text
  • Raghav goes to
  • school …
  • Code
  • def square(x): return
  • x * x
  • At each step: the prefix and tokens generated so far form one sequence.
  • Current sequence
  • embeddings + positions
  • Causal Transformer
  • self-attention + MLP blocks
  • Vocabulary head
  • select the next token
  • Append the selected token to the same sequence.
  • Q, K and V come from this sequence; there is no separate source encoder.

Authored teaching diagram · Primary source

Does a prompt and an answer imply an encoder-decoder architecture?

No. In a decoder-only model, the prompt and the generated answer occupy successive positions in one sequence. Causal self-attention processes that sequence. A separate source encoder is not present. Text and code completion illustrate this directly. The two rows are separate illustrative completions, not measured model outputs or two inputs processed together. In every self-attention layer, Q, K and V come from the same current sequence. Read the final embedding to produce vocabulary logits, apply softmax and select a token. Decoder-only models can also translate, summarize and chat when suitably trained: the instruction and any source text remain inside the prompt. Those task names do not determine the architecture. A standard Transformer encoder-decoder instead has a separate source encoder whose outputs are available to the decoder through cross-attention. Discuss that distinction after the translation example. Capability depends on training and data; the architecture alone does not guarantee useful completions.

Teaching note

Does a prompt and an answer imply an encoder-decoder architecture? Read each joined strip as one sequence. Purple tokens are supplied; amber tokens are appended one at a time. Only tokens generated so far are inputs at the next step. Follow the shared computation below and the return arrow after token selection. The two rows are separate illustrative completions, not measured model outputs or two inputs processed together. In every self-attention layer, Q, K and V come from the same current sequence. Read the final embedding to produce vocabulary logits, apply softmax and select a token. Decoder-only models can also translate, summarize and chat when suitably trained: the instruction and any source text remain inside the prompt. Those task names do not determine the architecture. A standard Transformer encoder-decoder instead has a separate source encoder whose outputs are available to the decoder through cross-attention. Discuss that distinction after the translation example. Capability depends on training and data; the architecture alone does not guarantee useful completions.

Frame overflows: split the content

02 · Understand the supplied input

Part II · The whole input is available

Predict labels for a complete input sentence.. Complete sentence. Token or sentence labels. Use contextual embeddings to classify the input.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesPredict labels for a complete input sentence.Complete sentenceToken or sentence labelsUse contextual embeddings to classify the input.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Predict labels for a complete input sentence.. Complete sentence. Token or sentence labels. Use contextual embeddings to classify the input.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Predict labels for a complete input sentence.
  • Complete sentence
  • Token or sentence labels
  • Use contextual embeddings to classify the input.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part II · The whole input is available

Predict labels for a complete input sentence.

Teaching note

Part II · The whole input is available Pause and establish the requested output before advancing.

Frame overflows: split the content

02 · Understand the supplied input

Which words refer to people or places?

Raghav. goes. to. school. in. Delhi. PERSON. PLACERaghavgoestoschoolinDelhiPERSONPLACE

Assign entity labels to words in the complete input.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. PERSON. PLACE
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • PERSON
  • PLACE

Authored teaching diagram

Do we need to generate a new sentence?

No. This task assigns a label to each input position. Raghav is a person and Delhi is a place in this teaching example.

Teaching note

Do we need to generate a new sentence? Point to Raghav and Delhi in the same source sentence used on the opening slide.

Frame overflows: split the content

02 · Understand the supplied input

Predicting a label at each input position

Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Do we need to generate a new sentence?RaghavgoestoschoolinDelhiTransformer · whole-input self-attentione1e2e3e4e5e6PERSONOTHEROTHEROTHEROTHERPLACEDo we need to generate a new sentence?

Six input positions → six label decisions.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Do we need to generate a new sentence?
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Transformer · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • PERSON
  • OTHER
  • OTHER
  • OTHER
  • OTHER
  • PLACE
  • Do we need to generate a new sentence?

Authored teaching diagram · Primary source

What should the unhighlighted words receive?

OTHER, under our simplified label menu. Every position receives a label, including words that are not named entities. The four-class inventory is PERSON, ORGANIZATION, PLACE and OTHER. This sentence does not have an organization, but that class remains available at every position.

Teaching note

What should the unhighlighted words receive? Count all six outputs, not just the two entity names. The four-class inventory is PERSON, ORGANIZATION, PLACE and OTHER. This sentence does not have an organization, but that class remains available at every position.

Frame overflows: split the content

02 · Understand the supplied input

Self-attention over the complete input

6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. goes. to. school. in. Delhi6 × 6 · allowed connectionsRaghav goes to school in Delhi.key columnsRRggttssiiDDRaghav can read…RaghavgoestoschoolinDelhi

Filled cell = connection allowed, not an attention weight.

Read the diagram labels
  • 6 × 6 · allowed connections. Raghav goes to school in Delhi.. key columns. R. R. g. g. t. t. s. s. i. i. D. D. Raghav can read…. Raghav. goes. to. school. in. Delhi
  • 6 × 6 · allowed connections
  • Raghav goes to school in Delhi.
  • key columns
  • R
  • R
  • g
  • g
  • t
  • t
  • s
  • s
  • i
  • i
  • D
  • D
  • Raghav can read…
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi

Authored teaching diagram · Primary source

Why is it legitimate for Raghav’s state to use Delhi?

We supplied the whole input and ask for labels about it. Later source words are available information, not future target answers.

Teaching note

Why is it legitimate for Raghav’s state to use Delhi? Compare the full matrix with the familiar causal triangle. Do not introduce a new attention equation.

Frame overflows: split the content

02 · Understand the supplied input

Queries, keys and values in an encoder

Raghav’s embedding e1. q1 = e1 × WQ. Current embeddings stacked in E. K = E × WK. V = E × WV. m1 = Attention(q1, K, V) = softmax(q1KT / √dk) V. The projections are shared across positions within a head.Raghav’s embedding e1q1 = e1 × WQCurrent embeddings stacked in EK = E × WKV = E × WVm1 = Attention(q1, K, V) = softmax(q1KT / √dk) VThe projections are shared across positions within a head.

The attention calculation is unchanged; each position can now use the complete input.

Read the diagram labels
  • Raghav’s embedding e1. q1 = e1 × WQ. Current embeddings stacked in E. K = E × WK. V = E × WV. m1 = Attention(q1, K, V) = softmax(q1KT / √dk) V. The projections are shared across positions within a head.
  • Raghav’s embedding e1
  • q1 = e1 × WQ
  • Current embeddings stacked in E
  • K = E × WK
  • V = E × WV
  • m1 = Attention(q1, K, V) = softmax(q1KT / √dk) V
  • The projections are shared across positions within a head.

Authored teaching diagram · Primary source

Do we need a new attention operation to read the whole input?

No. The same query–key scores, softmax and weighted values apply. The mask now permits all supplied positions. We use row-vector notation: q₁ = e₁ × W_Q, K = E × W_K, V = E × W_V. e₁ is Raghav’s current embedding; E stacks all current input embeddings for this layer. The retrieved result m₁ is its attention message. Q, K and V in the projection subscripts identify the learned maps. Scaling uses the query/key width dₖ, not necessarily the model width.

Teaching note

Do we need a new attention operation to read the whole input? Trace the Raghav query and the source-wide keys and values. Within one head/layer, positions share the projection matrices. We use row-vector notation: q₁ = e₁ × W_Q, K = E × W_K, V = E × W_V. e₁ is Raghav’s current embedding; E stacks all current input embeddings for this layer. The retrieved result m₁ is its attention message. Q, K and V in the projection subscripts identify the learned maps. Scaling uses the query/key width dₖ, not necessarily the model width.

Frame overflows: split the content

02 · Understand the supplied input

One contextual vector per token

Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Raghav in context. Delhi in context. E ∈ ℝ6 × d. 6 token positions × d featuresRaghavgoestoschoolinDelhiTransformer · whole-input self-attentione1e2e3e4e5e6Raghav in contextDelhi in contextE ∈ ℝ6 × d6 token positions × d features

E has shape n × d.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Raghav in context. Delhi in context. E ∈ ℝ6 × d. 6 token positions × d features
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Transformer · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Raghav in context
  • Delhi in context
  • E ∈ ℝ6 × d
  • 6 token positions × d features

Authored teaching diagram · Primary source

Is e₁ just the original embedding for Raghav?

No. It is the final contextual representation at that position after the Transformer blocks, and can depend on the supplied sentence.

Teaching note

Is e₁ just the original embedding for Raghav? Distinguish the input token identity from its final vector. The drawn bars indicate multiple features, not semantic coordinates.

Frame overflows: split the content

02 · Understand the supplied input

A shared classifier for all token positions

Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. ONE shared classifier: zi = ei W + b → softmax over 4 labels. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Entity labels at every position: PERSON · ORGANIZATION · PLACE · OTHERRaghavgoestoschoolinDelhiTransformer · whole-input self-attentione1e2e3e4e5e6ONE shared classifier: zi = ei W + b → softmax over 4 labelsPERSONOTHEROTHEROTHEROTHERPLACEEntity labels at every position: PERSON · ORGANIZATION · PLACE · OTHER

One W and b, reused at all six positions.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. Transformer · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. ONE shared classifier: zi = ei W + b → softmax over 4 labels. PERSON. OTHER. OTHER. OTHER. OTHER. PLACE. Entity labels at every position: PERSON · ORGANIZATION · PLACE · OTHER
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Transformer · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • ONE shared classifier: zi = ei W + b → softmax over 4 labels
  • PERSON
  • OTHER
  • OTHER
  • OTHER
  • OTHER
  • PLACE
  • Entity labels at every position: PERSON · ORGANIZATION · PLACE · OTHER

Authored teaching diagram · Primary source

Are we training six unrelated classifiers?

No. A single shared classifier maps each d-dimensional state to the same four label scores. Softmax operates across the four classes at each position.

Teaching note

Are we training six unrelated classifiers? Trace both Raghav and Delhi through the same wide classifier band. Ask which parameters are shared.

Frame overflows: split the content

02 · Understand the supplied input

Four entity labels for each token

. ei. same W, same b. PERSON. ORGANIZATION. PLACE. OTHER. d features → 4 scores → 4 probabilitieseisame W, same bPERSONORGANIZATIONPLACEOTHERd features → 4 scores → 4 probabilities

Training target: one entity label per position.

Read the diagram labels
  • . ei. same W, same b. PERSON. ORGANIZATION. PLACE. OTHER. d features → 4 scores → 4 probabilities
  • ei
  • same W, same b
  • PERSON
  • ORGANIZATION
  • PLACE
  • OTHER
  • d features → 4 scores → 4 probabilities

Authored teaching diagram · Primary source

Does the classifier for Delhi still score PERSON?

Yes. Every position receives all four class scores. A gold label per token supervises the corresponding probability distribution. A typical supervised loss sums or averages token cross-entropies. The task head must be trained; merely attaching a new head does not teach labels.

Teaching note

Does the classifier for Delhi still score PERSON? Point at all four scores, then say that the same computation repeats for every input position. A typical supervised loss sums or averages token cross-entropies. The task head must be trained; merely attaching a new head does not teach labels.

Frame overflows: split the content

02 · Understand the supplied input

Encoder: a contextual representation for each token

Raghav. goes. to. school. in. Delhi. input e1. input e2. input e3. input e4. input e5. input e6. Whole-input self-attention + residual. MLP + residual. Every token can read. the complete input. all ↔ all. × L. blocks. . e1. . e2. . e3. . e4. . e5. . e6. Six input tokens → six contextual embeddings.RaghavgoestoschoolinDelhiinput e1input e2input e3input e4input e5input e6Whole-input self-attention + residualMLP + residualEvery token can readthe complete inputall ↔ all× Lblockse1e2e3e4e5e6Six input tokens → six contextual embeddings.

Use the embeddings for token labels, sentence readouts or similarity.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. input e1. input e2. input e3. input e4. input e5. input e6. Whole-input self-attention + residual. MLP + residual. Every token can read. the complete input. all ↔ all. × L. blocks. . e1. . e2. . e3. . e4. . e5. . e6. Six input tokens → six contextual embeddings.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • input e1
  • input e2
  • input e3
  • input e4
  • input e5
  • input e6
  • Whole-input self-attention + residual
  • MLP + residual
  • Every token can read
  • the complete input
  • all ↔ all
  • × L
  • blocks
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Six input tokens → six contextual embeddings.

Authored teaching diagram · Primary source

Must an encoder reduce the input to a single vector?

No. Its output is a sequence of contextual states. Pooling, a readout token or a token classifier is a separate downstream choice. BERT is an example of this encoder family. Its pretraining objectives are not needed to understand the arrangement here.

Teaching note

Must an encoder reduce the input to a single vector? Trace each of the six input positions to its updated output embedding. The full attention square means every supplied position can read every other position. The six separate output arrows retain one embedding per input position; this does not pool the sentence into a single vector. BERT is an example of this encoder family. Its pretraining objectives are not needed to understand the arrangement here.

Frame overflows: split the content

02 · Understand the supplied input

Predicting one label for a sentence

The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. Which state should we read?. One review → one sentiment labelThemoviewasreallyawfulEncodere1e2e3e4e5Which state should we read?One review → one sentiment label

The movie was really awful. → NEGATIVE

Read the diagram labels
  • The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. Which state should we read?. One review → one sentiment label
  • The
  • movie
  • was
  • really
  • awful
  • Encoder
  • e1
  • e2
  • e3
  • e4
  • e5
  • Which state should we read?
  • One review → one sentiment label

Authored teaching diagram

We have five contextual states. How do we get one class prediction?

Choose a whole-sequence readout, such as pooling the states or selecting a learned readout position, then apply a classifier.

Teaching note

We have five contextual states. How do we get one class prediction? Pause with all five states visible before suggesting a pooling operation.

Frame overflows: split the content

02 · Understand the supplied input

Sentence classification with mean pooling

The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. mean: c = (e1 + … + e5) / 5. c → classifier → NEGATIVEThemoviewasreallyawfulEncodere1e2e3e4e5mean: c = (e1 + … + e5) / 5c → classifier → NEGATIVE

The pooled vector is the input to the sentence classifier.

Read the diagram labels
  • The. movie. was. really. awful. Encoder. . e1. . e2. . e3. . e4. . e5. mean: c = (e1 + … + e5) / 5. c → classifier → NEGATIVE
  • The
  • movie
  • was
  • really
  • awful
  • Encoder
  • e1
  • e2
  • e3
  • e4
  • e5
  • mean: c = (e1 + … + e5) / 5
  • c → classifier → NEGATIVE

Authored teaching diagram · Primary source

What shape does mean pooling produce?

Averaging n vectors of width d produces one vector of width d. A classifier maps that vector to C class scores. Mean pooling is a possible design, not a guarantee of useful features. It must be used with suitable training. This example has no padding.

Teaching note

What shape does mean pooling produce? Follow the five state arrows into the mean and then into the classifier. Remember this pooling drawing for translation. Mean pooling is a possible design, not a guarantee of useful features. It must be used with suitable training. This example has no padding.

Frame overflows: split the content

02 · Understand the supplied input

Sentence classification with a CLS token

[CLS]. The. movie. was. really. awful. Encoder · every supplied position is available. . eCLS. . e1. . e2. . e3. . e4. . e5. Classifier → NEGATIVE[CLS]ThemoviewasreallyawfulEncoder · every supplied position is availableeCLSe1e2e3e4e5Classifier → NEGATIVE

[CLS] participates in ordinary whole-input attention.

Read the diagram labels
  • [CLS]. The. movie. was. really. awful. Encoder · every supplied position is available. . eCLS. . e1. . e2. . e3. . e4. . e5. Classifier → NEGATIVE
  • [CLS]
  • The
  • movie
  • was
  • really
  • awful
  • Encoder · every supplied position is available
  • eCLS
  • e1
  • e2
  • e3
  • e4
  • e5
  • Classifier → NEGATIVE

Authored teaching diagram · Primary source

How can a token at the beginning collect information from later words?

The whole-input mask allows its query to read every supplied position, including the later words. We select its final contextual state for the task head. CLS also has keys and values and uses the ordinary per-head/layer projection matrices. The encoder now returns six states for five review words plus CLS; the classifier reads one.

Teaching note

How can a token at the beginning collect information from later words? Trace the added input position through the same stack. Follow only its final state to the sentiment head. CLS also has keys and values and uses the ordinary per-head/layer projection matrices. The encoder now returns six states for five review words plus CLS; the classifier reads one.

Frame overflows: split the content

02 · Understand the supplied input

Where does the CLS embedding come from?

Embedding table · trainable. . token ID. d numbers. . [CLS]. . Raghav. . goes. [CLS] has its own row, like a word.. How does that row start?. From scratch: small random values, once.. Pretrained: load the saved row.. Look up eCLS, then add position. Encoder + sentence → updated eCLS. Same table row for every sentence; the encoder output depends on the sentence.Embedding table · trainabletoken IDd numbers[CLS]Raghavgoes[CLS] has its own row, like a word.How does that row start?From scratch: small random values, once.Pretrained: load the saved row.Look up eCLS, then add positionEncoder + sentence → updated eCLSSame table row for every sentence; the encoder output depends on the sentence.

[CLS] has a token ID and a trainable row in the embedding table.

Read the diagram labels
  • Embedding table · trainable. . token ID. d numbers. . [CLS]. . Raghav. . goes. [CLS] has its own row, like a word.. How does that row start?. From scratch: small random values, once.. Pretrained: load the saved row.. Look up eCLS, then add position. Encoder + sentence → updated eCLS. Same table row for every sentence; the encoder output depends on the sentence.
  • Embedding table · trainable
  • token ID
  • d numbers
  • [CLS]
  • Raghav
  • goes
  • [CLS] has its own row, like a word.
  • How does that row start?
  • From scratch: small random values, once.
  • Pretrained: load the saved row.
  • Look up eCLS, then add position
  • Encoder + sentence → updated eCLS
  • Same table row for every sentence; the encoder output depends on the sentence.

Authored teaching diagram · Primary source

Do we initialize CLS with a new random vector for every sentence?

No. In this text encoder, [CLS] is a special token with its own ID. That ID selects a trainable row of d numbers from the same embedding table used for ordinary tokens. From scratch, initialise that row with small random values once; a pretrained model supplies its saved row. Add positional information, then send the sequence through the encoder. The table is a model parameter, not a fresh random draw for each example. Backpropagation and optimizer steps learn its CLS row. This is the standard special-token lookup design for text; a separate trainable vector is another implementation choice. The coloured cells are schematic vector coordinates, not measured values.

Teaching note

Do we initialize CLS with a new random vector for every sentence? Point to the ordinary token rows, then the CLS row. Follow the lookup arrow into the encoder input. The initial row is shared across sentences at any fixed set of model parameters; the updated CLS embedding depends on the supplied sentence. The table is a model parameter, not a fresh random draw for each example. Backpropagation and optimizer steps learn its CLS row. This is the standard special-token lookup design for text; a separate trainable vector is another implementation choice. The coloured cells are schematic vector coordinates, not measured values.

Frame overflows: split the content

02 · Understand the supplied input

CLS learns through the prediction loss

Forward pass: the CLS embedding table row and review-word embeddings enter the encoder. Select the updated CLS embedding, pass it into a schematic neural classifier and softmax, and obtain predicted probabilities y hat. The supplied positive label y and probabilities define cross-entropy loss minus log 0.8, approximately 0.223. Dashed backward arrows show gradients through softmax, classifier, encoder and the input table lookup. The optimizer updates trainable parameters. The displayed probabilities are an illustrative example, not model measurements.FORWARD PASS → compute the prediction and lossinput[CLS] roweCLSEncodera good moviereview wordsupdatedeCLSNeural classifier2 class scoressoftmaxpredictionŷ[0.2, 0.8]NEG, POSL(ŷ, y)given labely = POSITIVE−log(0.8)≈ 0.223← BACKWARD PASS: loss gradients flow through the same operationsOptimizer updates: CLS table row + encoder + classifier parameters.

Compare prediction ŷ with label y; use the loss to train the CLS representation.

Read the diagram labels
  • Forward pass: the CLS embedding table row and review-word embeddings enter the encoder. Select the updated CLS embedding, pass it into a schematic neural classifier and softmax, and obtain predicted probabilities y hat. The supplied positive label y and probabilities define cross-entropy loss minus log 0.8, approximately 0.223. Dashed backward arrows show gradients through softmax, classifier, encoder and the input table lookup. The optimizer updates trainable parameters. The displayed probabilities are an illustrative example, not model measurements.
  • FORWARD PASS → compute the prediction and loss
  • input
  • [CLS] row
  • eCLS
  • Encoder
  • a good movie
  • review words
  • updated
  • eCLS
  • Neural classifier
  • 2 class scores
  • softmax
  • prediction
  • ŷ
  • [0.2, 0.8]
  • NEG, POS
  • L(ŷ, y)
  • given label
  • y = POSITIVE
  • −log(0.8)
  • ≈ 0.223
  • ← BACKWARD PASS: loss gradients flow through the same operations
  • Optimizer updates: CLS table row + encoder + classifier parameters.

Authored teaching diagram · Primary source

Do we provide a target embedding for CLS, or a true label y?

We supply the true class label y. The CLS table row and review words pass through the encoder; the updated CLS embedding enters a neural classifier. Softmax gives the predicted class probabilities ŷ. The loss compares ŷ with y. Backpropagation computes gradients through the classifier and encoder to the original CLS embedding-table row. Here ŷ denotes the full predicted probability vector, while y is the supplied class label, not another prediction. The loss uses the probability assigned to the true class; no argmax is inserted in the differentiable training path. The classifier drawing is a schematic MLP with one hidden layer; a linear classifier is also possible. Neuron counts do not specify the encoder width. Input embeddings receive positional information; all encoder positions participate even though only the CLS output is drawn. The dashed lane is an offset view of gradients through the same computational graph, not a direct shortcut from loss to CLS. This shows full fine-tuning of unfrozen parameters. The optimizer changes model parameters, not the stored value of a per-example updated CLS activation. Frozen parameters would not update. Natural logarithms are used, and the probabilities are illustrative.

Teaching note

Do we provide a target embedding for CLS, or a true label y? Follow the solid forward arrows left to right. For the positive review, the illustrative probabilities are NEG 0.2 and POS 0.8; cross-entropy is −log(0.8), about 0.223. Then follow the dashed lower lane from right to left. Finally distinguish gradient computation from the optimizer step. Here ŷ denotes the full predicted probability vector, while y is the supplied class label, not another prediction. The loss uses the probability assigned to the true class; no argmax is inserted in the differentiable training path. The classifier drawing is a schematic MLP with one hidden layer; a linear classifier is also possible. Neuron counts do not specify the encoder width. Input embeddings receive positional information; all encoder positions participate even though only the CLS output is drawn. The dashed lane is an offset view of gradients through the same computational graph, not a direct shortcut from loss to CLS. This shows full fine-tuning of unfrozen parameters. The optimizer changes model parameters, not the stored value of a per-example updated CLS activation. Frozen parameters would not update. Natural logarithms are used, and the probabilities are illustrative.

Frame overflows: split the content

02 · Understand the supplied input

Token classification and sentence classification

A LABEL PER TOKEN. Raghav goes to school in Delhi. Encoder. . e1. PERSON. . e6. PLACE. …. shared head. ONE SENTENCE LABEL. [CLS] Raghav goes to school in Delhi. Encoder. . eCLS. one classifier → one labelA LABEL PER TOKENRaghav goes to school in DelhiEncodere1PERSONe6PLACE…shared headONE SENTENCE LABEL[CLS] Raghav goes to school in DelhiEncodereCLSone classifier → one label

Read every state for NER; read CLS for one sentence label.

Read the diagram labels
  • A LABEL PER TOKEN. Raghav goes to school in Delhi. Encoder. . e1. PERSON. . e6. PLACE. …. shared head. ONE SENTENCE LABEL. [CLS] Raghav goes to school in Delhi. Encoder. . eCLS. one classifier → one label
  • A LABEL PER TOKEN
  • Raghav goes to school in Delhi
  • Encoder
  • e1
  • PERSON
  • e6
  • PLACE
  • …
  • shared head
  • ONE SENTENCE LABEL
  • [CLS] Raghav goes to school in Delhi
  • Encoder
  • eCLS
  • one classifier → one label

Authored teaching diagram · Primary source

What changes when the whole sentence needs one label?

The task determines the readout and head. NER reads all ordinary token states through shared parameters. Sentence classification selects the final CLS state. For example, one could define a topic classification dataset; this drawing does not invent a measured topic prediction.

Teaching note

What changes when the whole sentence needs one label? Compare the identical Raghav sentence on both sides. The sentence-classification label is deliberately unspecified until a dataset defines the task. For example, one could define a topic classification dataset; this drawing does not invent a measured topic prediction.

Frame overflows: split the content

02 · Understand the supplied input

Six words, eight features per word

Raghav goes to school in Delhi. Encoder · choose d = 8 for this example. 1. 2. 3. 4. 5. 6. 7. 8. Raghav. goes. to. school. in. Delhi. 6 rows = 6 token positions. 8 columns = 8 features. E has shape 6 × 8. Each row is one updated embedding. The cells stand for numbers.Raghav goes to school in DelhiEncoder · choose d = 8 for this example12345678RaghavgoestoschoolinDelhi6 rows = 6 token positions8 columns = 8 featuresE has shape 6 × 8Each row is one updated embedding. The cells stand for numbers.

6 counts token positions; 8 counts features inside each embedding.

Read the diagram labels
  • Raghav goes to school in Delhi. Encoder · choose d = 8 for this example. 1. 2. 3. 4. 5. 6. 7. 8. Raghav. goes. to. school. in. Delhi. 6 rows = 6 token positions. 8 columns = 8 features. E has shape 6 × 8. Each row is one updated embedding. The cells stand for numbers.
  • Raghav goes to school in Delhi
  • Encoder · choose d = 8 for this example
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • 6 rows = 6 token positions
  • 8 columns = 8 features
  • E has shape 6 × 8
  • Each row is one updated embedding. The cells stand for numbers.

Authored teaching diagram · Primary source

What does a single row of this matrix represent?

One word’s updated embedding. We choose a toy model width of eight, so six input words produce six rows of eight numbers: E has shape 6 × 8. A feature column is not a word or a class label. This is one sentence with no batch dimension. The width eight is chosen only to make the shape easy to draw. It does not come from the six-word sentence or the four entity classes.

Teaching note

What does a single row of this matrix represent? Read the Raghav row across all eight cells, then count all six rows. The cells represent vector coordinates, not actual values. This is one sentence with no batch dimension. The width eight is chosen only to make the shape easy to draw. It does not come from the six-word sentence or the four entity classes.

Frame overflows: split the content

02 · Understand the supplied input

NER: follow each word from input to scores

Six words are aligned with six input embeddings. Each embedding has eight features in this toy example. Whole-input self-attention and MLP blocks with residual updates return six updated eight-feature embeddings. One shared classifier multiplies every updated embedding by the same 8 by 4 weight matrix and adds the same four-component bias. It returns four scores per word, in the order PERSON, ORGANIZATION, PLACE, OTHER. The twenty-four output boxes are score slots, not measured values. Input embeddings include position information.Raghavinput e1goesinput e2toinput e3schoolinput e4ininput e5Delhiinput e6Whole-input self-attention → MLPresidual updates · repeated encoder blocksupdated e1updated e2updated e3updated e4updated e5updated e6Same classifier at every position: × W (8 × 4), then + b (4)Each group of four boxes holds one word’s PERSON, ORG, PLACE and OTHER scores.

6 × 8 input features → 6 × 8 updated features → 6 × 4 class scores.

Read the diagram labels
  • Six words are aligned with six input embeddings. Each embedding has eight features in this toy example. Whole-input self-attention and MLP blocks with residual updates return six updated eight-feature embeddings. One shared classifier multiplies every updated embedding by the same 8 by 4 weight matrix and adds the same four-component bias. It returns four scores per word, in the order PERSON, ORGANIZATION, PLACE, OTHER. The twenty-four output boxes are score slots, not measured values. Input embeddings include position information.
  • Raghav
  • input e1
  • goes
  • input e2
  • to
  • input e3
  • school
  • input e4
  • in
  • input e5
  • Delhi
  • input e6
  • Whole-input self-attention → MLP
  • residual updates · repeated encoder blocks
  • updated e1
  • updated e2
  • updated e3
  • updated e4
  • updated e5
  • updated e6
  • Same classifier at every position: × W (8 × 4), then + b (4)
  • Each group of four boxes holds one word’s PERSON, ORG, PLACE and OTHER scores.

Authored teaching diagram · Primary source

What happens before we multiply by the classifier weights?

Each word has an input embedding, with positional information. Whole-input self-attention and MLP blocks update those embeddings using the sentence. The encoder still returns one eight-feature vector per word. The shared classifier then multiplies each updated vector by W and adds b, producing four class scores. We retain eᵢ at position i and label its input and updated stages explicitly. The model width eight is illustrative. W has shape 8 × 4; b has four components and is added, not multiplied. In matrix form, scores = E W + b, with b broadcast across the six rows. The four output slots per word use PERSON, ORGANIZATION, PLACE, OTHER order; they are not numerical predictions. The next slide rearranges those six sets of scores into rows and applies softmax independently within each row. Positional information and normalization are omitted from the drawing.

Teaching note

What happens before we multiply by the classifier weights? Follow Raghav straight down, then Delhi. At every stage there are six token positions. The encoder changes their features while preserving width eight. The task head changes the feature width from eight to four class scores; its parameters are shared across all six positions. We retain eᵢ at position i and label its input and updated stages explicitly. The model width eight is illustrative. W has shape 8 × 4; b has four components and is added, not multiplied. In matrix form, scores = E W + b, with b broadcast across the six rows. The four output slots per word use PERSON, ORGANIZATION, PLACE, OTHER order; they are not numerical predictions. The next slide rearranges those six sets of scores into rows and applies softmax independently within each row. Positional information and normalization are omitted from the drawing.

Frame overflows: split the content

02 · Understand the supplied input

NER: turn each set of scores into a label

One shared classifier reads each of the six embeddings. E: 6 × 8. W: 8 × 4; b: 4. Scores: 6 × 4. PERSON. ORG. PLACE. OTHER. Raghav. PERSON. goes. OTHER. to. OTHER. school. OTHER. in. OTHER. Delhi. PLACE. chosen label. Softmax within each row → 4 probabilities → one label per word.One shared classifier reads each of the six embeddingsE: 6 × 8W: 8 × 4; b: 4Scores: 6 × 4PERSONORGPLACEOTHERRaghavPERSONgoesOTHERtoOTHERschoolOTHERinOTHERDelhiPLACEchosen labelSoftmax within each row → 4 probabilities → one label per word.

The classifier changes 8 features into 4 scores at every position.

Read the diagram labels
  • One shared classifier reads each of the six embeddings. E: 6 × 8. W: 8 × 4; b: 4. Scores: 6 × 4. PERSON. ORG. PLACE. OTHER. Raghav. PERSON. goes. OTHER. to. OTHER. school. OTHER. in. OTHER. Delhi. PLACE. chosen label. Softmax within each row → 4 probabilities → one label per word.
  • One shared classifier reads each of the six embeddings
  • E: 6 × 8
  • W: 8 × 4; b: 4
  • Scores: 6 × 4
  • PERSON
  • ORG
  • PLACE
  • OTHER
  • Raghav
  • PERSON
  • goes
  • OTHER
  • to
  • OTHER
  • school
  • OTHER
  • in
  • OTHER
  • Delhi
  • PLACE
  • chosen label
  • Softmax within each row → 4 probabilities → one label per word.

Authored teaching diagram · Primary source

Why is the result 6 × 4 rather than one vector of four scores?

We classify every input position. The same W of shape 8 × 4 and bias of length four are used for all six rows: scores = E W + b has shape 6 × 4. Softmax across the four columns gives one distribution per word. Dark cells indicate the illustrative highest-scoring class; their shading is not a numerical score or probability. Raghav is labelled PERSON, Delhi PLACE, and the other words OTHER in this teaching example. The table depicts score slots and chosen labels, not a trained model run.

Teaching note

Why is the result 6 × 4 rather than one vector of four scores? The previous slide showed four score slots beneath every word. Here each word is a row and each class is a column. Point across the Raghav row, then the Delhi row. Both use the same classifier parameters. ORG abbreviates ORGANIZATION. Dark cells indicate the illustrative highest-scoring class; their shading is not a numerical score or probability. Raghav is labelled PERSON, Delhi PLACE, and the other words OTHER in this teaching example. The table depicts score slots and chosen labels, not a trained model run.

Frame overflows: split the content

02 · Understand the supplied input

CLS: select one row from seven

Add [CLS] to the same six-word sentence. 7 input positions → encoder → 7 updated embeddings. 1. 2. 3. 4. 5. 6. 7. 8. [CLS]. Raghav. goes. to. school. in. Delhi. 7 rows × 8 features = 7 × 8. Select the updated eCLS row. 1 row × 8 features = 1 × 8. Selection keeps all eight features.Add [CLS] to the same six-word sentence7 input positions → encoder → 7 updated embeddings12345678[CLS]RaghavgoestoschoolinDelhi7 rows × 8 features = 7 × 8Select the updated eCLS row1 row × 8 features = 1 × 8Selection keeps all eight features.

7 × 8 becomes 1 × 8 when we select the updated CLS embedding.

Read the diagram labels
  • Add [CLS] to the same six-word sentence. 7 input positions → encoder → 7 updated embeddings. 1. 2. 3. 4. 5. 6. 7. 8. [CLS]. Raghav. goes. to. school. in. Delhi. 7 rows × 8 features = 7 × 8. Select the updated eCLS row. 1 row × 8 features = 1 × 8. Selection keeps all eight features.
  • Add [CLS] to the same six-word sentence
  • 7 input positions → encoder → 7 updated embeddings
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • [CLS]
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • 7 rows × 8 features = 7 × 8
  • Select the updated eCLS row
  • 1 row × 8 features = 1 × 8
  • Selection keeps all eight features.

Authored teaching diagram · Primary source

Does selecting CLS shrink every embedding to one feature?

No. It selects one position and keeps all eight features in that row. Six ordinary words plus CLS produce seven updated embeddings: 7 × 8. Reading just the CLS row gives 1 × 8. CLS can contain information from the other words because its embedding has already passed through whole-input attention. This is row selection, not averaging the rows or compressing the feature dimension.

Teaching note

Does selecting CLS shrink every embedding to one feature? Hold the eight columns fixed while highlighting the CLS row. The other rows still exist; this task head simply does not read them directly. CLS can contain information from the other words because its embedding has already passed through whole-input attention. This is row selection, not averaging the rows or compressing the feature dimension.

Frame overflows: split the content

02 · Understand the supplied input

A neural classifier reads the updated CLS embedding

CLS token identity can be represented by a vocabulary-length one-hot vector. Embedding lookup selects its learned dense input embedding; the encoder combines it with the sentence and returns an updated dense CLS embedding. That eight-feature vector enters a six-unit hidden layer with a nonlinear activation, then three output neurons. Illustrative logits 2, 1, 0 become probabilities 0.665, 0.245, 0.090 for EDUCATION, TRAVEL, OTHER after softmax. The one-hot identity is not the classifier input. Vocabulary size, embedding width, hidden width and number of classes are distinct.[CLS] token identity0010…0one-hot · |V| entriesEmbedding lookupinput eCLSEncoder + sentenceeCLSupdatedUpdated eCLSHidden layer3 logitsProbabilitiessoftmax2.00.665EDUCATION1.00.245TRAVEL0.00.090OTHER8 dense features → 6 hidden units → 3 logits → 3 probabilitiesExample scores; the hidden layer has a nonlinear activation.

The one-hot vector identifies CLS; its updated embedding is used for classification.

Read the diagram labels
  • CLS token identity can be represented by a vocabulary-length one-hot vector. Embedding lookup selects its learned dense input embedding; the encoder combines it with the sentence and returns an updated dense CLS embedding. That eight-feature vector enters a six-unit hidden layer with a nonlinear activation, then three output neurons. Illustrative logits 2, 1, 0 become probabilities 0.665, 0.245, 0.090 for EDUCATION, TRAVEL, OTHER after softmax. The one-hot identity is not the classifier input. Vocabulary size, embedding width, hidden width and number of classes are distinct.
  • [CLS] token identity
  • 0
  • 0
  • 1
  • 0
  • …
  • 0
  • one-hot · |V| entries
  • Embedding lookup
  • input eCLS
  • Encoder + sentence
  • eCLS
  • updated
  • Updated eCLS
  • Hidden layer
  • 3 logits
  • Probabilities
  • softmax
  • 2.0
  • 0.665
  • EDUCATION
  • 1.0
  • 0.245
  • TRAVEL
  • 0.0
  • 0.090
  • OTHER
  • 8 dense features → 6 hidden units → 3 logits → 3 probabilities
  • Example scores; the hidden layer has a nonlinear activation.

Authored teaching diagram · Primary source

Does the classifier receive a one-hot CLS vector?

No. A vocabulary-length one-hot vector represents the identity of [CLS] for embedding lookup. That lookup returns its dense input embedding. After the encoder processes CLS with the sentence, the updated CLS vector enters the classifier. Our example has eight input features, six hidden units with a nonlinear activation, and three output logits. The one-hot vector has |V| entries, one for each vocabulary token; it is not a three-class target vector. Implementations normally index the embedding table directly rather than materialize one-hot vectors. Positions are included before the encoder. The schematic classifier is an MLP: h = activation(e_CLS W₁ + b₁), with W₁ of shape 8 × 6 and b₁ of length 6; logits = h W₂ + b₂, with W₂ of shape 6 × 3 and b₂ of length 3. This is a valid alternative to a linear 8 × 3 head. The logit example [2, 1, 0] gives softmax probabilities approximately [0.665, 0.245, 0.090], summing to one before rounding. These are authored numbers, not model measurements. A labelled dataset and an annotation rule define the topic classes. The model width, hidden width, vocabulary size and class count are independent choices.

Teaching note

Does the classifier receive a one-hot CLS vector? Trace the top strip from token identity through lookup and encoding, then follow the arrow to the eight input neurons. Count six hidden units and three output neurons. Read each logit and its probability beside EDUCATION, TRAVEL and OTHER. The softmax operates across all three logits. The one-hot vector has |V| entries, one for each vocabulary token; it is not a three-class target vector. Implementations normally index the embedding table directly rather than materialize one-hot vectors. Positions are included before the encoder. The schematic classifier is an MLP: h = activation(e_CLS W₁ + b₁), with W₁ of shape 8 × 6 and b₁ of length 6; logits = h W₂ + b₂, with W₂ of shape 6 × 3 and b₂ of length 3. This is a valid alternative to a linear 8 × 3 head. The logit example [2, 1, 0] gives softmax probabilities approximately [0.665, 0.245, 0.090], summing to one before rounding. These are authored numbers, not model measurements. A labelled dataset and an annotation rule define the topic classes. The model width, hidden width, vocabulary size and class count are independent choices.

Frame overflows: split the content

02 · Understand the supplied input

One encoder, several ways to use its embeddings

Encoder. Select updated eCLS. one label. Encoder. Use updated e1 … en. one label / token. Encoder. . Pool token embeddings. or select updated eCLS. representation. Readout = choosing or combining the encoder’s updated embeddings.EncoderSelect updated eCLSone labelEncoderUse updated e1 … enone label / tokenEncoderPool token embeddingsor select updated eCLSrepresentationReadout = choosing or combining the encoder’s updated embeddings.

Selecting updated CLS and pooling token embeddings are two readout choices.

Read the diagram labels
  • Encoder. Select updated eCLS. one label. Encoder. Use updated e1 … en. one label / token. Encoder. . Pool token embeddings. or select updated eCLS. representation. Readout = choosing or combining the encoder’s updated embeddings.
  • Encoder
  • Select updated eCLS
  • one label
  • Encoder
  • Use updated e1 … en
  • one label / token
  • Encoder
  • Pool token embeddings
  • or select updated eCLS
  • representation
  • Readout = choosing or combining the encoder’s updated embeddings.

Authored teaching diagram · Primary source

Does readout mean another token or another operation after CLS?

Readout is the general choice of which encoder outputs to use and how to combine them. Selecting the updated e_CLS is one readout. Pooling token embeddings is another. Token classification uses all ordinary token embeddings. There is no extra readout token or mandatory processing step implied by the word here. The diagram abbreviates the task heads in the first two rows. A sentence or token classifier is still needed to produce labels. Retrieval can use a suitably trained representation, potentially with a learned projection and normalization; neither raw pooling nor raw CLS is automatically useful for similarity. The training objective and evaluation must match the use.

Teaching note

Does readout mean another token or another operation after CLS? Read each concrete choice aloud. The top row selects updated CLS for a sentence classifier. The middle row sends all token embeddings through a shared token classifier. The bottom row forms a whole-input representation by pooling or by selecting updated CLS. The diagram abbreviates the task heads in the first two rows. A sentence or token classifier is still needed to produce labels. Retrieval can use a suitably trained representation, potentially with a learned projection and normalization; neither raw pooling nor raw CLS is automatically useful for similarity. The training objective and evaluation must match the use.

Frame overflows: split the content

03 · Represent one sequence, generate another

Part III · Sequence-to-sequence prediction

Encode the English input and generate a Hindi translation.. English source. Hindi target. Represent the source and condition target generation.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesEncode the English input and generate a Hindi translation.English sourceHindi targetRepresent the source and condition target generation.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Encode the English input and generate a Hindi translation.. English source. Hindi target. Represent the source and condition target generation.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Encode the English input and generate a Hindi translation.
  • English source
  • Hindi target
  • Represent the source and condition target generation.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part III · Sequence-to-sequence prediction

Encode the English input and generate a Hindi translation.

Teaching note

Part III · Sequence-to-sequence prediction Pause and establish the requested output before advancing.

Frame overflows: split the content

03 · Represent one sequence, generate another

Example: English-to-Hindi translation

Raghav goes to school in Delhi.. राघव दिल्ली में स्कूल जाता है।. Encode the source sentence and generate the target sequence.Raghav goes to school in Delhi.राघव दिल्ली में स्कूल जाता है।Encode the source sentence and generate the target sequence.

राघव दिल्ली में स्कूल जाता है।

Read the diagram labels
  • Raghav goes to school in Delhi.. राघव दिल्ली में स्कूल जाता है।. Encode the source sentence and generate the target sequence.
  • Raghav goes to school in Delhi.
  • राघव दिल्ली में स्कूल जाता है।
  • Encode the source sentence and generate the target sequence.

Authored teaching diagram · Primary source

Do contextual source vectors alone produce the Hindi sentence?

They represent the input. To generate a variable-length target sentence, we also need a mechanism that predicts target tokens and a stopping token. The translation is authored. Displayed Hindi words are teaching units, not a claim about real model tokenization.

Teaching note

Do contextual source vectors alone produce the Hindi sentence? Notice Delhi moves earlier in the Hindi sentence. Avoid suggesting one fixed source position maps to each output position. The translation is authored. Displayed Hindi words are teaching units, not a claim about real model tokenization.

Frame overflows: split the content

03 · Represent one sequence, generate another

How does Hindi generation begin?

The Hindi labels [शुरू] and [समाप्त] stand for special beginning-of-sequence and end-of-sequence token IDs. They are not ordinary Hindi words in the translated sentence. Supply the start ID, look up its trainable target embedding, add its position and begin decoder computation. Generate six displayed Hindi word tokens and stop on the end ID. These are teaching labels, not a claim about the exact spellings or special token conventions of a particular tokenizer.[शुरू] means “begin generating” · a special token ID[शुरू]Target embedding table: special rowe1Look up e1, add p1, then run the decoder with the English source.We supply [शुरू]. The first predicted Hindi word is राघव.[शुरू] राघव दिल्ली में स्कूल जाता है [समाप्त]The markers control generation; the translation contains only the six Hindi words.

[शुरू] starts the process; [समाप्त] tells us to stop.

Read the diagram labels
  • The Hindi labels [शुरू] and [समाप्त] stand for special beginning-of-sequence and end-of-sequence token IDs. They are not ordinary Hindi words in the translated sentence. Supply the start ID, look up its trainable target embedding, add its position and begin decoder computation. Generate six displayed Hindi word tokens and stop on the end ID. These are teaching labels, not a claim about the exact spellings or special token conventions of a particular tokenizer.
  • [शुरू] means “begin generating” · a special token ID
  • [शुरू]
  • Target embedding table: special row
  • e1
  • Look up e1, add p1, then run the decoder with the English source.
  • We supply [शुरू]. The first predicted Hindi word is राघव.
  • [शुरू] राघव दिल्ली में स्कूल जाता है [समाप्त]
  • The markers control generation; the translation contains only the six Hindi words.

Authored teaching diagram

Is शुरू the first word of the translation?

No. [शुरू] is our Hindi display label for a special beginning-of-sequence token, often called START or BOS. We supply its token ID and look up its embedding e₁ in the target embedding table. Its decoder output predicts the first Hindi word राघव. [समाप्त] labels the end token; it is not printed as part of the translated sentence. The marker labels are a teaching convention, not a tokenizer configuration change. A real model uses its configured decoder start ID, which can be a BOS, EOS, pad or other designated ID. Here we use a simple BOS/EOS teaching design. The start row is a learned parameter loaded from a checkpoint or initialized before training. Each displayed Hindi word is treated as one token for clarity; actual tokenizers may split it. The target embedding is updated through the decoder while its table row is fixed during inference.

Teaching note

Is शुरू the first word of the translation? Point to the brackets: these mark special tokens. Trace the start ID to its embedding-table row, then read the complete example output. We do not supply the complete target sentence during generation. The marker labels are a teaching convention, not a tokenizer configuration change. A real model uses its configured decoder start ID, which can be a BOS, EOS, pad or other designated ID. Here we use a simple BOS/EOS teaching design. The start row is a learned parameter loaded from a checkpoint or initialized before training. Each displayed Hindi word is treated as one token for clarity; actual tokenizers may split it. The target embedding is updated through the decoder while its table row is fixed during inference.

Frame overflows: split the content

03 · Represent one sequence, generate another

Step 1: encode the source sentence

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Keep one contextual vector per source position.

The same encoder states we used for NER.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Keep one contextual vector per source position.

Authored teaching diagram · Primary source

What new operation have we introduced so far?

None. English tokens receive embeddings and positions, then the encoder returns one contextual state per source position.

Teaching note

What new operation have we introduced so far? Keep the source strip, encoder band and vector positions fixed for all the following builds.

Frame overflows: split the content

03 · Represent one sequence, generate another

First attempt: compress the source into one vector

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. One fixed source vector, c ∈ ℝdRaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Pool → cOne fixed source vector, c ∈ ℝd

c = Pool(e₁, …, e₆)

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. One fixed source vector, c ∈ ℝd
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Pool → c
  • One fixed source vector, c ∈ ℝd

Authored teaching diagram · Primary source

Where have we already used this operation?

In whole-sentence classification. Here the pooled vector will condition generation instead of feeding a class head. This first design gives the decoder one pooled source vector. We will choose how the decoder accesses the source in the following section.

Teaching note

Where have we already used this operation? Reveal the pooling arrows below the unchanged encoder states. This first design gives the decoder one pooled source vector. We will choose how the decoder accesses the source in the following section.

Frame overflows: split the content

03 · Represent one sequence, generate another

Conditioning the decoder on the source vector

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi tokenRaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Pool → cCausal decodertarget prefix[शुरू] राघव …next Hindi token

p(yₜ | y<ₜ, c)

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi token
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Pool → c
  • Causal decoder
  • target prefix
  • [शुरू] राघव …
  • next Hindi token

Authored teaching diagram · Primary source

What information does the decoder need at this step?

It needs the target tokens so far and the source context c. For now, imagine the decoder has access to c at every step; the next diagrams show three ways to provide it. The diagram states a conditional dependency. It does not specify an extra token or a particular source interface. We first show the dependency, then compare concrete implementations.

Teaching note

What information does the decoder need at this step? Add the familiar causal decoder beneath c. Show the target prefix entering from the left. The diagram states a conditional dependency. It does not specify an extra token or a particular source interface. We first show the dependency, then compare concrete implementations.

Frame overflows: split the content

03 · Represent one sequence, generate another

A fixed source vector at each generation step

source. target prefix. next token. same c. +. [शुरू]. राघव. same c. +. [शुरू] राघव. दिल्ली. same c. +. [शुरू] राघव दिल्ली. में. p(yt | y<t, c)sourcetarget prefixnext tokensame c+[शुरू]राघवsame c+[शुरू] राघवदिल्लीsame c+[शुरू] राघव दिल्लीमेंp(yt | y<t, c)

The decoder has access to c at every step.

Read the diagram labels
  • source. target prefix. next token. same c. +. [शुरू]. राघव. same c. +. [शुरू] राघव. दिल्ली. same c. +. [शुरू] राघव दिल्ली. में. p(yt | y<t, c)
  • source
  • target prefix
  • next token
  • same c
  • +
  • [शुरू]
  • राघव
  • same c
  • +
  • [शुरू] राघव
  • दिल्ली
  • same c
  • +
  • [शुरू] राघव दिल्ली
  • में
  • p(yt | y<t, c)

Authored teaching diagram · Primary source

What changes between these three predictions?

The source context c stays fixed. The target prefix grows, so the decoder state and next-token prediction can change. This frame describes p(yₜ | y<ₜ, c). The following frames implement that dependency in three different ways; the plus sign on this frame means both inputs are available, not necessarily numerical addition. Continue through स्कूल, जाता and है, then predict [समाप्त]. The outputs are teaching examples.

Teaching note

What changes between these three predictions? Read the three rows in order. Keep pointing back to the same c. This frame describes p(yₜ | y<ₜ, c). The following frames implement that dependency in three different ways; the plus sign on this frame means both inputs are available, not necessarily numerical addition. Continue through स्कूल, जाता and है, then predict [समाप्त]. The outputs are teaching examples.

Frame overflows: split the content

03 · Represent one sequence, generate another

Option 1: add context to every target embedding

English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Project: u = c WP. WP: dc × d. Add to every position: et + u. Add pt → target self-attention → MLP. Vocabulary head → softmax → choose मेंEnglish → Encoder → e1 … e6 → Poolcdc numbers[शुरू] राघव दिल्लीe1, e2, e3 · each has d numbersProject: u = c WPWP: dc × dAdd to every position: et + uAdd pt → target self-attention → MLPVocabulary head → softmax → choose में

Project c to the decoder width, then broadcast the same vector across target positions.

Read the diagram labels
  • English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Project: u = c WP. WP: dc × d. Add to every position: et + u. Add pt → target self-attention → MLP. Vocabulary head → softmax → choose में
  • English → Encoder → e1 … e6 → Pool
  • c
  • dc numbers
  • [शुरू] राघव दिल्ली
  • e1, e2, e3 · each has d numbers
  • Project: u = c WP
  • WP: dc × d
  • Add to every position: et + u
  • Add pt → target self-attention → MLP
  • Vocabulary head → softmax → choose में

Authored teaching diagram

Can we add c directly to a target embedding?

Only if their dimensions match. Here c has d_c features and each target embedding has d features. Learn W_P with shape d_c × d, compute u = c W_P, then add u to each eₜ before positional information and the causal decoder. This is a concrete teaching architecture, not a required definition of an encoder-decoder. The decoder input has the same number of positions as the target prefix. The source encoder, projection and decoder are trained together on next-target-token loss. The source vector is recomputed for each source sentence and held fixed while that sentence is decoded. Normalisation and biases are omitted for clarity.

Teaching note

Can we add c directly to a target embedding? Trace the same source encoder and pool into c, then follow the projection and the repeated addition at the target positions. This is a concrete teaching architecture, not a required definition of an encoder-decoder. The decoder input has the same number of positions as the target prefix. The source encoder, projection and decoder are trained together on next-target-token loss. The source vector is recomputed for each source sentence and held fixed while that sentence is decoded. Normalisation and biases are omitted for clarity.

Frame overflows: split the content

03 · Represent one sequence, generate another

Add a position vector at each target position

[शुरू]. e1 + u. p1. +. e1 + u + p1. राघव. e2 + u. p2. +. e2 + u + p2. दिल्ली. e3 + u. p3. +. e3 + u + p3. pi is a position vector with d numbers, just like ei and u.. Same source context u; a different pi for each position.. Add matching coordinates. The result still has d features.[शुरू]e1 + up1+e1 + u + p1राघवe2 + up2+e2 + u + p2दिल्लीe3 + up3+e3 + u + p3pi is a position vector with d numbers, just like ei and u.Same source context u; a different pi for each position.Add matching coordinates. The result still has d features.

eᵢ identifies the token; u carries source context; pᵢ marks the position.

Read the diagram labels
  • [शुरू]. e1 + u. p1. +. e1 + u + p1. राघव. e2 + u. p2. +. e2 + u + p2. दिल्ली. e3 + u. p3. +. e3 + u + p3. pi is a position vector with d numbers, just like ei and u.. Same source context u; a different pi for each position.. Add matching coordinates. The result still has d features.
  • [शुरू]
  • e1 + u
  • p1
  • +
  • e1 + u + p1
  • राघव
  • e2 + u
  • p2
  • +
  • e2 + u + p2
  • दिल्ली
  • e3 + u
  • p3
  • +
  • e3 + u + p3
  • pi is a position vector with d numbers, just like ei and u.
  • Same source context u; a different pi for each position.
  • Add matching coordinates. The result still has d features.

Authored teaching diagram

What does “add positions” actually add?

In this example, pᵢ is a d-dimensional position vector. We add it coordinate by coordinate to the context-conditioned embedding eᵢ + u. शुरू receives p₁, राघव receives p₂ and दिल्ली receives p₃. The decoder therefore receives e₁ + u + p₁, e₂ + u + p₂ and e₃ + u + p₃. This is an additive positional-encoding design. Position vectors may be learned or fixed sinusoidal vectors; their construction is not needed here. Other Transformer designs incorporate positions differently, for example through rotations of queries and keys. We index illustrated positions from one for teaching. Add the positional information once, after the displayed source-context fusion. For concatenation, add pᵢ after W_F has returned the fused vector to width d. For prepending, the source prefix occupies position 1, so the three target positions receive p₂, p₃ and p₄.

Teaching note

What does “add positions” actually add? Follow each vertical column. The source context u is shared; the token embedding and position vector differ. Point at each plus node and count d features before and after addition. This is an additive positional-encoding design. Position vectors may be learned or fixed sinusoidal vectors; their construction is not needed here. Other Transformer designs incorporate positions differently, for example through rotations of queries and keys. We index illustrated positions from one for teaching. Add the positional information once, after the displayed source-context fusion. For concatenation, add pᵢ after W_F has returned the fused vector to width d. For prepending, the source prefix occupies position 1, so the three target positions receive p₂, p₃ and p₄.

Frame overflows: split the content

03 · Represent one sequence, generate another

Causal self-attention over the Hindi prefix

Causal target self-attention. Query rows शुरू, Raghav and Delhi can read target key/value columns up to their own position. The future token mein is blocked and not supplied. Queries, keys and values are projections of the target-side current embeddings, which already include the added source context and positional information.The same attention operation, now among Hindi target positions.Keys / values: positions we can read[शुरू]राघवदिल्लीमें[शुरू]✓×××राघव✓✓××दिल्ली✓✓✓×Rows = queries · filled cells = allowedQ, K, V all come from target ei[शुरू]reads itselfराघवreads शुरू and itselfदिल्लीreads all three supplied positionsमें is the next prediction. Its embedding is not an input yet.

Target queries read target keys and values. The causal mask keeps future tokens hidden.

Read the diagram labels
  • Causal target self-attention. Query rows शुरू, Raghav and Delhi can read target key/value columns up to their own position. The future token mein is blocked and not supplied. Queries, keys and values are projections of the target-side current embeddings, which already include the added source context and positional information.
  • The same attention operation, now among Hindi target positions.
  • Keys / values: positions we can read
  • [शुरू]
  • राघव
  • दिल्ली
  • में
  • [शुरू]
  • ✓
  • ×
  • ×
  • ×
  • राघव
  • ✓
  • ✓
  • ×
  • ×
  • दिल्ली
  • ✓
  • ✓
  • ✓
  • ×
  • Rows = queries · filled cells = allowed
  • Q, K, V all come from target ei
  • [शुरू]
  • reads itself
  • राघव
  • reads शुरू and itself
  • दिल्ली
  • reads all three supplied positions
  • में is the next prediction. Its embedding is not an input yet.

Authored teaching diagram

Is the decoder using self-attention again?

Yes. शुरू, राघव and दिल्ली exchange information through causal self-attention. Each target position supplies a query, key and value from its current embedding. शुरू can read itself; राघव can also read शुरू; दिल्ली can read all three. The next token में is not supplied yet. These embeddings already contain u and positional information. Their queries, keys and values are all target-side projections: this is self-attention. There is no separate query-to-English-token lookup in this fixed-context design. During parallel training, future supplied target positions are masked; during this generation step, they are absent. The same rule applies after concatenation and projection. With a prepended source vector, that extra earlier position also provides a key and value.

Teaching note

Is the decoder using self-attention again? Read each matrix row as one receiving target position. The gray में column is a teaching placeholder for a future token, not an actual input at this generation step. These embeddings already contain u and positional information. Their queries, keys and values are all target-side projections: this is self-attention. There is no separate query-to-English-token lookup in this fixed-context design. During parallel training, future supplied target positions are masked; during this generation step, they are absent. The same rule applies after concatenation and projection. With a prepended source vector, that extra earlier position also provides a key and value.

Frame overflows: split the content

03 · Represent one sequence, generate another

Updating the embedding at दिल्ली

One self-attention head at target position 3. Use the current Delhi embedding for q3 and all three supplied target embeddings for keys and values. Scaled query-key scores with a causal mask become softmax weights; the weighted values form message m3. Project with W_O, add the residual, then use the MLP and its residual. With multiple heads, concatenate their messages before W_O. The final updated Delhi embedding predicts the next token mein.Here ei means the current target embedding, after adding u and pi.दिल्ली: current e3[शुरू] राघव दिल्ली: current e1, e2, e3q3 = e3 WQki = ei WK; vi = ei WVMatch q3 with k1, k2, k3Mask future positionsSoftmax → α1, α2, α3Message m3 = α1 v1 + α2 v2 + α3 v3valuesΔe3 = m3 WO → add to e3 → MLP + residual → updated e3

The message m₃ updates दिल्ली; its final embedding will predict में.

Read the diagram labels
  • One self-attention head at target position 3. Use the current Delhi embedding for q3 and all three supplied target embeddings for keys and values. Scaled query-key scores with a causal mask become softmax weights; the weighted values form message m3. Project with W_O, add the residual, then use the MLP and its residual. With multiple heads, concatenate their messages before W_O. The final updated Delhi embedding predicts the next token mein.
  • Here ei means the current target embedding, after adding u and pi.
  • दिल्ली: current e3
  • [शुरू] राघव दिल्ली: current e1, e2, e3
  • q3 = e3 WQ
  • ki = ei WK; vi = ei WV
  • Match q3 with k1, k2, k3
  • Mask future positions
  • Softmax → α1, α2, α3
  • Message m3 = α1 v1 + α2 v2 + α3 v3
  • values
  • Δe3 = m3 WO → add to e3 → MLP + residual → updated e3

Authored teaching diagram

Where does the attention message go?

Project the current Delhi embedding into q₃. Compare it with the keys of the supplied target positions, scale the scores by the square root of the key width, apply the causal mask and softmax, then combine their values into m₃. Project the message with W_O to get Δe₃ and add it to the current embedding. The MLP and its residual follow. The displayed equations describe one head. Multiple heads compute their own messages; concatenate them before W_O. Layer normalization is omitted from the drawing. Here eᵢ denotes the current embedding entering the attention sublayer: in the first block this is the prepared eᵢ + u + pᵢ from the previous slide. Scores use q₃ kᵢ transpose divided by sqrt(d_k). Repeat the block updates before reading the last position through the vocabulary head. The third embedding belongs to दिल्ली, not to the token being predicted.

Teaching note

Where does the attention message go? Point from q₃ to the three keys, then from the weighted values to m₃. Trace the same message, delta and residual convention from the earlier attention lecture. The displayed equations describe one head. Multiple heads compute their own messages; concatenate them before W_O. Layer normalization is omitted from the drawing. Here eᵢ denotes the current embedding entering the attention sublayer: in the first block this is the prepared eᵢ + u + pᵢ from the previous slide. Scores use q₃ kᵢ transpose divided by sqrt(d_k). Repeat the block updates before reading the last position through the vocabulary head. The third embedding belongs to दिल्ली, not to the token being predicted.

Frame overflows: split the content

03 · Represent one sequence, generate another

From target embeddings to the next Hindi token

[शुरू]. e1 + u + p1. राघव. e2 + u + p2. दिल्ली. e3 + u + p3. . Causal self-attention + residual → MLP + residual. Repeat the blocks. Each position can read itself and earlier positions.. updated e1. updated e2. updated e3. Vocabulary head → |V| logits. Softmax. Choose: में. Read the last target embedding; append में and repeat.[शुरू]e1 + u + p1राघवe2 + u + p2दिल्लीe3 + u + p3Causal self-attention + residual → MLP + residualRepeat the blocks. Each position can read itself and earlier positions.updated e1updated e2updated e3Vocabulary head → |V| logitsSoftmaxChoose: मेंRead the last target embedding; append में and repeat.

The same decoder and vocabulary head are reused at the next generation step.

Read the diagram labels
  • [शुरू]. e1 + u + p1. राघव. e2 + u + p2. दिल्ली. e3 + u + p3. . Causal self-attention + residual → MLP + residual. Repeat the blocks. Each position can read itself and earlier positions.. updated e1. updated e2. updated e3. Vocabulary head → |V| logits. Softmax. Choose: में. Read the last target embedding; append में and repeat.
  • [शुरू]
  • e1 + u + p1
  • राघव
  • e2 + u + p2
  • दिल्ली
  • e3 + u + p3
  • Causal self-attention + residual → MLP + residual
  • Repeat the blocks. Each position can read itself and earlier positions.
  • updated e1
  • updated e2
  • updated e3
  • Vocabulary head → |V| logits
  • Softmax
  • Choose: में
  • Read the last target embedding; append में and repeat.

Authored teaching diagram

What is the “last embedding,” and why read that one?

The decoder updates every supplied position through causal self-attention and MLP blocks. Here the last target token is दिल्ली, so we read its updated e₃. The vocabulary head turns it into |V| scores; softmax gives a vocabulary distribution, from which we choose the example next token में. The subscript of e₃ identifies the third target token, not the token being predicted. The earlier updated embeddings still exist; generation uses the last target position to choose the next token. |V| is the target vocabulary size. Softmax does not itself choose the token. In the addition design, every new target input also receives the same source vector u for this source sentence. Concatenation and prepending change input preparation; the causal blocks and last-target-position readout follow the same pattern. Normalization and the individual attention-head projections are omitted from this summary.

Teaching note

What is the “last embedding,” and why read that one? Trace all three prepared inputs to all three updated outputs. Highlight the last one, then follow it through logits, softmax and token selection. Append में, obtain its token embedding and position vector, and run the next step. The subscript of e₃ identifies the third target token, not the token being predicted. The earlier updated embeddings still exist; generation uses the last target position to choose the next token. |V| is the target vocabulary size. Softmax does not itself choose the token. In the addition design, every new target input also receives the same source vector u for this source sentence. Concatenation and prepending change input preparation; the causal blocks and last-target-position readout follow the same pattern. Normalization and the individual attention-head projections are omitted from this summary.

Frame overflows: split the content

03 · Represent one sequence, generate another

Option 2: concatenate context as extra features

English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Concatenate features: [et ∥ c]. same c beside each et · d + dc features. Project: [et ∥ c] WF → d features. WF: (d + dc) × d. Add pt → causal blocks → updated e3 → logits → softmax → मेंEnglish → Encoder → e1 … e6 → Poolcdc numbers[शुरू] राघव दिल्लीe1, e2, e3 · each has d numbersConcatenate features: [et ∥ c]same c beside each et · d + dc featuresProject: [et ∥ c] WF → d featuresWF: (d + dc) × dAdd pt → causal blocks → updated e3 → logits → softmax → में

Concatenation increases the feature width; it does not add a token position.

Read the diagram labels
  • English → Encoder → e1 … e6 → Pool. c. dc numbers. [शुरू] राघव दिल्ली. e1, e2, e3 · each has d numbers. Concatenate features: [et ∥ c]. same c beside each et · d + dc features. Project: [et ∥ c] WF → d features. WF: (d + dc) × d. Add pt → causal blocks → updated e3 → logits → softmax → में
  • English → Encoder → e1 … e6 → Pool
  • c
  • dc numbers
  • [शुरू] राघव दिल्ली
  • e1, e2, e3 · each has d numbers
  • Concatenate features: [et ∥ c]
  • same c beside each et · d + dc features
  • Project: [et ∥ c] WF → d features
  • WF: (d + dc) × d
  • Add pt → causal blocks → updated e3 → logits → softmax → में

Authored teaching diagram

Does concatenating c create another token?

Not here: [eₜ ∥ c] concatenates features within each target position. Its width is d + d_c. A learned W_F of shape (d + d_c) × d maps it back to the decoder width. Then add the d-dimensional position vector pₜ and use the causal decoder, last-target-position readout, vocabulary logits and softmax just shown. Apply the same fusion map at every target position. Train it with the source encoder and decoder through the translation loss. With a linear fusion, [eₜ ∥ c]W_F can be split into a learned transform of eₜ plus a learned transform of c; these are related conditioning designs, not claims of strictly different expressive power. This diagram is an authored teaching construction.

Teaching note

Does concatenating c create another token? Point to the two feature groups, then the projection back to d. Count the target positions: there are still three. Apply the same fusion map at every target position. Train it with the source encoder and decoder through the translation loss. With a linear fusion, [eₜ ∥ c]W_F can be split into a learned transform of eₜ plus a learned transform of c; these are related conditioning designs, not claims of strictly different expressive power. This diagram is an authored teaching construction.

Frame overflows: split the content

03 · Represent one sequence, generate another

Option 3: prepend context as an extra input position

English → Encoder → e1 … e6 → Pool. c. dc numbers. u = c WP · d numbers. Target embeddings: [शुरू] राघव दिल्ली. . u · source. + p1. . e1 · शुरू. + p2. . e2 · राघव. + p3. . e3 · दिल्ली. + p4. Causal decoder → updated target embeddings → select last e3. Vocabulary head → softmax → choose में. p1 for u; p2, p3, p4 for targetsEnglish → Encoder → e1 … e6 → Poolcdc numbersu = c WP · d numbersTarget embeddings: [शुरू] राघव दिल्लीu · source+ p1e1 · शुरू+ p2e2 · राघव+ p3e3 · दिल्ली+ p4Causal decoder → updated target embeddings → select last e3Vocabulary head → softmax → choose मेंp1 for u; p2, p3, p4 for targets

The target reads the source-prefix vector through ordinary causal self-attention.

Read the diagram labels
  • English → Encoder → e1 … e6 → Pool. c. dc numbers. u = c WP · d numbers. Target embeddings: [शुरू] राघव दिल्ली. . u · source. + p1. . e1 · शुरू. + p2. . e2 · राघव. + p3. . e3 · दिल्ली. + p4. Causal decoder → updated target embeddings → select last e3. Vocabulary head → softmax → choose में. p1 for u; p2, p3, p4 for targets
  • English → Encoder → e1 … e6 → Pool
  • c
  • dc numbers
  • u = c WP · d numbers
  • Target embeddings: [शुरू] राघव दिल्ली
  • u · source
  • + p1
  • e1 · शुरू
  • + p2
  • e2 · राघव
  • + p3
  • e3 · दिल्ली
  • + p4
  • Causal decoder → updated target embeddings → select last e3
  • Vocabulary head → softmax → choose में
  • p1 for u; p2, p3, p4 for targets

Authored teaching diagram

Why place u before शुरू rather than after the target prefix?

Project c to a d-dimensional vector u and put it before the target embeddings. Every target query can then attend to that earlier position under the causal mask. If u were placed after the current prediction position, the causal mask would prevent that position from reading it. The continuous prefix is a vector, not a vocabulary word to predict. Assign positions consistently, including the extra slot, and omit a next-token loss on the source slot in this teaching construction. Train the projection, source encoder and decoder on target next-token loss. Unlike the later cross-attention architecture, this uses one pooled source vector and causal self-attention. The appendix enlarges the same prefix construction.

Teaching note

Why place u before शुरू rather than after the target prefix? Count four input positions. Add p₁ to u, p₂ to the शुरू embedding, p₃ to राघव and p₄ to दिल्ली. Keep e₁, e₂ and e₃ as target-token indices. Read updated e₃, at the fourth total input position, through the vocabulary head and softmax to predict में. The continuous prefix is a vector, not a vocabulary word to predict. Assign positions consistently, including the extra slot, and omit a next-token loss on the source slot in this teaching construction. Train the projection, source encoder and decoder on target next-token loss. Unlike the later cross-attention architecture, this uses one pooled source vector and causal self-attention. The appendix enlarges the same prefix construction.

Frame overflows: split the content

03 · Represent one sequence, generate another

An encoder-decoder generates a target from a source

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi token. ENCODER-DECODER · source context + target prefixRaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Pool → cCausal decodertarget prefix[शुरू] राघव …next Hindi tokenENCODER-DECODER · source context + target prefix

Represent the source + generate the target.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Pool → c. Causal decoder. target prefix. [शुरू] राघव …. next Hindi token. ENCODER-DECODER · source context + target prefix
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Pool → c
  • Causal decoder
  • target prefix
  • [शुरू] राघव …
  • next Hindi token
  • ENCODER-DECODER · source context + target prefix

Authored teaching diagram · Primary source

What are the two sequences in this architecture?

The complete supplied English source and the growing Hindi target. Their positions, lengths and vocabularies need not be identical.

Teaching note

What are the two sequences in this architecture? Name the family now, after following the computation through both stacks.

Frame overflows: split the content

03 · Represent one sequence, generate another

The decoder cannot revisit the source

Eight source embeddings are pooled into one fixed-width vector c. The decoder receives c and a growing target prefix, with no direct path back to the source embeddings. Different target steps need who, where and when details. If pooling fails to preserve Delhi, the source detail cannot be retrieved later. The Hindi phrases are illustrative fragments, not a full token-by-token translation.Raghav goes to school in Delhi every morning.Encodere1e2e3e4e5e6e7e88 updated embeddingspoolcfixed widthDecoder + output headGrowing target prefixDifferent target steps need different source detailsWho?Raghav → राघवWhere?Delhi → दिल्लीWhen?every morning → हर सुबहIf c loses “Delhi”, the decoder cannot look back to recover it.

The decoder relies on c for all source information.

Read the diagram labels
  • Eight source embeddings are pooled into one fixed-width vector c. The decoder receives c and a growing target prefix, with no direct path back to the source embeddings. Different target steps need who, where and when details. If pooling fails to preserve Delhi, the source detail cannot be retrieved later. The Hindi phrases are illustrative fragments, not a full token-by-token translation.
  • Raghav goes to school in Delhi every morning.
  • Encoder
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • e7
  • e8
  • 8 updated embeddings
  • pool
  • c
  • fixed width
  • Decoder + output head
  • Growing target prefix
  • Different target steps need different source details
  • Who?
  • Raghav → राघव
  • Where?
  • Delhi → दिल्ली
  • When?
  • every morning → हर सुबह
  • If c loses “Delhi”, the decoder cannot look back to recover it.

Authored teaching diagram · Primary source

So what is the problem with using one vector?

A fixed vector can work. The restriction is that every source detail needed later must survive pooling into c. If c does not preserve Delhi, a later decoder step cannot query the original Delhi embedding to recover it. Each displayed word is one token for this teaching example, so the source has eight positions. The Hindi items are translated fragments illustrating different information needs, not consecutive next-token predictions. Pooling produces a fixed-width vector even when the source grows longer. This creates a representation and learning burden; it does not prove that a fixed vector must forget information. The decoder can use c differently as its target prefix changes. It may guess a missing detail from learned patterns, but cannot retrieve it from the source. The next architecture keeps all source embeddings available for a fresh lookup.

Teaching note

So what is the problem with using one vector? Trace encoder outputs through pooling into c. Show the separate growing target prefix. Then ask what source information is needed for who, where and when. There is no direct edge from the source embeddings to the decoder. Each displayed word is one token for this teaching example, so the source has eight positions. The Hindi items are translated fragments illustrating different information needs, not consecutive next-token predictions. Pooling produces a fixed-width vector even when the source grows longer. This creates a representation and learning burden; it does not prove that a fixed vector must forget information. The decoder can use c differently as its target prefix changes. It may guess a missing detail from learned patterns, but cannot retrieve it from the source. The next architecture keeps all source embeddings available for a fresh lookup.

Frame overflows: split the content

04 · Cross-attention to the source

Part IV · Cross-attention to the source

Retain one encoder output per source position.. Updated target embedding. queries. Source keys and values. Let each target position query the source directly.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesRetain one encoder output per source position.Updated target embeddingqueriesSource keys and valuesLet each target position query the source directly.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Retain one encoder output per source position.. Updated target embedding. queries. Source keys and values. Let each target position query the source directly.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Retain one encoder output per source position.
  • Updated target embedding
  • queries
  • Source keys and values
  • Let each target position query the source directly.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part IV · Cross-attention to the source

Retain one encoder output per source position.

Teaching note

Part IV · Cross-attention to the source Pause and establish the requested output before advancing.

Frame overflows: split the content

04 · Cross-attention to the source

Retaining the source embeddings for the decoder

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Keep one contextual vector per source position.

Keep e₁ … e₆ available to the decoder.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Keep one contextual vector per source position.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Keep one contextual vector per source position.

Authored teaching diagram · Primary source

What could we retain instead of just c?

The full sequence of encoder output states. The decoder can make a source lookup suited to each current target position.

Teaching note

What could we retain instead of just c? Return to the identical source drawing. Remove the pooling constraint while leaving the encoder unchanged.

Frame overflows: split the content

04 · Cross-attention to the source

Source attention when predicting राघव

To predict the first Hindi token Raghav from शुरू, retain all six English encoder outputs. Illustrative hand-chosen weights 0.72, 0.04, 0.03, 0.09, 0.04, 0.08 put most weight on the source name Raghav. These are explanatory numbers, not measured model attention or guaranteed alignments.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e60.720.040.030.090.040.08An illustrative lookup: put most weight on the English name Raghav.Target prefix: [शुरू] → example next Hindi token: राघवIllustrative attention weights; their sum is 1.

All six English embeddings are available; the target query can assign more weight to Raghav.

Read the diagram labels
  • To predict the first Hindi token Raghav from शुरू, retain all six English encoder outputs. Illustrative hand-chosen weights 0.72, 0.04, 0.03, 0.09, 0.04, 0.08 put most weight on the source name Raghav. These are explanatory numbers, not measured model attention or guaranteed alignments.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • An illustrative lookup: put most weight on the English name Raghav.
  • Target prefix: [शुरू] → example next Hindi token: राघव
  • Illustrative attention weights; their sum is 1.

Authored teaching diagram

When predicting राघव, which source position might receive more attention?

For an intuitive example, give the English Raghav position weight 0.72 and distribute the remaining 0.28 over the other five positions. This illustrates a plausible source lookup for the first Hindi token. These numbers are hand-chosen, not measured attention. The weights sum to one. Real attention heads can use broader or different patterns; a name need not receive the largest weight in every head or layer. Crucially, the query for predicting the first Hindi token comes from the शुरू position. Feeding राघव to predict itself would leak the target. The following slides explain the source of the query, keys, values and weights.

Teaching note

When predicting राघव, which source position might receive more attention? Keep the encoder drawing fixed from the previous slide. Read all six weights, then point to शुरू as the supplied target prefix and राघव as the upcoming prediction. The weights sum to one. Real attention heads can use broader or different patterns; a name need not receive the largest weight in every head or layer. Crucially, the query for predicting the first Hindi token comes from the शुरू position. Feeding राघव to predict itself would leak the target. The following slides explain the source of the query, keys, values and weights.

Frame overflows: split the content

04 · Cross-attention to the source

Step 1: encode the English sentence

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Six encoder outputs, one for each English position.. Keep these contextual embeddings available to the decoder.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6Six encoder outputs, one for each English position.Keep these contextual embeddings available to the decoder.

Retain all six encoder outputs before starting the target decoder.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. Six encoder outputs, one for each English position.. Keep these contextual embeddings available to the decoder.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • Six encoder outputs, one for each English position.
  • Keep these contextual embeddings available to the decoder.

Authored teaching diagram

What information is available from the source?

The English encoder produces six contextual embeddings, one per supplied position. Keep them separately so the decoder can attend to different source positions. This sequence expands one head in one decoder block. English source indices and Hindi target indices are separate. Source e₁ belongs to Raghav; the target e₁ introduced next belongs to [शुरू].

Teaching note

What information is available from the source? Stop at the blue encoder outputs. Do not introduce target queries or source keys and values yet. This sequence expands one head in one decoder block. English source indices and Hindi target indices are separate. Source e₁ belongs to Raghav; the target e₁ introduced next belongs to [शुरू].

Frame overflows: split the content

04 · Cross-attention to the source

Step 2: initialise the target at [शुरू]

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. Initial target embedding. राघव has not been generated; the supplied target is only [शुरू].RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6TARGET: begin with the special token [शुरू][शुरू]input e1 + p1lookup + positionInitial target embeddingराघव has not been generated; the supplied target is only [शुरू].

Look up the start-token embedding and add its position vector.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. Initial target embedding. राघव has not been generated; the supplied target is only [शुरू].
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • TARGET: begin with the special token [शुरू]
  • [शुरू]
  • input e1 + p1
  • lookup + position
  • Initial target embedding
  • राघव has not been generated; the supplied target is only [शुरू].

Authored teaching diagram

Where does the first target embedding come from?

The special start token has an entry in the target embedding table. Look up its input e₁ and add p₁. No Hindi word has been generated yet. [शुरू] is the teaching label for the special beginning-of-sequence token. It is not an ordinary word in the translation. During inference, lookup reads the learned table row; it does not train that row.

Teaching note

Where does the first target embedding come from? Keep the encoder outputs fixed. Reveal only the start token, embedding lookup and position addition. [शुरू] is the teaching label for the special beginning-of-sequence token. It is not an ordinary word in the translation. During inference, lookup reads the learned table row; it does not train that row.

Frame overflows: split the content

04 · Cross-attention to the source

Step 3: update [शुरू] with target self-attention

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. Target Q, K, V → m1 → Δe1 → e1 + Δe1. Only [शुरू] is available: self-attention reads this one target position.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6TARGET: begin with the special token [शुरू][शुरू]input e1 + p1lookup + positionCausal self-attention+ residualupdated e1Target Q, K, V → m1 → Δe1 → e1 + Δe1Only [शुरू] is available: self-attention reads this one target position.

Causal self-attention updates the target embedding before cross-attention.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. Target Q, K, V → m1 → Δe1 → e1 + Δe1. Only [शुरू] is available: self-attention reads this one target position.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • TARGET: begin with the special token [शुरू]
  • [शुरू]
  • input e1 + p1
  • lookup + position
  • Causal self-attention
  • + residual
  • updated e1
  • Target Q, K, V → m1 → Δe1 → e1 + Δe1
  • Only [शुरू] is available: self-attention reads this one target position.

Authored teaching diagram

With only [शुरू] supplied, what can target self-attention read?

Only the start position itself. Its current embedding supplies the self-attention query, key and value. The attention message is projected and added through the residual connection to update e₁. With a single unmasked key, the attention weight is one, but the value and output projections can still produce a nonzero update. Once more Hindi tokens are supplied, causal self-attention can use their preceding target context. Layer normalization is omitted. The target MLP follows cross-attention in this decoder block.

Teaching note

With only [शुरू] supplied, what can target self-attention read? Trace input e₁ plus position through self-attention and the residual to updated e₁. There is no token prediction at this stage. With a single unmasked key, the attention weight is one, but the value and output projections can still produce a nonzero update. Once more Hindi tokens are supplied, causal self-attention can use their preceding target context. Layer normalization is omitted. The target MLP follows cross-attention in this decoder block.

Frame overflows: split the content

04 · Cross-attention to the source

Step 4: form a new query for cross-attention

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Cross-attention uses its own query projection WQ.. Its input is e1 after the target self-attention update.RaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6TARGET: begin with the special token [शुरू][शुरू]input e1 + p1lookup + positionCausal self-attention+ residualupdated e1q1 = e1 WQCross-attention uses its own query projection WQ.Its input is e1 after the target self-attention update.

Self-attention and cross-attention have separate learned query projections.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Cross-attention uses its own query projection WQ.. Its input is e1 after the target self-attention update.
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • TARGET: begin with the special token [शुरू]
  • [शुरू]
  • input e1 + p1
  • lookup + position
  • Causal self-attention
  • + residual
  • updated e1
  • q1 = e1 WQ
  • Cross-attention uses its own query projection WQ.
  • Its input is e1 after the target self-attention update.

Authored teaching diagram

Is this the query used in the preceding target self-attention?

No. Cross-attention applies its own W_Q to the updated target embedding. The self-attention query was formed within the preceding sublayer from its input embedding. The two sublayers have distinct parameters and read different sequences. W_Q here means the cross-attention query projection; the self-attention sublayer has a different W_Q. The local symbol is reused for the same mathematical role, not to imply shared weights. In implementations with pre-normalization, the projection takes a normalized version of this updated target stream. Normalization is omitted from the drawing.

Teaching note

Is this the query used in the preceding target self-attention? Reveal the new arrow from updated e₁ to q₁. Keep the e₁ subscript because this is still the start position. The cross-attention query will be matched against English source keys. W_Q here means the cross-attention query projection; the self-attention sublayer has a different W_Q. The local symbol is reused for the same mathematical role, not to imply shared weights. In implementations with pre-normalization, the projection takes a normalized version of this updated target stream. Normalization is omitted from the drawing.

Frame overflows: split the content

04 · Cross-attention to the source

Step 5: match the target query to English keys

Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. k1, v1. k2, v2. k3, v3. k4, v4. k5, v5. k6, v6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Source: ki = ei WK; vi = ei WV. Compare q1 with all six source keys: q1 kiT / √dkRaghavgoestoschoolinDelhiENCODER · whole-input self-attentione1e2e3e4e5e6k1, v1k2, v2k3, v3k4, v4k5, v5k6, v6TARGET: begin with the special token [शुरू][शुरू]input e1 + p1lookup + positionCausal self-attention+ residualupdated e1q1 = e1 WQSource: ki = ei WK; vi = ei WVCompare q1 with all six source keys: q1 kiT / √dk

Cross-attention: Q from the updated target; K and V from encoder outputs.

Read the diagram labels
  • Raghav. goes. to. school. in. Delhi. ENCODER · whole-input self-attention. . e1. . e2. . e3. . e4. . e5. . e6. k1, v1. k2, v2. k3, v3. k4, v4. k5, v5. k6, v6. TARGET: begin with the special token [शुरू]. [शुरू]. . input e1 + p1. lookup + position. . Causal self-attention. + residual. updated e1. q1 = e1 WQ. Source: ki = ei WK; vi = ei WV. Compare q1 with all six source keys: q1 kiT / √dk
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • ENCODER · whole-input self-attention
  • e1
  • e2
  • e3
  • e4
  • e5
  • e6
  • k1, v1
  • k2, v2
  • k3, v3
  • k4, v4
  • k5, v5
  • k6, v6
  • TARGET: begin with the special token [शुरू]
  • [शुरू]
  • input e1 + p1
  • lookup + position
  • Causal self-attention
  • + residual
  • updated e1
  • q1 = e1 WQ
  • Source: ki = ei WK; vi = ei WV
  • Compare q1 with all six source keys: q1 kiT / √dk

Authored teaching diagram

How are the English keys and values constructed?

Apply the cross-attention W_K and W_V to each encoder output to obtain six source keys and six values. Compare q₁ with each key using a scaled dot product. The next slide normalizes the scores and combines the values. Cross-attention has its own learned W_Q, W_K and W_V, separate from target and source self-attention parameters. Source keys and values can be computed in advance; this reveal order is for teaching and does not impose an execution order. Keys determine source-position weights; values supply the message. The query at [शुरू] supports prediction of राघव after the remaining decoder computation.

Teaching note

How are the English keys and values constructed? Reveal the six key/value pairs and the query-to-key connections. Follow the English paths down into K and V, then the updated target path into Q. Cross-attention has its own learned W_Q, W_K and W_V, separate from target and source self-attention parameters. Source keys and values can be computed in advance; this reveal order is for teaching and does not impose an execution order. Keys determine source-position weights; values supply the message. The query at [शुरू] supports prediction of राघव after the remaining decoder computation.

Frame overflows: split the content

04 · Cross-attention to the source

Computing the weighted source message

Illustrative scaled dot-product scores are the natural logarithms of weights 0.72, 0.04, 0.03, 0.09, 0.04 and 0.08. Scores are displayed rounded to two decimals. Softmax produces those weights, which multiply source values; sum them to obtain cross-attention message c1. Project the message, add it to the target stream, apply MLP and residual processing, and use the final vocabulary head to predict the illustrative Hindi token Raghav. The attention weights are source-position weights, not Hindi vocabulary probabilities.Match the target q1 with each English key: q1 kiT / √dkRaghavkey k1-0.33goeskey k2-3.22tokey k3-3.51schoolkey k4-2.41inkey k5-3.22Delhikey k6-2.53Softmax across the six source scores → weights that sum to 10.72 × v10.04 × v20.03 × v30.09 × v40.04 × v50.08 × v6Add weighted source values → message c1c1 WO + target e1 → MLP + residual → vocabulary head → राघव

Keys determine the weights; values supply the information sent to the target.

Read the diagram labels
  • Illustrative scaled dot-product scores are the natural logarithms of weights 0.72, 0.04, 0.03, 0.09, 0.04 and 0.08. Scores are displayed rounded to two decimals. Softmax produces those weights, which multiply source values; sum them to obtain cross-attention message c1. Project the message, add it to the target stream, apply MLP and residual processing, and use the final vocabulary head to predict the illustrative Hindi token Raghav. The attention weights are source-position weights, not Hindi vocabulary probabilities.
  • Match the target q1 with each English key: q1 kiT / √dk
  • Raghav
  • key k1
  • -0.33
  • goes
  • key k2
  • -3.22
  • to
  • key k3
  • -3.51
  • school
  • key k4
  • -2.41
  • in
  • key k5
  • -3.22
  • Delhi
  • key k6
  • -2.53
  • Softmax across the six source scores → weights that sum to 1
  • 0.72 × v1
  • 0.04 × v2
  • 0.03 × v3
  • 0.09 × v4
  • 0.04 × v5
  • 0.08 × v6
  • Add weighted source values → message c1
  • c1 WO + target e1 → MLP + residual → vocabulary head → राघव

Authored teaching diagram

How do the illustrative weights become information the decoder can use?

Score q₁ against all six English keys, scale by sqrt(d_k), and apply softmax across source positions. Multiply each source value by its weight and add the six weighted vectors to get c₁. Project that message and add it to the target embedding; after the remaining decoder computation, the vocabulary head can predict राघव. To make the arithmetic internally consistent, the authored scaled scores are log(0.72), log(0.04), log(0.03), log(0.09), log(0.04), log(0.08); displayed scores are rounded to two decimals. Softmax of the unrounded scores reproduces the weights exactly. These are not trained model scores. c₁ is a cross-attention message, analogous to mᵢ in self-attention. For multiple heads, concatenate their messages before W_O. The decoder retains the target stream through the residual; c₁ alone is neither the final embedding nor a vocabulary distribution.

Teaching note

How do the illustrative weights become information the decoder can use? Trace source scores downward through softmax to the weighted values. The larger 0.72 weights the Raghav value more heavily in this illustration. Keep source-attention weights distinct from output-vocabulary probabilities. To make the arithmetic internally consistent, the authored scaled scores are log(0.72), log(0.04), log(0.03), log(0.09), log(0.04), log(0.08); displayed scores are rounded to two decimals. Softmax of the unrounded scores reproduces the weights exactly. These are not trained model scores. c₁ is a cross-attention message, analogous to mᵢ in self-attention. For multiple heads, concatenate their messages before W_O. The decoder retains the target stream through the residual; c₁ alone is neither the final embedding nor a vocabulary distribution.

Frame overflows: split the content

04 · Cross-attention to the source

The source message updates शुरू’s embedding

Track शुरू through embedding lookup, position addition and target causal self-attention. Its current target embedding e1 branches: W_Q makes a query, while the residual keeps e1. English encoder outputs supply keys and values. Cross-attention returns illustrative message c1, emphasizing the source Raghav value. W_O projects c1 into delta e1; the plus node adds that delta to the retained current target embedding. This updates the शुरू position, not the English source embeddings. The MLP and next-token prediction follow on the next frame. One head is shown; normalization is omitted.[शुरू] → lookup input e1 → add p1 → causal self-attention + residualTARGET · शुरू positioncurrent target e1q1 = e1 WQSOURCE · English encoder outputse1 … e6 → K and VOne key and value per English positionCross-attention(q1, K, V)Message c1 = 0.72 v1 + … + 0.08 v6Embedding update: Δe1 = c1 WOKeep the current e1+updated e1 = e1 + Δe1The source message changes शुरू’s embedding.

Keep the current target embedding, then add the update retrieved from the source.

Read the diagram labels
  • Track शुरू through embedding lookup, position addition and target causal self-attention. Its current target embedding e1 branches: W_Q makes a query, while the residual keeps e1. English encoder outputs supply keys and values. Cross-attention returns illustrative message c1, emphasizing the source Raghav value. W_O projects c1 into delta e1; the plus node adds that delta to the retained current target embedding. This updates the शुरू position, not the English source embeddings. The MLP and next-token prediction follow on the next frame. One head is shown; normalization is omitted.
  • [शुरू] → lookup input e1 → add p1 → causal self-attention + residual
  • TARGET · शुरू position
  • current target e1
  • q1 = e1 WQ
  • SOURCE · English encoder outputs
  • e1 … e6 → K and V
  • One key and value per English position
  • Cross-attention(q1, K, V)
  • Message c1 = 0.72 v1 + … + 0.08 v6
  • Embedding update: Δe1 = c1 WO
  • Keep the current e1
  • +
  • updated e1 = e1 + Δe1
  • The source message changes शुरू’s embedding.

Authored teaching diagram

What exactly changes when शुरू attends to the English sentence?

शुरू begins with an embedding-table lookup and positional information. Target self-attention and its residual produce the current e₁. That embedding makes q₁ while a residual path keeps e₁. Cross-attention retrieves the source message c₁; W_O maps it to Δe₁. Add Δe₁ to the retained e₁, giving the शुरू position an embedding informed by the English source. This is one head in one decoder block; with multiple heads, concatenate their source messages before W_O. Layer normalization is omitted. The residual adds to the current target representation entering cross-attention, which has already passed through target self-attention; it does not jump back to the raw embedding-table row. The message c₁ plays the role of the earlier attention message mᵢ. The shown 0.72 weight is illustrative, not measured attention. A forward pass updates this contextual representation; it does not by itself train or overwrite the embedding table. Here the source is read through K/V, not through an added fixed pooled u.

Teaching note

What exactly changes when शुरू attends to the English sentence? Follow the शुरू stream, then pause at the fork: one copy makes the query and the other follows the long residual arrow. Trace the source message through W_O into the plus node. The result is still the embedding at शुरू. This is one head in one decoder block; with multiple heads, concatenate their source messages before W_O. Layer normalization is omitted. The residual adds to the current target representation entering cross-attention, which has already passed through target self-attention; it does not jump back to the raw embedding-table row. The message c₁ plays the role of the earlier attention message mᵢ. The shown 0.72 weight is illustrative, not measured attention. A forward pass updates this contextual representation; it does not by itself train or overwrite the embedding table. Here the source is read through K/V, not through an added fixed pooled u.

Frame overflows: split the content

04 · Cross-attention to the source

Read the updated शुरू embedding to predict राघव

The source-updated शुरू embedding continues through the MLP and its residual. After the repeated decoder blocks, the vocabulary head reads the final embedding at शुरू to produce logits, softmax probabilities, and the illustrative token Raghav. Append Raghav at target position 2 and obtain its input embedding from the target embedding table; the next step can predict Delhi. शुरू is not replaced by Raghav. Source encoder outputs stay available across steps.The शुरू embedding after cross-attention.e1 after source updateMLP + residualupdated शुरू e1Repeat the decoder blocks; then read the final embedding at शुरू.Vocabulary head → logitsSoftmax → probabilitiesChoose: राघवAppend राघव: [शुरू] राघव → next prediction: दिल्लीe1 still belongs to शुरू. The new token राघव gets input embedding e2.

शुरू keeps its position; its final embedding is used to predict the next token.

Read the diagram labels
  • The source-updated शुरू embedding continues through the MLP and its residual. After the repeated decoder blocks, the vocabulary head reads the final embedding at शुरू to produce logits, softmax probabilities, and the illustrative token Raghav. Append Raghav at target position 2 and obtain its input embedding from the target embedding table; the next step can predict Delhi. शुरू is not replaced by Raghav. Source encoder outputs stay available across steps.
  • The शुरू embedding after cross-attention.
  • e1 after source update
  • MLP + residual
  • updated शुरू e1
  • Repeat the decoder blocks; then read the final embedding at शुरू.
  • Vocabulary head → logits
  • Softmax → probabilities
  • Choose: राघव
  • Append राघव: [शुरू] राघव → next prediction: दिल्ली
  • e1 still belongs to शुरू. The new token राघव gets input embedding e2.

Authored teaching diagram

Does the updated शुरू embedding become the embedding of राघव?

No. It remains the contextual embedding at शुरू. After the remaining decoder blocks, a vocabulary head and softmax turn it into next-token probabilities. Choosing राघव appends a new token at position 2, with its own embedding-table lookup e₂. The next generation step can then predict दिल्ली. The decoder repeats causal self-attention, cross-attention, and MLP sublayers with their residuals across its blocks. The top row shows the continuation of the block enlarged on the preceding slide. Read the final target embedding after all blocks. Each new token receives its input embedding and positional information. English encoder outputs, and each layer’s source keys and values, can be reused while the target prefix grows. Predictions are illustrative.

Teaching note

Does the updated शुरू embedding become the embedding of राघव? Trace the updated शुरू embedding through the MLP, final vocabulary scores, probabilities and chosen token. Point to both positions in the new prefix: शुरू remains first, राघव is second. The decoder repeats causal self-attention, cross-attention, and MLP sublayers with their residuals across its blocks. The top row shows the continuation of the block enlarged on the preceding slide. Read the final target embedding after all blocks. Each new token receives its input embedding and positional information. English encoder outputs, and each layer’s source keys and values, can be reused while the target prefix grows. Predictions are illustrative.

Frame overflows: split the content

04 · Cross-attention to the source

The roles of keys and values

Source E = [e1 … e6]. Target embedding et. V = E WV. K = E WK. qt = et WQ. scores: qt KT / √dk. softmax → αt. αt V = ct. valuesSource E = [e1 … e6]Target embedding etV = E WVK = E WKqt = et WQscores: qt KT / √dksoftmax → αtαt V = ctvalues

Query-key scores determine the weights used to combine the values.

Read the diagram labels
  • Source E = [e1 … e6]. Target embedding et. V = E WV. K = E WK. qt = et WQ. scores: qt KT / √dk. softmax → αt. αt V = ct. values
  • Source E = [e1 … e6]
  • Target embedding et
  • V = E WV
  • K = E WK
  • qt = et WQ
  • scores: qt KT / √dk
  • softmax → αt
  • αt V = ct
  • values

Authored teaching diagram · Primary source

Why do we need both keys and values?

The query is compared with keys to produce scores. Softmax turns those scores into weights. Those weights mix the value vectors to produce cₜ. E contains final encoder embeddings. eₜ is the current target embedding after causal self-attention and before the cross-attention update. cₜ is the retrieved message, playing the same role as mᵢ earlier. Project the message to an embedding update, add it to the target stream, then apply the MLP and its residual. Different layers and heads generally use different projection parameters.

Teaching note

Why do we need both keys and values? Follow the K path into the scores, then the separate V path into the weighted sum. The query affects the sum through its weights. E contains final encoder embeddings. eₜ is the current target embedding after causal self-attention and before the cross-attention update. cₜ is the retrieved message, playing the same role as mᵢ earlier. Project the message to an embedding update, add it to the target stream, then apply the MLP and its residual. Different layers and heads generally use different projection parameters.

Frame overflows: split the content

04 · Cross-attention to the source

The cross-attention calculation

One target query, all source keys and values. qt = et WQ. ki = ei WK. vi = ei WV. αti = softmax over i (qt kiT / √dk). ct = Σi αti viOne target query, all source keys and valuesqt = et WQki = ei WKvi = ei WVαti = softmax over i (qt kiT / √dk)ct = Σi αti vi

cₜ is the cross-attention message, just as mᵢ was the self-attention message.

Read the diagram labels
  • One target query, all source keys and values. qt = et WQ. ki = ei WK. vi = ei WV. αti = softmax over i (qt kiT / √dk). ct = Σi αti vi
  • One target query, all source keys and values
  • qt = et WQ
  • ki = ei WK
  • vi = ei WV
  • αti = softmax over i (qt kiT / √dk)
  • ct = Σi αti vi

Authored teaching diagram · Primary source

Which dimension does this softmax normalize?

The source-position index i. For each target query, the weights over all permitted source positions sum to one. One head is shown. Multiple head messages are concatenated and projected back to model width before combining with the target stream. All equations use row-vector notation.

Teaching note

Which dimension does this softmax normalize? Connect each equation to its corresponding arrow in the preceding diagram. Separate scores, weights and the weighted message. One head is shown. Multiple head messages are concatenated and projected back to model width before combining with the target stream. All equations use row-vector notation.

Frame overflows: split the content

04 · Cross-attention to the source

Multi-head cross-attention uses the same mechanism

Source embeddings E. Target embeddings D. head 1. E → K,V. D → Q. Attention → C1. head 2. E → K,V. D → Q. Attention → C2. head 3. E → K,V. D → Q. Attention → C3. Concatenate(C1, C2, C3) → WO. Qj = D WQ(j) · Kj = E WK(j) · Vj = E WV(j). Cj = softmax(Qj KjT / √dk) VjSource embeddings ETarget embeddings Dhead 1E → K,VD → QAttention → C1head 2E → K,VD → QAttention → C2head 3E → K,VD → QAttention → C3Concatenate(C1, C2, C3) → WOQj = D WQ(j) · Kj = E WK(j) · Vj = E WV(j)Cj = softmax(Qj KjT / √dk) Vj

Each head uses target queries and source keys and values.

Read the diagram labels
  • Source embeddings E. Target embeddings D. head 1. E → K,V. D → Q. Attention → C1. head 2. E → K,V. D → Q. Attention → C2. head 3. E → K,V. D → Q. Attention → C3. Concatenate(C1, C2, C3) → WO. Qj = D WQ(j) · Kj = E WK(j) · Vj = E WV(j). Cj = softmax(Qj KjT / √dk) Vj
  • Source embeddings E
  • Target embeddings D
  • head 1
  • E → K,V
  • D → Q
  • Attention → C1
  • head 2
  • E → K,V
  • D → Q
  • Attention → C2
  • head 3
  • E → K,V
  • D → Q
  • Attention → C3
  • Concatenate(C1, C2, C3) → WO
  • Qj = D WQ(j) · Kj = E WK(j) · Vj = E WV(j)
  • Cj = softmax(Qj KjT / √dk) Vj

Authored teaching diagram · Primary source

What changes when we use several cross-attention heads?

Each head has its own learned query, key and value projections. Every head still takes queries from D (current target embeddings) and keys and values from E (final source embeddings). Concatenate the head outputs, then apply W_O. For head j: Qⱼ = D W_Q⁽ʲ⁾, Kⱼ = E W_K⁽ʲ⁾, Vⱼ = E W_V⁽ʲ⁾. Cⱼ = softmax(QⱼKⱼᵀ / √dₖ)Vⱼ. W_O maps the concatenated messages to model width before the residual update.

Teaching note

What changes when we use several cross-attention heads? Trace the same E and D into all three heads, then join their outputs. Students already know the multi-head calculation; stop after identifying the two origins. For head j: Qⱼ = D W_Q⁽ʲ⁾, Kⱼ = E W_K⁽ʲ⁾, Vⱼ = E W_V⁽ʲ⁾. Cⱼ = softmax(QⱼKⱼᵀ / √dₖ)Vⱼ. W_O maps the concatenated messages to model width before the residual update.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Hindi generation: self-attention, then cross-attention

Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.1 · Hindi self-attention2 · English cross-attentionSelf query from e1 at [शुरू]Keys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]1.00m1 = Σ αi vi · Hindi valuesΔe1 = m1 WO; add to current e1Updated e1 after self-attentionRaghavgoestoschoolinDelhiq1 = updated e1 WQ (cross-attention)q1 KT / √dk → softmax over English0.720.040.030.090.040.08c1 = Σ αi vi · English valuesCross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocksStep 1: vocabulary head → softmax → राघव · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e2 at राघवKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.35राघव0.65m2 = Σ αi vi · Hindi valuesΔe2 = m2 WO; add to current e2Updated e2 after self-attentionRaghavgoestoschoolinDelhiq2 = updated e2 WQ (cross-attention)q2 KT / √dk → softmax over English0.050.030.040.100.080.70c2 = Σ αi vi · English valuesCross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocksStep 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e3 at दिल्लीKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.15राघव0.25दिल्ली0.60m3 = Σ αi vi · Hindi valuesΔe3 = m3 WO; add to current e3Updated e3 after self-attentionRaghavgoestoschoolinDelhiq3 = updated e3 WQ (cross-attention)q3 KT / √dk → softmax over English0.040.040.030.090.700.10c3 = Σ αi vi · English valuesCross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocksStep 3: vocabulary head → softmax → में · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e4 at मेंKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.10राघव0.15दिल्ली0.45में0.30m4 = Σ αi vi · Hindi valuesΔe4 = m4 WO; add to current e4Updated e4 after self-attentionRaghavgoestoschoolinDelhiq4 = updated e4 WQ (cross-attention)q4 KT / √dk → softmax over English0.080.060.040.700.050.07c4 = Σ αi vi · English valuesCross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocksStep 4: vocabulary head → softmax → स्कूल · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e5 at स्कूलKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.08राघव0.12दिल्ली0.20में0.25स्कूल0.35m5 = Σ αi vi · Hindi valuesΔe5 = m5 WO; add to current e5Updated e5 after self-attentionRaghavgoestoschoolinDelhiq5 = updated e5 WQ (cross-attention)q5 KT / √dk → softmax over English0.040.750.040.080.040.05c5 = Σ αi vi · English valuesCross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocksStep 5: vocabulary head → softmax → जाता · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e6 at जाताKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.15में0.15स्कूल0.25जाता0.30m6 = Σ αi vi · Hindi valuesΔe6 = m6 WO; add to current e6Updated e6 after self-attentionRaghavgoestoschoolinDelhiq6 = updated e6 WQ (cross-attention)q6 KT / √dk → softmax over English0.100.350.050.200.150.15c6 = Σ αi vi · English valuesCross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocksStep 6: vocabulary head → softmax → है · append to Hindi prefix1 · Hindi self-attention2 · English cross-attentionSelf query from e7 at हैKeys and values: supplied Hindi positionsSame six English encoder outputsSource keys K and values V[शुरू]0.05राघव0.10दिल्ली0.10में0.15स्कूल0.15जाता0.20है0.25m7 = Σ αi vi · Hindi valuesΔe7 = m7 WO; add to current e7Updated e7 after self-attentionRaghavgoestoschoolinDelhiq7 = updated e7 WQ (cross-attention)q7 KT / √dk → softmax over English0.160.180.160.170.160.17c7 = Σ αi vi · English valuesCross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocksStep 7: vocabulary head → softmax → [समाप्त] · stop

Hindi context grows; English context stays fixed. Both sets of weights are illustrative.

Read the diagram labels
  • Seven Hindi generation steps. Start with the special token [शुरू], then append राघव, दिल्ली, में, स्कूल, जाता, है, and stop when [समाप्त] is chosen. Each panel shows Hindi self-attention over the growing supplied prefix, illustrative target weights, the self-attention message and residual update, a separate cross-attention query, six English source weights, and the source message followed by the target update and vocabulary prediction. Self-attention and cross-attention have distinct learned projections and independently normalized weights. Both sets of numbers are authored teaching examples, not measured alignments.
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e1 at [शुरू]
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 1.00
  • m1 = Σ αi vi · Hindi values
  • Δe1 = m1 WO; add to current e1
  • Updated e1 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q1 = updated e1 WQ (cross-attention)
  • q1 KT / √dk → softmax over English
  • 0.72
  • 0.04
  • 0.03
  • 0.09
  • 0.04
  • 0.08
  • c1 = Σ αi vi · English values
  • Cross update: Δe1 = c1 WO → add to e1 → MLP + residual; finish blocks
  • Step 1: vocabulary head → softmax → राघव · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e2 at राघव
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.35
  • राघव
  • 0.65
  • m2 = Σ αi vi · Hindi values
  • Δe2 = m2 WO; add to current e2
  • Updated e2 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q2 = updated e2 WQ (cross-attention)
  • q2 KT / √dk → softmax over English
  • 0.05
  • 0.03
  • 0.04
  • 0.10
  • 0.08
  • 0.70
  • c2 = Σ αi vi · English values
  • Cross update: Δe2 = c2 WO → add to e2 → MLP + residual; finish blocks
  • Step 2: vocabulary head → softmax → दिल्ली · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e3 at दिल्ली
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.15
  • राघव
  • 0.25
  • दिल्ली
  • 0.60
  • m3 = Σ αi vi · Hindi values
  • Δe3 = m3 WO; add to current e3
  • Updated e3 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q3 = updated e3 WQ (cross-attention)
  • q3 KT / √dk → softmax over English
  • 0.04
  • 0.04
  • 0.03
  • 0.09
  • 0.70
  • 0.10
  • c3 = Σ αi vi · English values
  • Cross update: Δe3 = c3 WO → add to e3 → MLP + residual; finish blocks
  • Step 3: vocabulary head → softmax → में · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e4 at में
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.10
  • राघव
  • 0.15
  • दिल्ली
  • 0.45
  • में
  • 0.30
  • m4 = Σ αi vi · Hindi values
  • Δe4 = m4 WO; add to current e4
  • Updated e4 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q4 = updated e4 WQ (cross-attention)
  • q4 KT / √dk → softmax over English
  • 0.08
  • 0.06
  • 0.04
  • 0.70
  • 0.05
  • 0.07
  • c4 = Σ αi vi · English values
  • Cross update: Δe4 = c4 WO → add to e4 → MLP + residual; finish blocks
  • Step 4: vocabulary head → softmax → स्कूल · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e5 at स्कूल
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.08
  • राघव
  • 0.12
  • दिल्ली
  • 0.20
  • में
  • 0.25
  • स्कूल
  • 0.35
  • m5 = Σ αi vi · Hindi values
  • Δe5 = m5 WO; add to current e5
  • Updated e5 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q5 = updated e5 WQ (cross-attention)
  • q5 KT / √dk → softmax over English
  • 0.04
  • 0.75
  • 0.04
  • 0.08
  • 0.04
  • 0.05
  • c5 = Σ αi vi · English values
  • Cross update: Δe5 = c5 WO → add to e5 → MLP + residual; finish blocks
  • Step 5: vocabulary head → softmax → जाता · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e6 at जाता
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.15
  • में
  • 0.15
  • स्कूल
  • 0.25
  • जाता
  • 0.30
  • m6 = Σ αi vi · Hindi values
  • Δe6 = m6 WO; add to current e6
  • Updated e6 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q6 = updated e6 WQ (cross-attention)
  • q6 KT / √dk → softmax over English
  • 0.10
  • 0.35
  • 0.05
  • 0.20
  • 0.15
  • 0.15
  • c6 = Σ αi vi · English values
  • Cross update: Δe6 = c6 WO → add to e6 → MLP + residual; finish blocks
  • Step 6: vocabulary head → softmax → है · append to Hindi prefix
  • 1 · Hindi self-attention
  • 2 · English cross-attention
  • Self query from e7 at है
  • Keys and values: supplied Hindi positions
  • Same six English encoder outputs
  • Source keys K and values V
  • [शुरू]
  • 0.05
  • राघव
  • 0.10
  • दिल्ली
  • 0.10
  • में
  • 0.15
  • स्कूल
  • 0.15
  • जाता
  • 0.20
  • है
  • 0.25
  • m7 = Σ αi vi · Hindi values
  • Δe7 = m7 WO; add to current e7
  • Updated e7 after self-attention
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • q7 = updated e7 WQ (cross-attention)
  • q7 KT / √dk → softmax over English
  • 0.16
  • 0.18
  • 0.16
  • 0.17
  • 0.16
  • 0.17
  • c7 = Σ αi vi · English values
  • Cross update: Δe7 = c7 WO → add to e7 → MLP + residual; finish blocks
  • Step 7: vocabulary head → softmax → [समाप्त] · stop

Authored teaching diagram

What can each attention operation read at this generation step?

Target self-attention uses only the supplied Hindi prefix, including the current last position. Its weighted message m is projected and added to that target embedding. A separate cross-attention W_Q then forms a query from the updated embedding. English encoder outputs supply the six source keys and values. Their weighted message c produces a second target update. The final vocabulary head predicts the next Hindi token. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Teaching note

What can each attention operation read at this generation step? Use the seven buttons or advance through the sequence. Read the Hindi tokens and purple weights first. Follow the message and residual to the updated embedding, then the arrow into the cross-attention query. Read the six English weights separately before the source update and token prediction. The sequence is [शुरू] → राघव → दिल्ली → में → स्कूल → जाता → है → [समाप्त]. To predict स्कूल, the supplied Hindi positions are [शुरू], राघव, दिल्ली and में; the last query is at position 4. The self-attention row has four weights while cross-attention still has six. At the first step, the sole self-attention weight is 1. Each target row and each source row is independently normalized. The two α labels are local weights over different positions, not shared parameters or probabilities over the output vocabulary. Self-attention and cross-attention have distinct W_Q, W_K, W_V and W_O. The left panel summarizes self-attention scoring, causal masking and softmax with illustrative weights. Values on the left come from target embeddings; values on the right come from encoder outputs. Input target embeddings include positions; normalization is omitted. No next token is predicted between the two attention updates. The final vocabulary head is used only after the MLP and remaining decoder blocks. Source encoder outputs and projected source keys/values can be reused. These numbers are authored examples, not measured alignments. Stop when [समाप्त] is chosen.

Frame overflows: split the content

04 · Cross-attention to the source

Fixed context and query-dependent context

ONE FIXED SOURCE VECTOR. SOURCE STATES RETAINED. e1 … e6 → Pool → c. e1 … e6 → K, V. c. predict y1. q1. lookup. c1. c. predict y2. q2. lookup. c2. c. predict y3. q3. lookup. c3ONE FIXED SOURCE VECTORSOURCE STATES RETAINEDe1 … e6 → Pool → ce1 … e6 → K, Vcpredict y1q1lookupc1cpredict y2q2lookupc2cpredict y3q3lookupc3

Source states stay fixed; the query-dependent message changes.

Read the diagram labels
  • ONE FIXED SOURCE VECTOR. SOURCE STATES RETAINED. e1 … e6 → Pool → c. e1 … e6 → K, V. c. predict y1. q1. lookup. c1. c. predict y2. q2. lookup. c2. c. predict y3. q3. lookup. c3
  • ONE FIXED SOURCE VECTOR
  • SOURCE STATES RETAINED
  • e1 … e6 → Pool → c
  • e1 … e6 → K, V
  • c
  • predict y1
  • q1
  • lookup
  • c1
  • c
  • predict y2
  • q2
  • lookup
  • c2
  • c
  • predict y3
  • q3
  • lookup
  • c3

Authored teaching diagram · Primary source

Do we re-encode the English sentence for each Hindi token?

No. We can reuse its encoder states. The changing target query produces a potentially different weighted source message at each target step.

Teaching note

Do we re-encode the English sentence for each Hindi token? Contrast the one-vector source interface with retaining all source states. Do not imply that fixed-context decoding has an unchanging target state.

Frame overflows: split the content

04 · Cross-attention to the source

One decoder block, two attention operations

Target embeddings + positions. Causal self-attention. Cross-attention. Feed-forward network. Q. Encoder states. K,V. target state → vocabulary headTarget embeddings + positionsCausal self-attentionCross-attentionFeed-forward networkQEncoder statesK,Vtarget state → vocabulary head

Residuals + normalization omitted for clarity.

Read the diagram labels
  • Target embeddings + positions. Causal self-attention. Cross-attention. Feed-forward network. Q. Encoder states. K,V. target state → vocabulary head
  • Target embeddings + positions
  • Causal self-attention
  • Cross-attention
  • Feed-forward network
  • Q
  • Encoder states
  • K,V
  • target state → vocabulary head

Authored teaching diagram · Primary source

Why does the decoder need both attention operations?

Causal self-attention supplies target-prefix context. Cross-attention retrieves from the complete source. The feed-forward sublayer then transforms each target position.

Teaching note

Why does the decoder need both attention operations? Trace the target stream down the block, stopping at the side connection from the source.

Frame overflows: split the content

05 · Attention patterns and applications

Part V · Attention patterns and applications

Compare what each attention operation can read.. Full, causal and cross-attention. Task predictions. Connect the attention pattern to the required output.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesCompare what each attention operation can read.Full, causal and cross-attentionTask predictionsConnect the attention pattern to the required output.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Compare what each attention operation can read.. Full, causal and cross-attention. Task predictions. Connect the attention pattern to the required output.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Compare what each attention operation can read.
  • Full, causal and cross-attention
  • Task predictions
  • Connect the attention pattern to the required output.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part V · Attention patterns and applications

Compare what each attention operation can read.

Teaching note

Part V · Attention patterns and applications Pause and establish the requested output before advancing.

Frame overflows: split the content

05 · Attention patterns and applications

Attention patterns and their applications

Three allowed-access matrices with concrete tasks. Encoder: six English query rows by six English key/value columns, full access, followed by a topic summary and classification head. Decoder-only continuation: three supplied English prefix positions with a causal triangle, then vocabulary prediction of school. Cross-attention: two Hindi target positions, start and Raghav, querying all six English positions to help predict Delhi. Filled cells are permission, not numerical attention weights. Detailed word-labeled matrices follow.English example: Raghav goes to school in DelhiEncoder self-attentionClassify the topic: EDUCATIONEnglish → EnglishQ, K, V: Englishp(label | complete input)Decoder self-attentionContinue: Raghav goes to …Prefix → prefixQ, K, V: prefixp(school | Raghav goes to)Cross-attentionTranslate: [शुरू] राघव …Hindi → EnglishQ: Hindi · K,V: Englishp(दिल्ली | English, Hindi prefix)

Filled cells show allowed connections. The task probability comes from the final head.

Read the diagram labels
  • Three allowed-access matrices with concrete tasks. Encoder: six English query rows by six English key/value columns, full access, followed by a topic summary and classification head. Decoder-only continuation: three supplied English prefix positions with a causal triangle, then vocabulary prediction of school. Cross-attention: two Hindi target positions, start and Raghav, querying all six English positions to help predict Delhi. Filled cells are permission, not numerical attention weights. Detailed word-labeled matrices follow.
  • English example: Raghav goes to school in Delhi
  • Encoder self-attention
  • Classify the topic: EDUCATION
  • English → English
  • Q, K, V: English
  • p(label | complete input)
  • Decoder self-attention
  • Continue: Raghav goes to …
  • Prefix → prefix
  • Q, K, V: prefix
  • p(school | Raghav goes to)
  • Cross-attention
  • Translate: [शुरू] राघव …
  • Hindi → English
  • Q: Hindi · K,V: English
  • p(दिल्ली | English, Hindi prefix)

Authored teaching diagram

How does the attention pattern relate to the task?

An encoder can read the complete English input to build representations for a sentence classifier. A causal decoder reads the supplied prefix to support next-token prediction. Cross-attention lets Hindi target positions consult the English source while translating. The following slides label every row and column with the actual words. These are three illustrative task settings, not exclusive uses of each architecture. The encoder example has six positions, decoder continuation has three supplied positions, and cross-attention has two Hindi positions by six English positions. Filled cells are permitted access, not measured attention weights or probabilities. The decoder can use previous words and the current supplied token; it cannot see future target words. The encoder returns one vector per input position, with pooling or a selected token used to obtain a summary when the task needs one.

Teaching note

How does the attention pattern relate to the task? Compare the full English square, the causal prefix triangle and the Hindi-to-English rectangle. Read the task probability below each one. These are three illustrative task settings, not exclusive uses of each architecture. The encoder example has six positions, decoder continuation has three supplied positions, and cross-attention has two Hindi positions by six English positions. Filled cells are permitted access, not measured attention weights or probabilities. The decoder can use previous words and the current supplied token; it cannot see future target words. The encoder returns one vector per input position, with pooling or a selected token used to obtain a summary when the task needs one.

Frame overflows: split the content

05 · Attention patterns and applications

Encoder attention for sentence classification

English encoder self-attention with actual word-labeled query rows and key/value columns. Every source position can read all six input positions. The highlighted school row can attend to Raghav, goes, to, school, in and Delhi. The encoder returns six contextual embeddings. For this illustrative sentence-topic classification task, mean pool the outputs, then use a classifier and softmax for p(y equals EDUCATION given the complete sentence x). Pooling is a task choice; an encoder does not inherently output just one summary.Input x: Raghav goes to school in DelhiColumns: English keys / valuesRaghavgoestoschoolinDelhiRaghavgoestoschoolinDelhiRows: English queries · all six words are visibleEncoder → 6 updated embeddingsMean pool → one sentence vectorTopic head → logits → softmaxp(y = EDUCATION | x)The encoder keeps all positions. Pooling gives the summary used by this task.

Example task: classify the sentence’s topic using all supplied words.

Read the diagram labels
  • English encoder self-attention with actual word-labeled query rows and key/value columns. Every source position can read all six input positions. The highlighted school row can attend to Raghav, goes, to, school, in and Delhi. The encoder returns six contextual embeddings. For this illustrative sentence-topic classification task, mean pool the outputs, then use a classifier and softmax for p(y equals EDUCATION given the complete sentence x). Pooling is a task choice; an encoder does not inherently output just one summary.
  • Input x: Raghav goes to school in Delhi
  • Columns: English keys / values
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • Rows: English queries · all six words are visible
  • Encoder → 6 updated embeddings
  • Mean pool → one sentence vector
  • Topic head → logits → softmax
  • p(y = EDUCATION | x)
  • The encoder keeps all positions. Pooling gives the summary used by this task.

Authored teaching diagram

Does the encoder itself reduce six words to one summary?

No. It returns six updated token embeddings. In this example we mean-pool them into one vector, then apply a topic classifier and softmax to obtain p(y | x), where x is the complete sentence and y is a topic label. EDUCATION is an illustrative label. Rows are positions whose queries receive messages; columns are positions supplying keys and values. Both axes come from the same English input. CLS selection is another summary choice, and token labeling could instead use every output embedding. This slide uses pooling to keep the six-token attention matrix literal. No trained prediction or numerical attention values are claimed.

Teaching note

Does the encoder itself reduce six words to one summary? Read the highlighted school row against all six column words. Follow all encoder outputs into mean pooling and the topic head. Rows are positions whose queries receive messages; columns are positions supplying keys and values. Both axes come from the same English input. CLS selection is another summary choice, and token labeling could instead use every output embedding. This slide uses pooling to keep the six-token attention matrix literal. No trained prediction or numerical attention values are claimed.

Frame overflows: split the content

05 · Attention patterns and applications

Causal attention for next-token prediction

Decoder causal self-attention over the supplied prefix Raghav goes to. Rows and columns are Raghav, goes and to. Raghav reads itself; goes reads Raghav and itself; to reads all three supplied positions. The future school token is absent. Finish the decoder blocks and use the final to embedding e3 for vocabulary logits and softmax, modeling p(school given Raghav goes to). Self-attention updates embeddings; the final vocabulary head produces the next-token distribution.Supplied prefix: Raghav goes to → next token: schoolColumns: supplied prefix keys / valuesRaghavgoestoRaghavgoestoCausal attention updates embeddingsFinish blocks → read last e3Vocabulary logits → softmaxThe “to” row can read Raghav, goes and to.p(school | Raghav goes to)Previous words are available; future words cannot be read.

The current and previous supplied words are visible; future words are not.

Read the diagram labels
  • Decoder causal self-attention over the supplied prefix Raghav goes to. Rows and columns are Raghav, goes and to. Raghav reads itself; goes reads Raghav and itself; to reads all three supplied positions. The future school token is absent. Finish the decoder blocks and use the final to embedding e3 for vocabulary logits and softmax, modeling p(school given Raghav goes to). Self-attention updates embeddings; the final vocabulary head produces the next-token distribution.
  • Supplied prefix: Raghav goes to → next token: school
  • Columns: supplied prefix keys / values
  • Raghav
  • goes
  • to
  • Raghav
  • goes
  • to
  • Causal attention updates embeddings
  • Finish blocks → read last e3
  • Vocabulary logits → softmax
  • The “to” row can read Raghav, goes and to.
  • p(school | Raghav goes to)
  • Previous words are available; future words cannot be read.

Authored teaching diagram

Can the “to” position attend to the previous word “goes”?

Yes. The last supplied position to can read Raghav, goes and itself. It cannot read school because school has not been supplied. After the decoder blocks, the vocabulary head reads the final e₃ and models p(school | Raghav goes to). Causal attention alone updates embeddings rather than emitting a word. More generally, next-token language modeling learns p(xₜ₊₁ | x₁,…,xₜ). During training, later tokens may be present in the batch tensor but the causal mask hides them from earlier queries. At generation time those future tokens are absent. The diagonal is allowed because the token at the receiving position is an input used to predict the following token.

Teaching note

Can the “to” position attend to the previous word “goes”? Read the three matrix rows in order. Trace the highlighted bottom row to all three input columns, then the last embedding to the vocabulary distribution. More generally, next-token language modeling learns p(xₜ₊₁ | x₁,…,xₜ). During training, later tokens may be present in the batch tensor but the causal mask hides them from earlier queries. At generation time those future tokens are absent. The diagonal is allowed because the token at the receiving position is an input used to predict the following token.

Frame overflows: split the content

05 · Attention patterns and applications

Cross-attention between Hindi and English

Cross-attention has Hindi query rows start and Raghav, and English key/value columns Raghav, goes, to, school, in, Delhi. Both target positions can read the complete English source. Queries use target embeddings after causal self-attention; keys and values use encoder outputs. The retrieved source message updates the target stream before the remaining decoder computation and vocabulary prediction. The model predicts p(Delhi given the English input x and supplied Hindi prefix start Raghav). This full rectangular access pattern is not a future-target leak: English is supplied, whereas future Hindi tokens are unavailable.English x: Raghav goes to school in DelhiColumns: English keys / valuesRaghavgoestoschoolinDelhi[शुरू]राघवRows: Hindi queries after causal self-attentionHindi prefix: [शुरू] राघवQ: Hindi · K and V: EnglishSource message → target updateBoth Hindi positions may read all six English positions.Translation task: p(दिल्ली | English x, Hindi prefix [शुरू] राघव)

Rows are Hindi target positions; columns are English source positions.

Read the diagram labels
  • Cross-attention has Hindi query rows start and Raghav, and English key/value columns Raghav, goes, to, school, in, Delhi. Both target positions can read the complete English source. Queries use target embeddings after causal self-attention; keys and values use encoder outputs. The retrieved source message updates the target stream before the remaining decoder computation and vocabulary prediction. The model predicts p(Delhi given the English input x and supplied Hindi prefix start Raghav). This full rectangular access pattern is not a future-target leak: English is supplied, whereas future Hindi tokens are unavailable.
  • English x: Raghav goes to school in Delhi
  • Columns: English keys / values
  • Raghav
  • goes
  • to
  • school
  • in
  • Delhi
  • [शुरू]
  • राघव
  • Rows: Hindi queries after causal self-attention
  • Hindi prefix: [शुरू] राघव
  • Q: Hindi · K and V: English
  • Source message → target update
  • Both Hindi positions may read all six English positions.
  • Translation task: p(दिल्ली | English x, Hindi prefix [शुरू] राघव)

Authored teaching diagram

Why is the cross-attention matrix full even though the decoder is causal?

The complete English source is already supplied, so each Hindi query may read all six source positions. Causality still applies within the Hindi stream. To predict दिल्ली, use the current embedding at राघव after Hindi self-attention, then query English keys and combine English values. The translation model represents p(yₜ | x, y₁,…,yₜ₋₁), with x the complete English sentence and the supplied Hindi prefix beginning with [शुरू]. The probability of a Hindi output token is produced by the final vocabulary head after target updates, not directly by the attention matrix. Attention weights normalize over English positions; vocabulary probabilities normalize over possible Hindi tokens. Future Hindi tokens are not available to these queries.

Teaching note

Why is the cross-attention matrix full even though the decoder is causal? Read the highlighted राघव row across the six English column words. Contrast the two languages on the axes and read the conditional probability aloud. The translation model represents p(yₜ | x, y₁,…,yₜ₋₁), with x the complete English sentence and the supplied Hindi prefix beginning with [शुरू]. The probability of a Hindi output token is produced by the final vocabulary head after target updates, not directly by the attention matrix. Attention weights normalize over English positions; vocabulary probabilities normalize over possible Hindi tokens. Future Hindi tokens are not available to these queries.

Frame overflows: split the content

06 · Comparing model architectures

Part VI · Comparing model architectures

Compare the architectures used for the example tasks.. Inputs and outputs. Architecture choice. Compare attention patterns and task heads.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model familiesCompare the architectures used for the example tasks.Inputs and outputsArchitecture choiceCompare attention patterns and task heads.01 · Generation02 · Input labels03 · Translation04 · Cross-attention05 · Patterns06 · Model families
Read the diagram labels
  • Compare the architectures used for the example tasks.. Inputs and outputs. Architecture choice. Compare attention patterns and task heads.. 01 · Generation. 02 · Input labels. 03 · Translation. 04 · Cross-attention. 05 · Patterns. 06 · Model families
  • Compare the architectures used for the example tasks.
  • Inputs and outputs
  • Architecture choice
  • Compare attention patterns and task heads.
  • 01 · Generation
  • 02 · Input labels
  • 03 · Translation
  • 04 · Cross-attention
  • 05 · Patterns
  • 06 · Model families

Authored teaching diagram

Part VI · Comparing model architectures

Compare the architectures used for the example tasks.

Teaching note

Part VI · Comparing model architectures Pause and establish the requested output before advancing.

Frame overflows: split the content

06 · Comparing model architectures

Encoder, decoder-only and encoder-decoder models

ENCODER. Encode the input. DECODER. Predict the next token. ENCODER-DECODER. Generate from a source. . . . . . . . Encoder. . . . . . . one vector per token. labels · retrieval · readouts. . . . . Decoder. next. append → repeat. chat · code · continuation. Encoder. . . . target prefix. source K,V. Decoder. next. source + target prefix. translation · summarizationENCODEREncode the inputDECODERPredict the next tokenENCODER-DECODERGenerate from a sourceEncoderone vector per tokenlabels · retrieval · readoutsDecodernextappend → repeatchat · code · continuationEncodertarget prefixsource K,VDecodernextsource + target prefixtranslation · summarization

Compare the input, attention connections and output of each architecture.

Read the diagram labels
  • ENCODER. Encode the input. DECODER. Predict the next token. ENCODER-DECODER. Generate from a source. . . . . . . . Encoder. . . . . . . one vector per token. labels · retrieval · readouts. . . . . Decoder. next. append → repeat. chat · code · continuation. Encoder. . . . target prefix. source K,V. Decoder. next. source + target prefix. translation · summarization
  • ENCODER
  • Encode the input
  • DECODER
  • Predict the next token
  • ENCODER-DECODER
  • Generate from a source
  • Encoder
  • one vector per token
  • labels · retrieval · readouts
  • Decoder
  • next
  • append → repeat
  • chat · code · continuation
  • Encoder
  • target prefix
  • source K,V
  • Decoder
  • next
  • source + target prefix
  • translation · summarization

Authored teaching diagram · Primary source

What did we add to connect understanding a source with generating a target?

A source encoder and a pathway from its representations into the target decoder. Our final arrangement uses cross-attention for that source pathway.

Teaching note

What did we add to connect understanding a source with generating a target? Compare all three architectures only now, after deriving them from the tasks.

Frame overflows: split the content

06 · Comparing model architectures

Choosing an architecture for a task

One label for the input. Encoder + readout. sentiment. One label per token. Encoder + shared head. NER. Continue a prefix. Decoder-only. text generation. Turn source X into target Y. Encoder-decoder. English → HindiOne label for the inputEncoder + readoutsentimentOne label per tokenEncoder + shared headNERContinue a prefixDecoder-onlytext generationTurn source X into target YEncoder-decoderEnglish → Hindi

These are common choices; other architectures can perform the same tasks.

Read the diagram labels
  • One label for the input. Encoder + readout. sentiment. One label per token. Encoder + shared head. NER. Continue a prefix. Decoder-only. text generation. Turn source X into target Y. Encoder-decoder. English → Hindi
  • One label for the input
  • Encoder + readout
  • sentiment
  • One label per token
  • Encoder + shared head
  • NER
  • Continue a prefix
  • Decoder-only
  • text generation
  • Turn source X into target Y
  • Encoder-decoder
  • English → Hindi

Authored teaching diagram · Primary source

Can a decoder-only model also produce a class label or translation?

Yes. It can generate labels or translations from a suitable prompt and training setup. These rows illustrate useful arrangements rather than exclusive capabilities.

Teaching note

Can a decoder-only model also produce a class label or translation? Ask for a task and have students trace its desired output shape.

Frame overflows: split the content

06 · Comparing model architectures

Choosing an architecture

Do I need to generate new tokens?. NO. YES. Encoder. representations. labels / embeddings. Separate source pathway?. NO. YES. Decoder-only. Encoder-decoderDo I need to generate new tokens?NOYESEncoderrepresentationslabels / embeddingsSeparate source pathway?NOYESDecoder-onlyEncoder-decoder

The output type and source pathway guide the choice of architecture.

Read the diagram labels
  • Do I need to generate new tokens?. NO. YES. Encoder. representations. labels / embeddings. Separate source pathway?. NO. YES. Decoder-only. Encoder-decoder
  • Do I need to generate new tokens?
  • NO
  • YES
  • Encoder
  • representations
  • labels / embeddings
  • Separate source pathway?
  • NO
  • YES
  • Decoder-only
  • Encoder-decoder

Authored teaching diagram · Primary source

Does having source text force us to use an encoder-decoder?

No. This branch asks whether the chosen design has a separate source pathway. A decoder-only model can instead consume source text within its prefix.

Teaching note

Does having source text force us to use an encoder-decoder? Ask whether the task needs new tokens. For a generated output, ask whether the design uses a separate source pathway. Representations can support labels, embeddings or token outputs.

Frame overflows: split the content

06 · Comparing model architectures

Review: attention, representations and training

01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?. Next lecture: what if our tokens are not words?01WHICH POSITIONS CAN ATTEND TO EACH OTHER?02WHICH EMBEDDINGS DOES THE TASK USE?03WHAT IS THE TRAINING TARGET?Next lecture: what if our tokens are not words?
Read the diagram labels
  • 01. WHICH POSITIONS CAN ATTEND TO EACH OTHER?. 02. WHICH EMBEDDINGS DOES THE TASK USE?. 03. WHAT IS THE TRAINING TARGET?. Next lecture: what if our tokens are not words?
  • 01
  • WHICH POSITIONS CAN ATTEND TO EACH OTHER?
  • 02
  • WHICH EMBEDDINGS DOES THE TASK USE?
  • 03
  • WHAT IS THE TRAINING TARGET?
  • Next lecture: what if our tokens are not words?

Authored teaching diagram

For NER, classification and translation: who reads whom, what do we read, and what trains it?

NER uses whole-input attention, all token states and token labels. Classification uses a whole-input readout and one sentence label. Translation uses source encoding, causal target states, cross-attention and next-target-token supervision.

Teaching note

For NER, classification and translation: who reads whom, what do we read, and what trains it? Let students reconstruct each example, then reveal the teaser about tokens that are not words.

Frame overflows: split the content

Appendix · Optional worked examples

Appendix: worked examples

Tensor shapes, training and a numerical attention example.. Shapes and training. Numerical examples. Optional details for discussion and practice.Tensor shapes, training and a numerical attention example.Shapes and trainingNumerical examplesOptional details for discussion and practice.
Read the diagram labels
  • Tensor shapes, training and a numerical attention example.. Shapes and training. Numerical examples. Optional details for discussion and practice.
  • Tensor shapes, training and a numerical attention example.
  • Shapes and training
  • Numerical examples
  • Optional details for discussion and practice.

Authored teaching diagram

Appendix: worked examples

Tensor shapes, training and a numerical attention example.

Teaching note

Appendix: worked examples Pause and establish the requested output before advancing.

Frame overflows: split the content

Appendix · Optional worked examples

Tensor shapes in cross-attention

E ∈ ℝn × d · source states. D ∈ ℝm × d · target states. d = model width. . K ∈ ℝn × dk. . V ∈ ℝn × dv. . Q ∈ ℝm × dk. . A = softmax(QKT / √dk) ∈ ℝm × n. . C = AV ∈ ℝm × dv. Example: m = 4, n = 6 → A is 4 × 6E ∈ ℝn × d · source statesD ∈ ℝm × d · target statesd = model widthK ∈ ℝn × dkV ∈ ℝn × dvQ ∈ ℝm × dkA = softmax(QKT / √dk) ∈ ℝm × nC = AV ∈ ℝm × dvExample: m = 4, n = 6 → A is 4 × 6

Rows are target queries; columns are source positions.

Read the diagram labels
  • E ∈ ℝn × d · source states. D ∈ ℝm × d · target states. d = model width. . K ∈ ℝn × dk. . V ∈ ℝn × dv. . Q ∈ ℝm × dk. . A = softmax(QKT / √dk) ∈ ℝm × n. . C = AV ∈ ℝm × dv. Example: m = 4, n = 6 → A is 4 × 6
  • E ∈ ℝn × d · source states
  • D ∈ ℝm × d · target states
  • d = model width
  • K ∈ ℝn × dk
  • V ∈ ℝn × dv
  • Q ∈ ℝm × dk
  • A = softmax(QKT / √dk) ∈ ℝm × n
  • C = AV ∈ ℝm × dv
  • Example: m = 4, n = 6 → A is 4 × 6

Authored teaching diagram · Primary source

Must source and target sequence lengths match?

No. For m target queries and n source positions, A has shape m × n. Each row sums to one. C = AV has shape m × dᵥ per head. With m = 4 and n = 6, the weight matrix is 4 × 6. This diagram uses a common model width for simplicity; learned projections can also map different source and target widths into a shared query/key width.

Teaching note

Must source and target sequence lengths match? Multiply the inner dimensions aloud. At generation time highlight one query row. This diagram uses a common model width for simplicity; learned projections can also map different source and target widths into a shared query/key width.

Frame overflows: split the content

Appendix · Optional worked examples

Conditioning with a continuous source prefix

c = Pool(source states). u = c WP · decoder width. u. [शुरू]. राघव. दिल्ली. continuous prefix. ordinary target embeddings + positions. Causal Transformer → next Hindi token. [CONTEXT] denotes an injected vector, not a new vocabulary word.c = Pool(source states)u = c WP · decoder widthu[शुरू]राघवदिल्लीcontinuous prefixordinary target embeddings + positionsCausal Transformer → next Hindi token[CONTEXT] denotes an injected vector, not a new vocabulary word.

The target can read the source prefix through causal self-attention.

Read the diagram labels
  • c = Pool(source states). u = c WP · decoder width. u. [शुरू]. राघव. दिल्ली. continuous prefix. ordinary target embeddings + positions. Causal Transformer → next Hindi token. [CONTEXT] denotes an injected vector, not a new vocabulary word.
  • c = Pool(source states)
  • u = c WP · decoder width
  • u
  • [शुरू]
  • राघव
  • दिल्ली
  • continuous prefix
  • ordinary target embeddings + positions
  • Causal Transformer → next Hindi token
  • [CONTEXT] denotes an injected vector, not a new vocabulary word.

Authored teaching diagram · Primary source

Is [CONTEXT] a word the decoder needs to predict?

No. Here it labels an injected continuous input vector u = c W_P. Prepend it to the positioned target embeddings, start with [शुरू], and read the last target position to predict the next word. Give the prefix slot a consistent position, and train this construction end to end on translation targets. The initial u is source-dependent and fixed across decoding steps; target hidden states continue to depend on the growing target prefix. The prefix slot itself has no language-model target in this construction.

Teaching note

Is [CONTEXT] a word the decoder needs to predict? Trace c through the width projection to the prefix slot, then through the ordinary causal stack. Give the prefix slot a consistent position, and train this construction end to end on translation targets. The initial u is source-dependent and fixed across decoding steps; target hidden states continue to depend on the growing target prefix. The prefix slot itself has no language-model target in this construction.

Frame overflows: split the content

Appendix · Optional worked examples

Next-token prediction during training

Source: Raghav goes to school in Delhi.. Decoder input. [शुरू]. राघव. दिल्ली. में. स्कूल. जाता. है. राघव. दिल्ली. में. स्कूल. जाता. है. [समाप्त]. Training target. L = −Σt log p(yt | y<t, x)Source: Raghav goes to school in Delhi.Decoder input[शुरू]राघवदिल्लीमेंस्कूलजाताहैराघवदिल्लीमेंस्कूलजाताहै[समाप्त]Training targetL = −Σt log p(yt | y<t, x)

Shift the target sequence by one position.

Read the diagram labels
  • Source: Raghav goes to school in Delhi.. Decoder input. [शुरू]. राघव. दिल्ली. में. स्कूल. जाता. है. राघव. दिल्ली. में. स्कूल. जाता. है. [समाप्त]. Training target. L = −Σt log p(yt | y<t, x)
  • Source: Raghav goes to school in Delhi.
  • Decoder input
  • [शुरू]
  • राघव
  • दिल्ली
  • में
  • स्कूल
  • जाता
  • है
  • राघव
  • दिल्ली
  • में
  • स्कूल
  • जाता
  • है
  • [समाप्त]
  • Training target
  • L = −Σt log p(yt | y<t, x)

Authored teaching diagram · Primary source

Can the decoder see the target it is currently supposed to predict?

Not at the prediction position: its input is the previous target token and its self-attention is causal. The source remains available through cross-attention. Teacher forcing supplies the known target prefix during training. With causal masking, all these next-token losses can be computed in one parallel pass.

Teaching note

Can the decoder see the target it is currently supposed to predict? Trace each input column to the different target below it, including [शुरू] and [समाप्त]. Teacher forcing supplies the known target prefix during training. With causal masking, all these next-token losses can be computed in one parallel pass.

Frame overflows: split the content

Appendix · Optional worked examples

Training versus generation

TRAINING. GENERATION. correct previous target token. previous model prediction. predict next token. predict next token. Known target prefixes supplied. Append prediction; repeatTRAININGGENERATIONcorrect previous target tokenprevious model predictionpredict next tokenpredict next tokenKnown target prefixes suppliedAppend prediction; repeat

The next-token objective is the same; prefix construction differs.

Read the diagram labels
  • TRAINING. GENERATION. correct previous target token. previous model prediction. predict next token. predict next token. Known target prefixes supplied. Append prediction; repeat
  • TRAINING
  • GENERATION
  • correct previous target token
  • previous model prediction
  • predict next token
  • predict next token
  • Known target prefixes supplied
  • Append prediction; repeat

Authored teaching diagram · Primary source

What is fed back at inference time?

The model’s selected prediction. In teacher-forced training we supply the observed previous target token instead.

Teaching note

What is fed back at inference time? Contrast how each prefix is built. Avoid decoding algorithms here.

Frame overflows: split the content

Appendix · Optional worked examples

Attention head ≠ classification head

Attention heads. Classification head. mix information. map states to label scores. Their counts are independent.Attention headsClassification headmix informationmap states to label scoresTheir counts are independent.

Attention mixes states; a task head predicts labels.

Read the diagram labels
  • Attention heads. Classification head. mix information. map states to label scores. Their counts are independent.
  • Attention heads
  • Classification head
  • mix information
  • map states to label scores
  • Their counts are independent.

Authored teaching diagram · Primary source

Do four entity labels imply four attention heads?

No. The number of attention heads and the number of classes are independent. They name different computations.

Teaching note

Do four entity labels imply four attention heads? Locate the attention heads inside the Transformer and the classifier after it.

Frame overflows: split the content

Appendix · Optional worked examples

A numerical cross-attention example

One query · four source keys · chosen scaled scores. Raghav. 0.2. 0.110. goes. 0.1. 0.100. school. 0.3. 0.122. Delhi. 2.0. 0.668. softmax across source positions → ct = Σi αti viOne query · four source keys · chosen scaled scoresRaghav0.20.110goes0.10.100school0.30.122Delhi2.00.668softmax across source positions → ct = Σi αti vi

Chosen scaled scores; probabilities recomputed with softmax.

Read the diagram labels
  • One query · four source keys · chosen scaled scores. Raghav. 0.2. 0.110. goes. 0.1. 0.100. school. 0.3. 0.122. Delhi. 2.0. 0.668. softmax across source positions → ct = Σi αti vi
  • One query · four source keys · chosen scaled scores
  • Raghav
  • 0.2
  • 0.110
  • goes
  • 0.1
  • 0.100
  • school
  • 0.3
  • 0.122
  • Delhi
  • 2.0
  • 0.668
  • softmax across source positions → ct = Σi αti vi

Authored teaching diagram · Primary source

Which source position gets the most weight?

Delhi. Softmax of scaled scores [0.2, 0.1, 0.3, 2.0] is approximately [0.110, 0.100, 0.122, 0.668]. These weights mix the four source value vectors. This deliberately reduced four-key example is separate from the six-word translation diagram. Scores already include division by √dₖ. No learned model or actual alignment is being claimed.

Teaching note

Which source position gets the most weight? Normalize the four exponentials, then point back to the familiar weighted sum. This deliberately reduced four-key example is separate from the six-word translation diagram. Scores already include division by √dₖ. No learned model or actual alignment is being claimed.

Frame overflows: split the content

Appendix · Optional worked examples

Computing the weighted sum of values

source. weight α. value vector V. Raghav. 0.110. [1, 0]. goes. 0.100. [0, 1]. school. 0.122. [1, 1]. Delhi. 0.668. [2, 1]. ct = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1]. ct ≈ [1.57, 0.89]sourceweight αvalue vector VRaghav0.110[1, 0]goes0.100[0, 1]school0.122[1, 1]Delhi0.668[2, 1]ct = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1]ct ≈ [1.57, 0.89]

The weights mix values into cₜ ≈ [1.57, 0.89].

Read the diagram labels
  • source. weight α. value vector V. Raghav. 0.110. [1, 0]. goes. 0.100. [0, 1]. school. 0.122. [1, 1]. Delhi. 0.668. [2, 1]. ct = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1]. ct ≈ [1.57, 0.89]
  • source
  • weight α
  • value vector V
  • Raghav
  • 0.110
  • [1, 0]
  • goes
  • 0.100
  • [0, 1]
  • school
  • 0.122
  • [1, 1]
  • Delhi
  • 0.668
  • [2, 1]
  • ct = 0.110[1,0] + 0.100[0,1] + 0.122[1,1] + 0.668[2,1]
  • ct ≈ [1.57, 0.89]

Authored teaching diagram · Primary source

What does cross-attention return after softmax?

A weighted sum of value vectors. With the four values shown, cₜ is approximately [1.57, 0.89]. It is a feature vector, not a class label or a word probability distribution. The exact softmax weights give cₜ = [1.567881402964, 0.889620530623]. The slide rounds the weights to three decimals and the result to two. Both agree at the displayed result precision. Values and scores are chosen for this calculation.

Teaching note

What does cross-attention return after softmax? Multiply each value by its weight, then add coordinate by coordinate. The Delhi value contributes most in this chosen example. The exact softmax weights give cₜ = [1.567881402964, 0.889620530623]. The slide rounds the weights to three decimals and the result to two. Both agree at the displayed result precision. Values and scores are chosen for this calculation.

Frame overflows: split the content

Lecture overview

Read or present

P: presentation / reading. Arrow keys or Space: advance in presentation. O: overview. Q: question. A: answer and follow-up questions. S: teaching notes. M: model-family map. N: next frame even when a control has focus. Home / End: first / last frame. Escape closes a dialog first, then exits presentation. Phone reading includes expandable diagram labels.

Presentation controls hide after two seconds of inactivity. Move the pointer, tap, or press a key to bring them back. Open menus and keyboard-focused controls stay visible.

Each next frame is a progressive teaching build. All frames appear in the reading version. The complete HTML works offline.