Notation used on this pagesymbols, meanings, shapes
    01

    The next token now depends on two sequences: the source and the answer so far.

    Translate one phrase

    Our earlier model continued a single sequence. A translation model has a complete source phrase to read and a new sequence to write. We will use the river bank and its French translation la rive. The task is still next-token prediction: after generating la, what should come next? The model can use both the English source and the French prefix. We will follow those two inputs separately until cross-attention brings them together.

    English, already known

    the river bank

    Two French words for bank

    la banque (money)  or  la rive (river)?

    Only the English sentence can settle it: river rules out banque. The decoder has written la and must choose the next French token while reading the English source. Observed: la rive.

    <bos> starts the answer. <eos> marks the end of each sequence.

    The source has four tokens including its end marker. To predict rive, the decoder can read only <bos> la.

    Every representation has shape $1\times3$.

    • Separate learned embedding tables for source and target tokens.
    • A position vector is added to each token embedding. The width stays three.
    • One encoder self-attention block, one masked decoder self-attention block, and one cross-attention block.

    The coordinates are unnamed learned numbers. This toy omits feed-forward sublayers and LayerNorm, covered in Part 3.

    The checkpoint was fitted to two short training pairs, including this river example. It lets us inspect a working forward pass and one further learning step; it does not demonstrate general translation ability. We keep the model small so that every row, weight, and output probability can be displayed.

    02

    Each English token can read all four source positions.

    The encoder reads the whole source

    The source is available before we write the translation, so the encoder needs no left-to-right mask. Its bank row can read river, and its the row can read bank. The calculation is the self-attention we already know: form queries, keys, and values from the same input rows; mix the values; project the message; and add the update to the input. Here the superscript src means a source input row, and enc means its representation after the encoder.

    Source: the river bank <eos>. Work the row for bank, source position 3.

    source_rows = src_embed(source_ids) + src_pos[:4]

    Both operands have three coordinates, so their sum still has three. Stack the four source rows into a $4\times3$ matrix.

    All four source input rows

    Each row is a token lookup plus the vector for its own position. The source embedding and position tables are learned parameters; the summed rows are recomputed for this input.

    Every source query can read every source value.

    Q, K, V = source_rows @ W_Q, source_rows @ W_K, source_rows @ W_V
    A = (Q @ K.T / math.sqrt(3)).softmax(dim=-1)

    For source position $j$: $\ve{e_j^{\mathrm{src}}}+\vd{\Delta e_j^{\mathrm{enc}}}=\vp{e_j^{\mathrm{enc}}}$.

    The encoder keeps one row per source position. It does not compress the whole phrase into a single vector.

    These four encoded rows will be the decoder's source of information. With the source and model parameters fixed, they remain the same as the French prefix grows. Cross-attention can nevertheless read a different mixture of them for each new decoder query.

    03

    Before reading English, a decoder row first reads the available French prefix.

    The decoder reads the answer so far

    Suppose the French prefix is <bos> la. To predict its next token, the decoder uses the row at la, the last known target position. We start from the target embedding table and target position vectors. Masked self-attention then lets la read itself and the start marker. The English source enters in the next block, cross-attention. We label the target input row $e_i^{\mathrm{tgt}}$ and the row after this first decoder block $e_i^{\mathrm{self}}$.

    Source stays the river bank <eos>. Target prefix: <bos> la.

    target_rows = tgt_embed(prefix_ids) + tgt_pos[:len(prefix_ids)]

    The row at la will predict the next French token. The unknown token cannot supply its own query.

    Target row $i$ reads only target positions $j\le i$.

    scores = Q @ K.T / math.sqrt(3)
    A = scores.masked_fill(future, -torch.inf).softmax(dim=-1)

    At la: $\ve{e_2^{\mathrm{tgt}}}+\vd{\Delta e_2^{\mathrm{self}}}=\vp{e_2^{\mathrm{self}}}$.

    This updated French-prefix row supplies the cross-attention query.

    The source encoder and masked target block have different jobs. The encoder contextualises the English phrase; target self-attention contextualises the French prefix. Cross-attention gives the decoder a way to use the first while writing the second.

    04

    The query comes from French. The keys and values come from encoded English.

    The decoder asks the source

    At the French position la, the decoder needs information that helps predict the next French token. We can describe that need as a question: which parts of the English phrase would help complete the translation? The learned query is a numerical projection of the decoder row, not an English sentence. Each encoded source row supplies a key for matching that query and a value carrying information to return. These projections have their own learned matrices; they are different from the matrices used inside the encoder and target self-attention blocks.

    $\vq{q_2^{\mathrm{cross}}}=\vp{e_2^{\mathrm{self}}}W_Q^{\mathrm{cross}}$

    The French la row retrieves source information to predict its next token.

    At French la: $\vq{q_2^{\mathrm{cross}}}=\vp{e_2^{\mathrm{self}}}W_Q^{\mathrm{cross}}$.

    Q = decoder_self @ W_Q_cross

    Source: the river bank <eos>. Both projections read the same encoded source rows.

    Keys match; values carry information. Same width, different vectors.

    K = encoded_source @ W_K_cross
    V = encoded_source @ W_V_cross

    One decoder query gives four scores, one per source position.

    Every score uses the same query, but a different source key. The table rounds its display to three decimals; the model uses the original values throughout.

    $s_{2j}=\dfrac{\vq{q_2^{\mathrm{cross}}}\cdot\vk{k_j^{\mathrm{cross}}}}{\sqrt{d_k}},\qquad d_k=3$

    Wider dot products tend to have a larger spread. Dividing by $\sqrt{d_k}$ helps keep softmax from becoming too sharp just because the vectors are wide.

    Decoder row: la. Softmax runs across the four English source positions.

    $\alpha_{2j}=\dfrac{\exp(s_{2j})}{\sum_{r=1}^{4}\exp(s_{2r})},\qquad\sum_{j=1}^{4}\alpha_{2j}=1$

    The largest weight in this trained toy's la row goes to the source position river. That weight is a contribution coefficient for a contextual value vector. It should not be read as a word-alignment label or a complete explanation of the model's prediction. Every encoded source row has already mixed information from the whole English phrase.

    05

    The reading weights bring an English-source message into the French decoder row.

    Source information changes the prediction

    We have chosen how much to read from each encoded source value. Multiplying each value by its weight and adding produces a message $m_2$ for the decoder's la position. The output projection turns that message into a contextual update, $\Delta e_2^{\mathrm{cross}}$, which we add to the decoder row. The resulting representation includes both the French prefix and information retrieved from the English source. The vocabulary head then predicts the next French token.

    English source: the river bank <eos>. Receiving decoder row: la.

    $m_2=\sum_{j=1}^{4}\alpha_{2j}\vv{v_j^{\mathrm{cross}}}$ is a three-number message for this query.

    The source values before weighting

    The message coordinates have learned numerical roles. We do not label them as three facts or three words. Their use is determined by the output projection and the rest of the model.

    $\vd{\Delta e_2^{\mathrm{cross}}}=m_2W_O^{\mathrm{cross}}$, then $\vp{e_2^{\mathrm{cross}}}=\vp{e_2^{\mathrm{self}}}+\vd{\Delta e_2^{\mathrm{cross}}}$.

    message = A @ V
    delta_e = message @ W_O_cross
    decoder_out = decoder_self + delta_e

    The learned output projection

    In a larger model, the value width and representation width can differ. The output projection must return the representation width before the residual addition.

    $\ell_2=\vp{e_2^{\mathrm{cross}}}W_{\mathrm{vocab}}+b_{\mathrm{vocab}}$ gives one logit per French vocabulary token.

    logits = decoder_out @ W_vocab + b_vocab

    Use the updated decoder row at la and the rive column of $W_{\mathrm{vocab}}$.

    Repeat with each vocabulary column and its bias to get all five logits.

    Source: the river bank. French prefix: <bos> la. Observed next token: rive.

    Attention weights four source positions; vocabulary softmax gives five next-token probabilities.

    We have completed a next-token prediction. Training will compare it with the observed token; generation will use a decoding rule to choose a token. Both start with this same forward calculation.

    06

    One teacher-forced pass predicts la, rive, and the end marker.

    Train on the observed translation

    During training, the observed French translation supplies the correct prefix for every prediction. This is teacher forcing. We shift the sequence by one position: the decoder inputs are <bos> la rive, and the targets are la rive <eos>. The causal target mask lets us compute all three rows in parallel without letting a row read its answer. Every row can read the complete English source through cross-attention. We average the three next-token losses, then differentiate that loss through the same forward graph.

    English source for every row: the river bank <eos>.

    decoder_ids = target_ids[:-1]  # BOS, la, rive
    labels = target_ids[1:]        # la, rive, EOS

    The la row predicts rive; the rive row predicts the end marker.

    $\vq{Q}:3\times3$, $\vk{K}:4\times3$; therefore $QK^\top/\sqrt{3}:3\times4$.

    Each row sums to one across the four English source columns. Cross-attention need not have a square matrix.

    The rectangular score matrix before softmax

    Softmax is applied separately to each row. The target self-attention matrix was $3\times3$ and triangularly masked. This $3\times4$ cross-attention matrix reads the four known source tokens, so it has no future-target triangle.

    logits = model(source_ids, decoder_ids)  # shape: 3 × 5
    loss = F.cross_entropy(logits, labels)

    Gradients reach both the source-value path and the query/key path, then the underlying embeddings and projections.

    The PyTorch lines show the training interface. This browser demonstration reads checkpoints and a verified update from the toy's numerical engine; it does not run PyTorch in the browser. The saved gradients were checked numerically. In a PyTorch implementation, autograd performs the backward pass and the optimizer uses those gradients.

    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

    The optimizer changes learned tables, projections, and the vocabulary head. A fresh forward pass recomputes $Q,K,V$, weights, and decoder rows.

    The learned parameters are the two embedding tables, their position tables, each attention block's $W_Q,W_K,W_V,W_O$, and the vocabulary head. The optimizer updates these stored parameters, not the cached query, key, value, or attention-weight arrays.

    The smaller mean loss shows that this particular step improved the fit to the observed river translation. It did not improve every prediction: the river example's end-marker loss rises slightly, and the fitted financial example's probability of banque falls. Generalisation needs separate data. The checkpoint's second fitted pair, the financial bank to la banque, is a check that the model can use its source input, not an evaluation of translation quality.

    07

    At generation time, the model must build its own French prefix.

    Generate one token at a time

    Training supplied the correct French prefix. Generation starts with only the source and <bos>. We use greedy decoding: choose the highest-probability next token, append it, and repeat. The examples below use the checkpoint after the displayed update. Each distribution is computed from the prefix the model has generated so far. The English source, encoded rows, and cross-attention keys and values stay fixed. The last decoder row, its query, and its weights can change as the French prefix grows.

    next_id = logits[-1].argmax()
    prefix = torch.cat([prefix, next_id[None]])

    Repeat until the chosen token is <eos>. With parameters fixed, cache the encoded source and its cross-attention $K,V$.

    How the source-reading weights change

    Each row belongs to a different decoder query over the same four source positions. Greedy decoding is only one rule for choosing a token; a sampling rule would still use the same model probabilities. Real applications also impose a length limit. This toy stops when it predicts its end marker.

    08

    To identify an attention block, ask where its query, keys, and values come from.

    Same operation, different sources

    The encoder, decoder self-attention, and cross-attention all use the same match, softmax, and weighted-value calculation. What changes is where the rows come from and which ones are available. Encoder self-attention reads the complete source. Decoder self-attention reads the allowed target prefix. Cross-attention lets those target representations read the encoded source. Its output updates the decoder representation at each receiving position.

    Self-attention uses one stream for $Q,K,V$. Cross-attention takes $Q$ from the decoder and $K,V$ from the encoded source.

    To predict rive, the decoder row at la combines the French prefix with information read from the encoded English phrase.

    In a vision-language model that uses cross-attention, the same arrangement can let text queries read keys and values computed from encoded image patches.

    The encoder-decoder arrangement and scaled dot-product attention follow Vaswani et al., Attention Is All You Need, especially Sections 3.1 and 3.2.3. The paper's full model uses multi-head attention, feed-forward sublayers, normalisation, and repeated blocks. Our three-coordinate toy isolates one pass through the attention blocks so that their inputs and outputs remain visible.