Attention and language · Notebook 5 · optional detailed lab

From a story to a trained predictor

One sentence. Seven training examples. Two models. Follow the data, every tensor and one learning step before generating a new token. The compact four-head classroom walkthrough is now in Part III.

Before you start

The small example uses B=2 examples, w=4 context slots and C=10 vocabulary items. Every number is generated by the supplied code, not drawn by hand. The final measured comparison uses a separate 6,000-story corpus and trained models.

Unzip the download, open a terminal in that folder and run:

python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
jupyter lab 05_training_and_inference_maps.ipynb

Then choose Run → Run All Cells. This notebook needs no GPU or data download. Notebooks 1 and 3 explain the real corpus and training runs. Keep wordlm.py and the other helpers beside the notebooks.

1. Data and tokens · Step 01 / 88Open optional worked slide ↗

The training map: stories

We begin with complete stories. The highlighted box supplies the data for everything that follows.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
1. Data and tokens · Step 02 / 88Open optional worked slide ↗

The data: complete short stories

The saved experiment uses 6,000 TinyStories documents. Each document is a complete story. This is an excerpt from source row 12400.

The data: complete short storiesTinyStories · source row 12400“Once upon a time, there lived a little bunny.”One excerpt from a complete story; 6,000 documents in this subset.First, look at whole stories. Then we will build training pairs.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
sample = evidence['sample']
print(sample['text_excerpt'])
print('Source row:', sample['row_idx'])
Printed output
Once upon a time, there lived a little bunny.
Source row: 12400
1. Data and tokens · Step 03 / 88Open optional worked slide ↗

Read one complete story

One document is one complete story, not one sentence. This is the shortest story in our saved 6,000-document subset; many others are longer.

Read one complete storyTinyStories · source row 1992490 · complete text52 words · 60 tokensOnce upon a time, there was an old hotel. The hotel was very big and hadmany rooms. One day, a big storm came and lightning struck the hotel. Thehotel was very old and could not handle the strike. The hotel fell down andeveryone had to go home. The end.Everything above is one document, from the opening to “The end.”
Same figure as the lecture. Full-story lengths are computed below. TinyStories · Ronen Eldan & Yuanzhi Li · source dataset · CDLA-Sharing-1.0. Display labels added; […] marks omissions.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
short_story = story_examples['examples'][0]
short_words = len(short_story['text'].split())
short_tokens = len(tokenize(short_story['text']))
assert (short_words, short_tokens) == (52, 60)
print(short_story['text'])
print(f'Full story: {short_words} words; {short_tokens} tokens.')
Printed output
Once upon a time, there was an old hotel. The hotel was very big and had many rooms. One day, a big storm came and lightning struck the hotel. The hotel was very old and could not handle the strike. The hotel fell down and everyone had to go home. The end.
Full story: 52 words; 60 tokens.
1. Data and tokens · Step 04 / 88Open optional worked slide ↗

Longer stories have several paragraphs

These beginnings and endings come from two other documents in the same subset. […] marks omitted text; the lengths count each whole story. Words are whitespace-separated items; tokens also separate punctuation and exclude start/end markers.

Longer stories have several paragraphsSource row 12400107 words · 133 tokensOnce upon a time, there lived a littlebunny.[…]He thanked the rock for the new andweird carrot and hopped away to findmore delicious snacks.Full story: 6 paragraphsSource row 588307248 words · 300 tokensSara and Tom were playing with theircrayons and paper.[…]She wanted to find a new friend whowould be nice and kind.Full story: 8 paragraphsTraining stories: 60–942 tokens; median 176 tokens.
Same figure as the lecture. Full-story lengths are computed below. TinyStories · Ronen Eldan & Yuanzhi Li · source dataset · CDLA-Sharing-1.0. Display labels added; […] marks omissions.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
long_stories = story_examples['examples'][1:]
story_lengths = []
for story in long_stories:
    counts = (len(story['text'].split()), len(tokenize(story['text'])))
    assert counts == (story['word_count'], story['token_count'])
    story_lengths.append(counts)
    print(f"Source row {story['row_idx']}: {counts[0]} words; {counts[1]} tokens")
    print(story['text'] + '\n')  # full texts here; excerpts in the figure
train_lengths = evidence['audit']['story_length_tokens']['train']
print('Training-story token lengths:', train_lengths)
Printed output
Source row 12400: 107 words; 133 tokens
Once upon a time, there lived a little bunny. The bunny loved carrots, so every day the bunny went outside to find carrots. One day, the bunny was looking for some carrots and came across something strange. It looked like a weird rock.

The bunny hopped closer and asked, "What are you?"

The rock replied, "I'm a weird carrot!"

The bunny was very surprised and said, "Weird carrots don't exist!"

But the rock said, "Oh yes they do! I'm a cool carrot that always rocks!"

The bunny was amazed. He thanked the rock for the new and weird carrot and hopped away to find more delicious snacks.

Source row 588307: 248 words; 300 tokens
Sara and Tom were playing with their crayons and paper. They liked to draw animals and flowers and cars. Sara was drawing a big red pepper, because she liked to eat peppers. Tom was drawing a blue fish, because he liked to swim.

"Look at my pepper!" Sara said to Tom. "It is very pretty and shiny. Do you like it?"

Tom looked at Sara's pepper. He did not like it. He thought it was boring and ugly. He wanted to make Sara feel bad. He was very rude.

"I don't like your pepper," Tom said to Sara. "It is not pretty and shiny. It is silly and dumb. Watch this!"

Tom took his black crayon and made a big mark on Sara's pepper. He drew a line across it and made a face with a tongue sticking out. He laughed and said, "Now your pepper is funny and silly!"

Sara saw what Tom did to her pepper. She felt very sad and angry. She liked her pepper and she worked hard to draw it. She did not think it was funny or silly. She thought it was mean and nasty.

"Tom, you are very rude!" Sara said to Tom. "You ruined my pepper! You are not my friend! Go away!"

Sara took her paper and crayons and ran to the other side of the room. She did not want to play with Tom anymore. She wanted to find a new friend who would be nice and kind.

Training-story token lengths: {'max': 942, 'median': 176, 'min': 60}
1. Data and tokens · Step 05 / 88Open optional worked slide ↗

The training map: splitting stories

Keep each story in one split before making overlapping windows. Only the training stories fit the vocabulary and model.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
1. Data and tokens · Step 06 / 88Open optional worked slide ↗

Stories stay together when we split the data

The split happens before overlapping windows are made. Training fits parameters and the vocabulary. Validation chooses settings and checkpoints. Test measures the frozen choice.

Stories stay together when we split the dataSplitWhole storiesUsed forTrain4,822Parameter and vocabulary learningValidation586Settings and checkpoint selectionTest592Final held-out measurement
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
audit = evidence['audit']
print(audit['documents'])
assert sum(audit['documents'].values()) == 6000
assert not any(audit['document_overlap'].values())
Printed output
{'test': 592, 'train': 4822, 'validation': 586}
1. Data and tokens · Step 07 / 88Open optional worked slide ↗

A sentence we can trace completely

We use this authored sentence for the arithmetic. Its small vocabulary and randomly initialized models are separate from the measured TinyStories experiment.

A sentence we can trace completelyLily found a red ball.A complete authored example, including its final punctuation.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
sentence = 'Lily found a red ball.'
print(sentence)
Printed output
Lily found a red ball.
1. Data and tokens · Step 08 / 88Open optional worked slide ↗

Tokenization

A token is one item the model reads or predicts. The tokenizer chooses these items before we assign integer IDs and look up learned embeddings.

TokenizationDoes “next token” always mean “next word”?redder!The same text can become different sequences of tokens.
Same figure as the lecture. Numeric values are computed from the code below. Background: Hugging Face tokenizer overview · Sennrich et al., 2016: subword units. Notebook rules: wordlm.py.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
tokenization_text = 'redder!'
print('The same text for three tokenization choices:', tokenization_text)
Printed output
The same text for three tokenization choices: redder!
1. Data and tokens · Step 09 / 88Open optional worked slide ↗

Words, characters or pieces of words

Whole-word vocabularies need entries for many word forms, while characters produce longer sequences. Subword methods such as byte pair encoding (BPE) learn reusable pieces from training text to balance these concerns.

Words, characters or pieces of wordsChoiceTokens for redder!CountWord + punctuation[redder] [!]2Character[r] [e] [d] [d] [e] [r] [!]7Subword (illustrative)[red] [der] [!]3Context length counts tokens. The same width can cover different amounts of text.The subword split is illustrative. Actual splits depend on the tokenizer.
Same figure as the lecture. Numeric values are computed from the code below. Background: Hugging Face tokenizer overview · Sennrich et al., 2016: subword units. Notebook rules: wordlm.py.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
tokenization_choices = {
    'Word + punctuation': tokenize(tokenization_text),
    'Character': list(tokenization_text),
    'Subword (illustrative)': ['red', 'der', '!'],
}
tokenization_counts = {name: len(parts) for name, parts in tokenization_choices.items()}
assert list(tokenization_counts.values()) == [2, 7, 3]
for name, parts in tokenization_choices.items():
    assert ''.join(parts) == tokenization_text
    print(name, parts, 'tokens:', len(parts))
# The subword split is hand-chosen, not the output of a trained BPE tokenizer.
Printed output
Word + punctuation ['redder', '!'] tokens: 2
Character ['r', 'e', 'd', 'd', 'e', 'r', '!'] tokens: 7
Subword (illustrative) ['red', 'der', '!'] tokens: 3
1. Data and tokens · Step 10 / 88Open optional worked slide ↗

The tokenizer used in this notebook

Words plus punctuation keep our calculations easy to inspect. Both models use the same rules and vocabulary during training and generation. We build the vocabulary from training stories only and map missing words to the unknown-token marker, UNK.

The tokenizer used in this notebookText: Lily can't find 12 balls![lily] [can't] [find] [12] [balls] [!]Lily and lily share a token. Spaces separate words but are not tokens here.The apostrophe stays inside can't. The digits in 12 stay together.The exclamation mark is a separate token, just like the full stop in our story.English teaching tokenizer. Original case and spacing are lost.
Same figure as the lecture. Numeric values are computed from the code below. Background: Hugging Face tokenizer overview · Sennrich et al., 2016: subword units. Notebook rules: wordlm.py.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
tokenizer_probe = "Lily can't find 12 balls!"
probe_tokens = tokenize(tokenizer_probe)
assert probe_tokens == ['lily', "can't", 'find', '12', 'balls', '!']
assert tokenize('LILY   found!') == ['lily', 'found', '!']
print(tokenizer_probe)
print(probe_tokens)
# This English teaching tokenizer drops case and spacing.
# Its word pattern is ASCII-based; it is not a general multilingual tokenizer.
Printed output
Lily can't find 12 balls!
['lily', "can't", 'find', '12', 'balls', '!']
1. Data and tokens · Step 11 / 88Open optional worked slide ↗

The training map: tokens and IDs

Text becomes tokens, then integer vocabulary IDs. Learned embedding vectors come later at token lookup.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
1. Data and tokens · Step 12 / 88Open optional worked slide ↗

Lowercase words and separate punctuation

This tokenizer normalizes Unicode, lowercases text and keeps punctuation as tokens. Five words plus the full stop give six tokens. A tokenizer decides what one prediction unit is.

Lowercase words and separate punctuationText: Lily found a red ball.Tokens:lilyfoundaredball.5 word tokens + 1 punctuation token = 6 ordinary tokens
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
pieces = tokenize(sentence)
print(pieces)
assert pieces == ['lily', 'found', 'a', 'red', 'ball', '.']
assert len(pieces) == 6
Printed output
['lily', 'found', 'a', 'red', 'ball', '.']
1. Data and tokens · Step 13 / 88Open optional worked slide ↗

Every vocabulary item gets an integer ID

Six ordinary tokens + four special tokens = C=10; their IDs are arbitrary labels. Here, BOS and EOS mark one complete story (a sequence), not each sentence. PAD means padding; BOS means beginning of sequence; EOS means end of sequence; UNK means unknown token. The benchmark uses a different, frequency-ranked vocabulary with C=4,000.

Every vocabulary item gets an integer IDIDTokenMeaning0<PAD>Padding1<BOS>Beginning of sequence2<EOS>End of sequence3<UNK>Unknown token4.PunctuationIDVocabulary item5a6ball7found8lily9red
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
words = list(SPECIAL_TOKENS) + sorted(set(pieces))
vocab = Vocabulary(words, {t:i for i,t in enumerate(words)}, {}, 10, 1)
C = len(words)
assert C == 10
print(vocab.stoi)
Printed output
{'<PAD>': 0, '<BOS>': 1, '<EOS>': 2, '<UNK>': 3, '.': 4, 'a': 5, 'ball': 6, 'found': 7, 'lily': 8, 'red': 9}
1. Data and tokens · Step 14 / 88Open optional worked slide ↗

Four special tokens have different jobs

blue is missing from our toy vocabulary, so this lookup returns [3], the UNK ID. boundaries=False means “do not add BOS or EOS”. The input ['blue'] is already a list of tokens; encode_tokens looks up their integer IDs, without creating embeddings. With boundaries=True, the same call returns [1, 3, 2]: BOS, UNK, EOS.

Four special tokens have different jobsIDTokenRole0PADFill unused input slots1BOSFirst context token of a story2EOSPredict that the story ends3UNKRepresent a missing vocabulary item
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
assert 'blue' not in vocab.stoi
token_ids = vocab.encode_tokens(['blue'], boundaries=False)
print(token_ids)  # [3]: UNK
assert token_ids == [3] == [vocab.unk_id]

with_boundaries = vocab.encode_tokens(['blue'], boundaries=True)
print('With BOS and EOS:', with_boundaries)  # [1, 3, 2]
assert with_boundaries == [vocab.bos_id, vocab.unk_id, vocab.eos_id]
Printed output
[3]
With BOS and EOS: [1, 3, 2]
1. Data and tokens · Step 15 / 88Open optional worked slide ↗

The training map: story boundaries

Add BOS and EOS to each story’s ID sequence. These boundaries let us create the first input and the final stopping target.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
1. Data and tokens · Step 16 / 88Open optional worked slide ↗

One BOS and EOS per document

In this notebook, each document is a complete story with one BOS at its start and one EOS at its end. We do not add extra markers between sentences. The Lily sentence is our entire toy document, so its six ordinary tokens become eight IDs including BOS and EOS. BOS provides context, and the remaining seven IDs are prediction targets. encode_tokens wraps the token list we pass it, without detecting sentences. We pass a complete story once before creating its training windows, which never cross document boundaries. Other datasets may use different boundary conventions.

One BOS and EOS per documentOne complete document: our one-sentence story<BOS>1lily8found7a5red9ball6.4<EOS>28 IDs total. Predict each ID after BOS: 7 supervised targets.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
ids = vocab.encode_tokens(pieces, boundaries=True)
print(ids)
assert ids == [1, 8, 7, 5, 9, 6, 4, 2]

# Two sentences still make one document with one pair of markers.
two_sentence_story = 'Lily found a ball. Lily found a red ball.'
two_sentence_ids = vocab.encode_tokens(tokenize(two_sentence_story), boundaries=True)
print(' '.join(vocab.decode_ids(two_sentence_ids, skip_special=False)))
assert two_sentence_ids.count(vocab.bos_id) == 1
assert two_sentence_ids.count(vocab.eos_id) == 1
assert two_sentence_ids.count(vocab.stoi['.']) == 2
Printed output
[1, 8, 7, 5, 9, 6, 4, 2]
<BOS> lily found a ball . lily found a red ball . <EOS>
2. Windows and batches · Step 17 / 88Open optional worked slide ↗

The training map: context and target

A window holds the previous w token IDs. Its target is the next observed ID in the story.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
2. Windows and batches · Step 18 / 88Open optional worked slide ↗

Token positions and vocabulary IDs

Python positions start at 0: position 4 contains red, whose vocabulary ID is 9. With a four-token window, this example reads positions 0 through 3 and predicts the token at position 4.

Token positions and vocabulary IDsStory position01234567Token<BOS>lilyfoundaredball.<EOS>Token ID18759642Input: positions 0, 1, 2, 3Target: position 4, ID 9 (red)
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
w = 4
target_position = 4
print(list(enumerate(ids)))
Printed output
[(0, 1), (1, 8), (2, 7), (3, 5), (4, 9), (5, 6), (6, 4), (7, 2)]
2. Windows and batches · Step 19 / 88Open optional worked slide ↗

One input window and its observed target

context_ids is the input for one example, often called x, and target_id is its single observed answer, often called y. The slice ids[0:4] stops before position 4, so red is outside this input.

One input window and its observed targetcontext_ids: one input xBOS lily found a[1, 8, 7, 5]target_id: one answer yredID 9ids[0:4] contains four IDsids[4] is one ID
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
context_ids = ids[target_position-w:target_position]
target_id = ids[target_position]
assert context_ids == [1, 8, 7, 5] and target_id == 9
print('context_ids:', context_ids, 'target_id:', target_id)
Printed output
context_ids: [1, 8, 7, 5] target_id: 9
2. Windows and batches · Step 20 / 88Open optional worked slide ↗

Lists for inputs and targets

We inspected the example with red as its target, but have not stored any pairs yet. To collect every example in order, start at position 1, where lily follows BOS.

Lists for inputs and targetscontexts = []Will hold one input list per exampletargets = []Will hold one answer per exampleStored so far: 0 input rows and 0 targets
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
contexts = []
targets = []
assert len(contexts) == len(targets) == 0
2. Windows and batches · Step 21 / 88Open optional worked slide ↗

The first pair has only BOS as history

At position t=1, the visible prefix is [1], meaning BOS. Three PAD IDs fill the unused slots, giving context_ids=[0, 0, 0, 1] and target_id=8 for lily.

The first pair has only BOS as historyt = 1 visible = [1] (BOS)3 PAD IDs + 1 known token ID = 4 input slotscontext_ids = [0, 0, 0, 1]target_id = 8PAD PAD PAD BOSlily
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
t = 1
visible = ids[max(0, t-w):t]
context_ids = [vocab.pad_id] * (w-len(visible)) + visible
target_id = ids[t]
assert context_ids == [0, 0, 0, 1] and target_id == 8
2. Windows and batches · Step 22 / 88Open optional worked slide ↗

Storing an input and its target

contexts.append(context_ids) adds the whole four-ID list as one row. targets.append(target_id) adds one answer at the same row index, so contexts[0] and targets[0] belong together.

Storing an input and its targetcontexts: a list of input liststargets: a list of IDs[[8][0, 0, 0, 1]row 0targets[0] = 8 (lily)]1 input row, 1 target. Same index means the same example.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
contexts.append(context_ids)
targets.append(target_id)
assert contexts == [[0, 0, 0, 1]] and targets == [8]
print('contexts =', contexts)
print('targets =', targets)
Printed output
contexts = [[0, 0, 0, 1]]
targets = [8]
2. Windows and batches · Step 23 / 88Open optional worked slide ↗

The next pair includes the previous observed target

At t=2, BOS and lily form the known prefix, and the observed next word is found. Computing this pair changes context_ids and target_id, while the stored lists still contain only the first pair.

The next pair includes the previous observed targett = 2 visible = [1, 8] (BOS lily)2 PAD IDs + 2 known token IDs = 4 input slotscontext_ids = [0, 0, 1, 8]target_id = 7PAD PAD BOS lilyfound
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
t = 2
visible = ids[max(0, t-w):t]
context_ids = [vocab.pad_id] * (w-len(visible)) + visible
target_id = ids[t]
assert context_ids == [0, 0, 1, 8] and target_id == 7
assert contexts == [[0, 0, 0, 1]] and targets == [8]
2. Windows and batches · Step 24 / 88Open optional worked slide ↗

The second pair becomes row 1 in both lists

The same two append calls add the new pair after the first one. Each row in contexts still lines up with exactly one entry in targets.

The second pair becomes row 1 in both listscontexts: a list of input liststargets: a list of IDs[[8, 7][0, 0, 0, 1],row 0targets[0] = 8 (lily)[0, 0, 1, 8]row 1targets[1] = 7 (found)]2 input rows, 2 targets. Same index means the same example.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
contexts.append(context_ids)
targets.append(target_id)
assert contexts == [[0, 0, 0, 1], [0, 0, 1, 8]]
assert targets == [8, 7]
print('contexts =', contexts)
print('targets =', targets)
Printed output
contexts = [[0, 0, 0, 1], [0, 0, 1, 8]]
targets = [8, 7]
2. Windows and batches · Step 25 / 88Open optional worked slide ↗

Building the remaining training pairs

Positions 1 and 2 are already stored, so this loop continues at 3 and stops before len(ids)=8. Each pass builds a fresh input list and appends it with the observed target, producing seven paired rows in total.

Building the remaining training pairsAlready storedThis loop addsTotal pairst = 1, 2t = 3, 4, 5, 6, 77
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
for t in range(3, len(ids)):
    visible = ids[max(0, t-w):t]
    context_ids = [vocab.pad_id] * (w-len(visible)) + visible
    target_id = ids[t]
    contexts.append(context_ids)
    targets.append(target_id)
assert len(contexts) == len(targets) == 7
assert contexts[3] == ids[0:4] and targets[3] == ids[4]
2. Windows and batches · Step 26 / 88Open optional worked slide ↗

Early prefixes need padding

These are the first four stored rows, with the IDs translated back to tokens. Row 3 is the red-target pair we inspected first, now stored in contexts[3] and targets[3].

Early prefixes need paddingStored rowInput: exactly four slotsNext target0PAD PAD PAD BOSlily1PAD PAD BOS lilyfound2PAD BOS lily founda3BOS lily found ared
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
for row in range(4):
    print(row, contexts[row], targets[row])
assert contexts[3] == [1, 8, 7, 5] and targets[3] == 9
Printed output
0 [0, 0, 0, 1] 8
1 [0, 0, 1, 8] 7
2 [0, 1, 8, 7] 5
3 [1, 8, 7, 5] 9
2. Windows and batches · Step 27 / 88Open optional worked slide ↗

Later prefixes drop their oldest tokens

The remaining three rows use full windows, keeping only the four tokens immediately before each target. The last target is EOS, and no window crosses into a different story.

Later prefixes drop their oldest tokensStored rowInput: exactly four slotsNext target4lily found a redball5found a red ball.6a red ball .<EOS>Every ordinary token plus EOS is predicted once: N = 6 + 1 = 7.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
for row in range(4, 7):
    print(row, contexts[row], targets[row])
assert targets[-1] == vocab.eos_id
Printed output
4 [8, 7, 5, 9] 6
5 [7, 5, 9, 6] 4
6 [5, 9, 6, 4] 2
2. Windows and batches · Step 28 / 88Open optional worked slide ↗

The paired lists become two integer tensors

The conversion preserves every token ID and row pairing. all_X has shape [7, 4], all_y has shape [7], and torch.long means integer IDs.

The paired lists become two integer tensorsRowcontexts[r] = all_X[r]targets[r] = all_y[r]Word0[0, 0, 0, 1]8lily1[0, 0, 1, 8]7found2[0, 1, 8, 7]5a3[1, 8, 7, 5]9red4[8, 7, 5, 9]6ball5[7, 5, 9, 6]4.6[5, 9, 6, 4]2<EOS>all_X: 7 rows × 4 IDsall_y: 7 targets N = 7
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
all_X = torch.tensor(contexts, dtype=torch.long)
all_y = torch.tensor(targets, dtype=torch.long)
N = len(all_y)
assert N == 7 and all_X.shape == (7, 4)
assert all_X.tolist() == contexts and all_y.tolist() == targets
assert not all_y.eq(vocab.pad_id).any()
2. Windows and batches · Step 29 / 88Open optional worked slide ↗

Six ordinary tokens give seven training examples

Each ordinary token is a target once, and EOS supplies one more target. BOS starts the history, while PAD fills empty input slots, so neither adds a target.

Six ordinary tokens give seven training examplesOrdinary tokens6EOS target1Training examples7+=Every ordinary token and each story ending supplies one target.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
ordinary_tokens = len(pieces)
examples_in_story = ordinary_tokens + 1
assert ordinary_tokens == 6 and examples_in_story == N == 7
2. Windows and batches · Step 30 / 88Open optional worked slide ↗

Training examples across the corpus

The saved training split has 964,338 ordinary tokens across 4,822 complete stories. Adding one EOS target per story gives 969,160 training examples.

Training examples across the corpusOrdinary tokens964,338One EOS per story4,822Training examples969,160+=Every ordinary token and each story ending supplies one target.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
train_tokens = 964_338
train_stories = 4_822
train_examples = train_tokens + train_stories
assert train_tokens == audit['oov']['train']['tokens']
assert train_stories == audit['documents']['train']
assert train_examples == 969160
2. Windows and batches · Step 31 / 88Open optional worked slide ↗

Example counts for each data split

For each split, add its ordinary-token count and its story count. These are counts of supervised examples, not optimizer steps.

Example counts for each data splitSplitOrdinary tokensOne EOS / storyTotal targetsTrain964,3384,822969,160Validation119,634586120,220Test119,801592120,393For one story: 6 + 1 = 7. Across training stories: 964,338 + 4,822 = 969,160.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
window_counts = {
    split: audit['oov'][split]['tokens'] + count
    for split, count in audit['documents'].items()
}
assert window_counts['train'] == 969160
assert window_counts['train'] == train_examples
print(window_counts)
Printed output
{'test': 120393, 'train': 969160, 'validation': 120220}
2. Windows and batches · Step 32 / 88Open optional worked slide ↗

Context length controls what is visible

History is everything already known. The context window is the suffix read for this prediction. Increasing w can recover an older clue, but also changes model size or computation.

Context length controls what is visiblewVisible suffixTarget count / story2red ball74found a red ball76BOS lily found a red ball7
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
history = ids[:6]  # BOS lily found a red ball
contexts_by_width = {width: history[-width:] for width in [2, 4, 6]}
print(contexts_by_width)
Printed output
{2: [9, 6], 4: [7, 5, 9, 6], 6: [1, 8, 7, 5, 9, 6]}
2. Windows and batches · Step 33 / 88Open optional worked slide ↗

The training map: one batch

We select B=2 examples, each with w=4 input slots. They form X with shape [2,4] and y with shape [2].

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
2. Windows and batches · Step 34 / 88Open optional worked slide ↗

A batch of two examples

Each batch row pairs four input token IDs in X with one observed next-token ID in y. The text columns decode the IDs, and the targets come from the story.

A batch of two examplesData rowBatch rowX: input IDsInput tokensy: IDTarget token20[0, 1, 8, 7]<PAD> <BOS> lily found5a31[1, 8, 7, 5]<BOS> lily found a9redDataset rows 2 and 3 become batch rows 0 and 1.PAD fills an empty slot. BOS marks the start of the story.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
selected = torch.tensor([2, 3])
X, y = all_X[selected], all_y[selected]
B = X.shape[0]
assert X.tolist() == [[0, 1, 8, 7], [1, 8, 7, 5]]
assert y.tolist() == [5, 9] and B == 2
2. Windows and batches · Step 35 / 88Open optional worked slide ↗

X and y contain integer token IDs

Tokenization and vocabulary lookup are complete: X and y contain integers, with no embedding coordinates yet. Next, X goes through embedding lookup, while y stays as target IDs for the loss.

X and y contain integer token IDsX: input token IDsBatch rowSlot 1Slot 2Slot 3Slot 40018711875X.shape = (2, 4)2 examples × 4 input slotsy: next-token IDs[5, 9]y.shape = (2,)2 targets, one per example
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
print(X.shape, X.dtype)
print(y.shape, y.dtype)
assert X.shape == (2, 4) and y.shape == (2,)
assert X.dtype == y.dtype == torch.long
Printed output
torch.Size([2, 4]) torch.int64
torch.Size([2]) torch.int64
2. Windows and batches · Step 36 / 88Open optional worked slide ↗

Seven examples, four batches

B=2 with drop_last=False gives batches of 2, 2, 2 and 1 per pass. With one update per batch, that is four optimizer steps. The benchmark samples 512 windows per update with replacement.

Seven examples, four batchesSequential batchExample indicesActual B10, 1222, 3234, 52461
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
batch_sizes = [len(all_y[start:start+B]) for start in range(0, N, B)]
assert batch_sizes == [2, 2, 2, 1]
print('One sequential pass:', batch_sizes)
print('Benchmark target presentations:', 6000 * 512)
Printed output
One sequential pass: [2, 2, 2, 1]
Benchmark target presentations: 3072000
2. Windows and batches · Step 37 / 88Open optional worked slide ↗

Batch and model dimensions

C includes the special tokens. Queries and keys use the same width dₖ, while values can use a different width dᵥ.

Batch and model dimensionsData dimensionsSymbolValueWhat it countsB2examples per batchw4input slots per exampleC10vocabulary itemsModel dimensionsSymbolValueWhat it countsd4embedding coordinatesh8prediction-head hidden unitsdₖ3query and key coordinatesdᵥ2value coordinates
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
d, h, d_k, d_v = 4, 8, 3, 2
torch.manual_seed(11)
mlp = FixedWindowMLP(C, w, d, h, vocab.pad_id)
torch.manual_seed(11)
attention = CausalAttentionLM(C, w, d, d_k, d_v, h, vocab.pad_id)
print(dict(B=B, w=w, C=C, d=d, h=h, d_k=d_k, d_v=d_v))
Printed output
{'B': 2, 'w': 4, 'C': 10, 'd': 4, 'h': 8, 'd_k': 3, 'd_v': 2}
3. The MLP forward pass · Step 38 / 88Open optional worked slide ↗

The full MLP training map

X has shape [2,4] and y has shape [2]. Input IDs enter token lookup, while the observed target goes directly to the loss. Each optimizer step updates all learned layers for the next batch.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
print('X:', tuple(X.shape), 'y:', tuple(y.shape))
Printed output
X: (2, 4) y: (2,)
3. The MLP forward pass · Step 39 / 88Open optional worked slide ↗

An ID selects one embedding row

PAD fills empty slots: padding_idx=0 initializes its row to zero and skips its embedding gradient, so it stays zero here. These are numeric zeros, not nulls, while BOS/EOS/UNK have random, trainable rows that can learn when used as inputs. The table has C=10 rows and d=4 columns. ID 7, found, selects one four-number row. This is the initial table, before training, and the coordinates have no assigned word meanings. Only PAD has this zero-row rule in our model. EOS ends a story and is only a target in these training windows, so its input embedding receives no task gradient here, although the separate output layer learns to predict EOS. UNK can receive an embedding gradient when a missing word maps to it in an input. Zeroing PAD is a model choice, not a rule for all special tokens. In the attention model, we also mask padded source positions because a zero embedding alone does not remove them from softmax. Reference: https://docs.pytorch.org/docs/stable/generated/torch.nn.Embedding.html

An ID selects one embedding rowT: C=10 rows × d=4 coordinates · initial weightsTokenIDcoord. 1coord. 2coord. 3coord. 4<PAD>00.0000.0000.0000.000found7-0.687-0.800-1.369-0.325lily8-0.4301.7800.057-0.642red90.1861.0100.146-0.639
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
embedding_table = mlp.token_embedding.weight
assert mlp.token_embedding.padding_idx == vocab.pad_id == 0
assert embedding_table[vocab.pad_id].eq(0).all()
found_id = vocab.stoi['found']
found_vector = embedding_table[found_id]
assert found_vector.shape == (4,)
3. The MLP forward pass · Step 40 / 88Open optional worked slide ↗

Input IDs select rows from the embedding table

The learned table T has shape [C,d] = [10,4]. Each ID in X[1] selects one complete row, giving four vectors in the same slot order. The next slide stacks this result for both batch examples.

Input IDs select rows from the embedding tableX[1]: four input IDsSlotIDToken01<BOS>18lily27found35aT: 10 rows × 4 coordinatesT rowc₁c₂c₃c₄1-0.18-1.500.14-0.528-0.431.780.06-0.647-0.69-0.80-1.37-0.325-0.15-1.101.370.69Four selected rows of the 10-row tableE[1]: four vectors in slot orderSlotc₁c₂c₃c₄0-0.18-1.500.14-0.521-0.431.780.06-0.642-0.69-0.80-1.37-0.323-0.15-1.101.370.69Both examples together: X [2,4] indexes T [10,4] to produce E [2,4,4].
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
selected_rows = embedding_table[X[1]]
assert selected_rows.shape == (w, d)
assert torch.equal(selected_rows[2], found_vector)
3. The MLP forward pass · Step 41 / 88Open optional worked slide ↗

The whole batch becomes a 2 × 4 × 4 tensor

Each of the two examples has four token slots. Each slot receives four embedding coordinates. PAD uses the fixed zero row. The same token ID selects the same row wherever it occurs.

The whole batch becomes a 2 × 4 × 4 tensorExample 0: PAD BOS lily foundTokenc1c2c3c4<PAD>0.0000.0000.0000.000<BOS>-0.182-1.4970.142-0.524lily-0.4301.7800.057-0.642found-0.687-0.800-1.369-0.325Example 1: BOS lily found aTokenc1c2c3c4<BOS>-0.182-1.4970.142-0.524lily-0.4301.7800.057-0.642found-0.687-0.800-1.369-0.325a-0.155-1.0961.3670.689
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
E_mlp = mlp.token_embedding(X)
assert E_mlp.shape == (2, 4, 4)
assert E_mlp[0, 0].eq(0).all()
assert torch.equal(E_mlp[0, 2], E_mlp[1, 1])  # lily
3. The MLP forward pass · Step 42 / 88Open optional worked slide ↗

The MLP map: joining the input vectors

Token lookup produced E with shape [B,w,d] = [2,4,4]. Each example now becomes one row of w×d=16 numbers.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
3. The MLP forward pass · Step 43 / 88Open optional worked slide ↗

Flatten the context axis, keep the batch axis

Each of the four tokens supplies four input nodes, giving 16 numbers in slot order. Both batch examples use this same 16 → 8 → 10 network separately. Flattening only rearranges the coordinates and learns no weights. We never concatenate one batch example with another. Swapping token positions changes which input connections receive each embedding.

Flatten the context axis, keep the batch axisExample 1: <BOS> lily found aSame weights for both batch examples16 input nodesflat: [2,16]8 hidden neuronshidden: [2,8]10 vocabulary outputslogits: [2,10]<BOS>Slot 1-0.182Slot 1, coordinate 1: -0.182-1.497Slot 1, coordinate 2: -1.4970.142Slot 1, coordinate 3: 0.142-0.524Slot 1, coordinate 4: -0.524lilySlot 2-0.430Slot 2, coordinate 1: -0.4301.780Slot 2, coordinate 2: 1.7800.057Slot 2, coordinate 3: 0.057-0.642Slot 2, coordinate 4: -0.642foundSlot 3-0.687Slot 3, coordinate 1: -0.687-0.800Slot 3, coordinate 2: -0.800-1.369Slot 3, coordinate 3: -1.369-0.325Slot 3, coordinate 4: -0.325aSlot 4-0.155Slot 4, coordinate 1: -0.155-1.096Slot 4, coordinate 2: -1.0961.367Slot 4, coordinate 3: 1.3670.689Slot 4, coordinate 4: 0.68901234567<PAD><BOS><EOS><UNK>.aballfoundlilyred4 token rows × 4 coordinates = 16 inputs. Flattening learns no weights.Next: a learned 16×8 connection matrix, ReLU, then an 8×10 classifier.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
flat = E_mlp.flatten(start_dim=1)
assert flat.shape == (2, 16)
assert torch.equal(flat[1, :4], E_mlp[1, 0])
3. The MLP forward pass · Step 44 / 88Open optional worked slide ↗

The MLP map: the hidden layer

The 16 input coordinates feed 8 hidden units. An affine operation followed by ReLU produces one hidden row per example.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
3. The MLP forward pass · Step 45 / 88Open optional worked slide ↗

The hidden layer: 16 inputs, 8 outputs

Each example enters the hidden layer as 16 numbers and leaves as 8 numbers. With two examples, the batch shape changes from [2,16] to [2,8]. The model setting h=8 chooses the hidden width. ReLU follows without changing this shape, then the classifier produces ten vocabulary scores per example.

The hidden layer: 16 inputs, 8 outputsExample 1: <BOS> lily found aOne network, applied separately to each example16 input nodesflat: [2,16]8 hidden neuronshidden: [2,8]10 vocabulary outputslogits: [2,10]<BOS>Slot 1Slot 1, coordinate 1: -0.182Slot 1, coordinate 2: -1.497Slot 1, coordinate 3: 0.142Slot 1, coordinate 4: -0.524lilySlot 2Slot 2, coordinate 1: -0.430Slot 2, coordinate 2: 1.780Slot 2, coordinate 3: 0.057Slot 2, coordinate 4: -0.642foundSlot 3Slot 3, coordinate 1: -0.687Slot 3, coordinate 2: -0.800Slot 3, coordinate 3: -1.369Slot 3, coordinate 4: -0.325aSlot 4Slot 4, coordinate 1: -0.155Slot 4, coordinate 2: -1.096Slot 4, coordinate 3: 1.367Slot 4, coordinate 4: 0.689<PAD><BOS><EOS><UNK>.aballfoundlilyredThis step: [2,16] → [2,8]2 = examples in the batch. 16 and 8 = numbers per example.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
pre_hidden = mlp.hidden_layer(flat)
assert pre_hidden.shape == (2, 8)
3. The MLP forward pass · Step 46 / 88Open optional worked slide ↗

ReLU keeps positive activations

At each of the same eight hidden neurons, ReLU keeps a positive sum and replaces a negative sum with zero. The labels show each value before and after ReLU, with batch shape still [2,8]. This is an activation within the hidden layer, not another set of eight learned neurons. The nonlinearity lets the network learn more than a single affine mapping.

ReLU keeps positive activationsExample 1: <BOS> lily found aSame weights for both batch examples16 input nodesflat: [2,16]8 hidden neuronshidden: [2,8]10 vocabulary outputslogits: [2,10]<BOS>Slot 1-0.182Slot 1, coordinate 1: -0.182-1.497Slot 1, coordinate 2: -1.4970.142Slot 1, coordinate 3: 0.142-0.524Slot 1, coordinate 4: -0.524lilySlot 2-0.430Slot 2, coordinate 1: -0.4301.780Slot 2, coordinate 2: 1.7800.057Slot 2, coordinate 3: 0.057-0.642Slot 2, coordinate 4: -0.642foundSlot 3-0.687Slot 3, coordinate 1: -0.687-0.800Slot 3, coordinate 2: -0.800-1.369Slot 3, coordinate 3: -1.369-0.325Slot 3, coordinate 4: -0.325aSlot 4-0.155Slot 4, coordinate 1: -0.155-1.096Slot 4, coordinate 2: -1.0961.367Slot 4, coordinate 3: 1.3670.689Slot 4, coordinate 4: 0.6890-0.336 → 0.0001-0.232 → 0.0002-0.715 → 0.0003-0.317 → 0.0004-0.252 → 0.00050.859 → 0.8596-0.410 → 0.0007-0.463 → 0.000<PAD><BOS><EOS><UNK>.aballfoundlilyredAt the same 8 neurons: z → max(0, z). Negative sums become zero.ReLU has no learned weights. The classifier receives the values on the right.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
hidden_mlp = torch.relu(pre_hidden)
assert hidden_mlp.shape == (2, 8)
assert hidden_mlp.ge(0).all()
3. The MLP forward pass · Step 47 / 88Open optional worked slide ↗

The MLP map: vocabulary scores

The hidden row produces one score for each of C=10 vocabulary items. With two examples, the logits have shape [2,10].

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
3. The MLP forward pass · Step 48 / 88Open optional worked slide ↗

Four input tokens, one next-token target

A logit is a raw next-token score, and the largest gives the model’s guess. Ground truth is the actual next token in the story, outside the four input slots. Both examples use the same MLP: four token IDs select four embeddings, which flatten to 16 numbers, pass through eight hidden activations, then produce ten vocabulary scores. For example 0, the context is <PAD> <BOS> lily found and the observed next token is a. For example 1, the context is <BOS> lily found a and the observed next token is red. These are initial, untrained scores. A correct guess here can happen by chance. Ground truth comes from the data, not from selecting the highest score. Softmax and the training loss follow later.

Four input tokens, one next-token targetStory: Lily found a red ball.Initial, untrained MLPExample 0: four input slots<PAD><BOS>lilyfoundMLP: 16 inputs → 8 hidden → 10 scoresCandidate tokenScore (logit)<PAD>-0.129<BOS>0.009<EOS>-0.428<UNK>0.136.0.137a-0.341ball0.199found0.119lily-0.057red0.296Model guess: redGround truth: aExample 1: four input slots<BOS>lilyfoundaMLP: 16 inputs → 8 hidden → 10 scoresCandidate tokenScore (logit)<PAD>-0.314<BOS>0.224<EOS>-0.067<UNK>0.299.0.218a-0.335ball0.030found0.365lily0.230red0.519Model guess: redGround truth: red
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
logits_mlp = mlp.vocab_head(hidden_mlp)
assert logits_mlp.shape == (2, 10)
assert torch.allclose(logits_mlp, mlp(X))
predicted_ids = logits_mlp.argmax(dim=-1)
assert y.tolist() == [vocab.stoi['a'], vocab.stoi['red']]
4. The attention forward pass · Step 49 / 88Open optional worked slide ↗

The full attention training map

The data, targets and loss stay the same. Attention changes how context reaches the hidden prediction layer. This one-layer implementation computes only the final query for each window.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
assert X.shape == (2, 4) and y.shape == (2,)
print('Both architectures predict the same two targets:', y.tolist())
Printed output
Both architectures predict the same two targets: [5, 9]
4. The attention forward pass · Step 50 / 88Open optional worked slide ↗

The attention map: input vectors

Attention adds a position vector to each token vector. E still has shape [2,4,4] before the query, key and value projections.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
4. The attention forward pass · Step 51 / 88Open optional worked slide ↗

Add a position vector to each token vector

The position table has four rows, one for each context slot. Token and position vectors both have width four, so addition preserves the width. PAD keys will be masked even though position addition can make their input row nonzero.

Add a position vector to each token vectorFinal slot of example 1c1c2c3c4Token: a-0.155-1.0961.3670.689Position: slot 41.005-0.248-0.0770.080Sum E0.850-1.3451.2890.769
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
positions = torch.arange(w)
token_rows = attention.token_embedding(X)
position_rows = attention.position_embedding(positions)
E = token_rows + position_rows[None, :, :]
assert E.shape == (2, 4, 4)
4. The attention forward pass · Step 52 / 88Open optional worked slide ↗

The attention map: queries, keys and values

Each window supplies one final query and four source keys and values. The next three calculations use the same input tensor E.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
4. The attention forward pass · Step 53 / 88Open optional worked slide ↗

The final known token supplies one query

For example 1 the final known token is a. Its four-number input row maps to three query coordinates. The unseen target red does not take part in this multiplication.

The final known token supplies one queryFinal input E[1, −1]: [ 0.850, -1.345, 1.289, 0.769 ]W_Q is 4 × 3 in row-vector notationFirst query coordinate: sum of four products0.075 + -0.404 + 0.054 + -0.230 = -0.505Query q: [ -0.505, 0.965, 0.077 ]
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
q = attention.W_Q(E[:, -1:, :])
query_terms = E[1, -1] * attention.W_Q.weight[0]
assert q.shape == (2, 1, 3)
assert torch.allclose(query_terms.sum(), q[1, 0, 0])
4. The attention forward pass · Step 54 / 88Open optional worked slide ↗

Each source has a matching key

All four source rows map to keys of width three. Query and key widths match because we will take their dot products. The same W_K applies at every source position and in both examples.

Each source has a matching keySource in example 1coord. 1coord. 2coord. 3<BOS>-0.147-0.2800.006lily0.0030.9420.597found-1.4840.621-0.258a0.027-0.7020.615
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
K = attention.W_K(E)
assert K.shape == (2, 4, 3)
print('Row-vector W_K shape:', tuple(attention.W_K.weight.T.shape))
Printed output
Row-vector W_K shape: (4, 3)
4. The attention forward pass · Step 55 / 88Open optional worked slide ↗

Each source also has information to send

W_V maps each source to a two-number value. These coordinates carry information into the weighted message. They do not enter the query-key dot product.

Each source also has information to sendSource in example 1coord. 1coord. 2<BOS>0.418-0.427lily-1.0600.357found-0.523-1.135a1.385-0.462
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
V = attention.W_V(E)
assert V.shape == (2, 4, 2)
print('Row-vector W_V shape:', tuple(attention.W_V.weight.T.shape))
Printed output
Row-vector W_V shape: (4, 2)
4. The attention forward pass · Step 56 / 88Open optional worked slide ↗

The attention map: source weights

Query-key scores determine where to read. Scaling, padding masks and softmax produce one row of four weights per example.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
4. The attention forward pass · Step 57 / 88Open optional worked slide ↗

Compare the query with all four keys

Each score is a three-term dot product divided by √3. The result has one row of four source scores per example. These are source-match scores, not vocabulary logits.

Compare the query with all four keysOne match, source found: 0.749 + 0.599 + -0.020Dot product 1.328 ÷ √3 = 0.767SourceBOSlilyfoundaRaw q·k-0.1950.9541.328-0.644Scaled-0.1130.5510.767-0.372
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
raw_scores = q @ K.transpose(-2, -1)
scores = raw_scores / math.sqrt(d_k)
dot_terms = q[1, 0] * K[1, 2]
assert torch.allclose(dot_terms.sum(), raw_scores[1, 0, 2])
assert scores.shape == (2, 1, 4)
4. The attention forward pass · Step 58 / 88Open optional worked slide ↗

Padding must receive zero attention weight

Example 0 has PAD in its first slot, so that score becomes −∞. All slots of example 1 are real tokens. Since these are final queries, their windows contain no future columns.

Padding must receive zero attention weightExample 0PADBOSlilyfoundBefore mask0.0420.351-0.223-0.150After mask−∞0.351-0.223-0.150exp(−∞) = 0, so the padded source contributes zero weight.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
pad_mask = X[:, None, :].eq(vocab.pad_id)
masked_scores = scores.masked_fill(pad_mask, float('-inf'))
assert torch.isneginf(masked_scores[0, 0, 0])
4. The attention forward pass · Step 59 / 88Open optional worked slide ↗

Softmax turns four source scores into weights

Subtract the largest score, exponentiate and divide by the row sum. This softmax runs over context slots. Later, a different softmax will run over vocabulary items.

Softmax turns four source scores into weightsExample 0PADBOSlilyfoundShifted exponentials0.0001.0000.5630.606Normalized A0.0000.4610.2600.279Divide by 2.169; the normalized weights sum to 1.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
source_exp = (masked_scores - masked_scores.amax(-1, keepdim=True)).exp()
A = source_exp / source_exp.sum(-1, keepdim=True)
assert torch.allclose(A, masked_scores.softmax(-1))
assert A[0, 0, 0] == 0 and torch.allclose(A.sum(-1), torch.ones(2, 1))
4. The attention forward pass · Step 60 / 88Open optional worked slide ↗

The attention map: the message

Each source weight multiplies a two-coordinate value row. Their sum gives one message of width dᵥ=2 per example.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
4. The attention forward pass · Step 61 / 88Open optional worked slide ↗

Each weight scales a complete value row

For example 1, multiply each two-number value by its source weight and add the four contributions. The result is one two-number message. Weighting does not change the value width.

Each weight scales a complete value rowSourceWeightValue 1Value 2Contribution 1Contribution 2<BOS>0.1630.418-0.4270.068-0.070lily0.317-1.0600.357-0.3360.113found0.394-0.523-1.135-0.206-0.447a0.1261.385-0.4620.175-0.058Add the four contribution rows: message = [ -0.299, -0.462 ]
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
contributions = A[1, 0, :, None] * V[1]
message = A @ V
assert message.shape == (2, 1, 2)
assert torch.allclose(contributions.sum(0), message[1, 0])
4. The attention forward pass · Step 62 / 88Open optional worked slide ↗

The attention map: the output projection

Wₒ maps the two-coordinate message into four coordinates. The residual addition then combines it with the original final input row.

trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
4. The attention forward pass · Step 63 / 88Open optional worked slide ↗

Project two message coordinates into four

The message has width two, but the original input has width four. The learned W_O is 2×4 in row-vector notation. Every output coordinate can combine both message coordinates.

Project two message coordinates into fourMessage: [ -0.299, -0.462 ]W_O: value → modelc1c2c3c4Value coordinate 1-0.4000.2510.146-0.567Value coordinate 2-0.157-0.499-0.191-0.086Update: [ 0.192, 0.155, 0.045, 0.209 ]
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
update = attention.W_O(message).squeeze(1)
assert update.shape == (2, 4)
assert torch.allclose(message[1, 0] @ attention.W_O.weight.T, update[1])
4. The attention forward pass · Step 64 / 88Open optional worked slide ↗

Add the update to the original final row

The residual addition now combines two vectors with the same width. The updated row enters an eight-unit ReLU layer and then the ten-class vocabulary head, just as in the MLP path.

Add the update to the original final rowExample 1c1c2c3c4Original final row0.850-1.3451.2890.769Attention update0.1920.1550.0450.209Updated final row1.042-1.1891.3340.9794 updated coordinates → 8 ReLU units → 10 vocabulary logits
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
final = E[:, -1, :] + update
hidden_att = torch.relu(attention.hidden_layer(final))
logits_att = attention.vocab_head(hidden_att)
assert torch.allclose(logits_att, attention(X), atol=1e-6)
assert torch.allclose(logits_att, attention.forward_details(X)['logits'], atol=1e-6)
5. Loss and learning · Step 65 / 88Open optional worked slide ↗

The training map: predictions and targets

Return to the MLP’s two rows of vocabulary logits. The observed IDs y=[5,9], meaning a and red, meet those logits at the loss.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
5. Loss and learning · Step 66 / 88Open optional worked slide ↗

Ten logits become ten next-token probabilities

For the MLP example with target red, apply vocabulary softmax. The denominator includes all ten classes. These values come from the seeded model before training. Even a correct guess here can occur by chance.

Ten logits become ten next-token probabilitiesWordProbabilityexp(z − max)<PAD>6.28%0.435<BOS>10.75%0.745<EOS>8.04%0.556<UNK>11.59%0.803.10.69%0.740WordProbabilityexp(z − max)a6.15%0.425ball8.86%0.613found12.38%0.857lily10.82%0.749red14.44%1.000For each word: exp(z − max) ÷ 6.924 = probability.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
word_exp = (logits_mlp[1] - logits_mlp[1].max()).exp()
p = word_exp / word_exp.sum()
assert torch.allclose(p, logits_mlp[1].softmax(-1))
assert torch.allclose(p.sum(), torch.tensor(1.0))
5. Loss and learning · Step 67 / 88Open optional worked slide ↗

Observed target and model prediction

Argmax returns the largest-probability vocabulary ID. The target remains red because the corpus contains red after this prefix. Training does not replace that observed target with the model’s guess.

Observed target and model predictionKnown input: BOS lily found aModel’s largest probability: redObserved next token: red (ID 9)p(red) = 0.1444; the loss measures this observed target.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
guess_id = int(p.argmax())
target_id = int(y[1])
print('Guess:', words[guess_id], 'target:', words[target_id])
print('Probability of target:', float(p[target_id].detach()))
Printed output
Guess: red target: red
Probability of target: 0.1444295048713684
5. Loss and learning · Step 68 / 88Open optional worked slide ↗

Per-example loss and batch mean

For each example, cross-entropy is −log of the probability assigned to its observed target. PyTorch accepts raw logits and performs log-softmax internally. The optimizer uses the mean across B=2 examples.

Per-example loss and batch meanExampleTargetTarget probability−log(probability)0a0.0702.6621red0.1441.935Batch loss = (2.662 + 1.935) / 2 = 2.298
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
losses = F.cross_entropy(logits_mlp, y, reduction='none')
loss = losses.mean()
manual_loss = -logits_mlp.log_softmax(-1)[torch.arange(B), y]
assert torch.allclose(losses, manual_loss)
5. Loss and learning · Step 69 / 88Open optional worked slide ↗

The training map: backpropagation

The batch mean loss supplies gradients for the learned parameters. Backpropagation computes these gradients before any weight changes.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
5. Loss and learning · Step 70 / 88Open optional worked slide ↗

Gradients from backpropagation

For vocabulary weight W[j,r], the batch-mean gradient is the mean of (p_r − 1[y=r]) × hidden_j. Here we inspect the weight feeding the red logit from hidden unit 0.

Gradients from backpropagationWeight: hidden unit 0 → red logitFor each example: (p(red) − target-is-red) × hidden[0]Average the two contributions because B = 2.Autograd = manual gradient = 0.011
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
mlp.zero_grad(set_to_none=True)
loss.backward()
manual_grad = ((logits_mlp.softmax(-1)[:, 9] - y.eq(9).float()) * hidden_mlp[:, 0]).mean()
grad = mlp.vocab_head.weight.grad[9, 0]
assert torch.allclose(grad, manual_grad)
5. Loss and learning · Step 71 / 88Open optional worked slide ↗

The training map: updating parameters

The optimizer uses the gradients to change the learned parameters. The next batch repeats the forward pass with those updated weights.

trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
5. Loss and learning · Step 72 / 88Open optional worked slide ↗

One optimizer step changes the stored weights

For this arithmetic demonstration we use SGD with learning rate 0.1: new weight = old weight − 0.1×gradient. The measured TinyStories runs use validation-tuned AdamW instead.

One optimizer step changes the stored weights-0.022 − 0.1 × (0.011) = -0.023old weight learning rate × gradient new weightOn this batch: loss 2.298 → 2.239The step updates all trainable parameters, not just this one weight.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
old_weight = mlp.vocab_head.weight[9, 0].detach().clone()
before_loss = float(loss.detach())
optimizer = torch.optim.SGD(mlp.parameters(), lr=0.1)
optimizer.step()
new_weight = mlp.vocab_head.weight[9, 0].detach().clone()
after_loss = float(F.cross_entropy(mlp(X), y).detach())
assert torch.allclose(new_weight, old_weight - 0.1*grad)
5. Loss and learning · Step 73 / 88Open optional worked slide ↗

Many batches, with validation between checkpoints

The actual trainer samples 512 training windows per step, computes the loss, backpropagates and updates parameters. Validation selects the best saved checkpoint. The test set does not choose a checkpoint.

Many batches, with validation between checkpointsStageData / actionTraining step512 sampled windows; forward, loss, backward, AdamWValidationEvaluate frozen weights; save lower validation lossFinishRestore best validation checkpoint; score test once
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full ATTENTION map?
trainV bypasses scoringkeep the final input rowtarget y [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookup+ positionE [B,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
benchmark = evidence['benchmark']
print('Maximum steps:', 6000, 'batch size:', 512)
print('Maximum target presentations:', 6000*512)
print('Checkpoint rule: best validation loss')
Printed output
Maximum steps: 6000 batch size: 512
Maximum target presentations: 3072000
Checkpoint rule: best validation loss
6. Inference · Step 74 / 88Open optional worked slide ↗

Held-out scoring and free-running generation

Held-out scoring uses observed X,y pairs and records loss. Generation has a prompt but no observed next target. Both freeze parameters. Neither calls backward or optimizer.step.

Held-out scoring and free-running generationModeInputObserved target?Parameter updates?TrainingTraining prefixYesYesHeld-out scoringTest prefixYesNoGenerationGrowing promptNoNo
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
trainy [B]TinyStoriescomplete documentsSplit storiestrain / validation / testTokenize +assign IDstrain-only vocabularyIDs + boundaries<BOS> … <EOS>Context and targetX [B,w] y [B]Token lookupE [B,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Cross-entropy(z, y)one target per windowloss.backward()gradients for θoptimizer.step()update learned θNext training batchreuse updated θThe next batch repeats the forward pass with updated parameters in every learned layer.
Python · run after the previous step
mlp.eval()
with torch.inference_mode():
    frozen_loss = F.cross_entropy(mlp(X), y)
print('Illustrative scoring loss:', float(frozen_loss))
Printed output
Illustrative scoring loss: 2.2394907474517822
6. Inference · Step 75 / 88Open optional worked slide ↗

The complete MLP generation loop

Load the same vocabulary and trained parameters. Prepare the latest window, predict, choose a token and append it. Stop on EOS or the generation budget. The only changing state is the history.

load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
print('Generation uses B=1; the parameters stay fixed.')
Printed output
Generation uses B=1; the parameters stay fixed.
6. Inference · Step 76 / 88Open optional worked slide ↗

Attention uses the same generation loop

The new final token supplies the next query. This notebook recomputes the current window and has no KV cache. Position indices describe the slots in that current window.

load weights for every learned layerV bypasses scoringkeep the final input rowPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookup+ positionE [1,w,d]Final q, allK and Vq [B,1,dₖ] · K,V: w rowsqKᵀ / √dₖscores [B,1,w]Mask PAD · softmaxweights A [B,1,w]A @ Vmessage [B,1,dᵥ]Output map Wₒupdate [B,1,d]Add originalfinal rowe′ = e + updateFinalupdated vector[B,d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
attention.eval()
print('Final query only; all available keys and values; no KV cache.')
Printed output
Final query only; all available keys and values; no KV cache.
6. Inference · Step 77 / 88Open optional worked slide ↗

The generation map: preparing the prompt

Generation uses the saved tokenizer and the latest w IDs. Our worked example follows the frozen MLP with B=1 and no observed next target.

load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
6. Inference · Step 78 / 88Open optional worked slide ↗

Prepare a prompt using the same rules

Use the training tokenizer and vocabulary on the unfinished prompt. Prepend BOS once, and leave EOS out because generation has not finished.

Prepare a prompt using the same rulesPrompt: Lily foundToken IDs: [8, 7]With BOS: [1, 8, 7] (BOS lily found)EOS is absent because the prompt is unfinished.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
prompt = 'Lily found'
prompt_ids = vocab.encode_tokens(tokenize(prompt), boundaries=False)
history = [vocab.bos_id] + prompt_ids
assert history == [1, 8, 7]
6. Inference · Step 79 / 88Open optional worked slide ↗

The prompt fills the same four input slots

Keep the most recent w=4 IDs from the history. This prompt has only three IDs including BOS, so one PAD ID fills the unused slot on the left.

The prompt fills the same four input slotsPrompt: Lily foundHistory: BOS lily foundModel input: [PAD BOS lily found] = [0, 1, 8, 7]B = 1; w = 4; no target supplied; parameters are frozen.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
kept = history[-w:]
context_ids = [vocab.pad_id] * (w-len(kept)) + kept
inference_X = torch.tensor([context_ids])
assert inference_X.tolist() == [[0, 1, 8, 7]]
6. Inference · Step 80 / 88Open optional worked slide ↗

A frozen model scores the next token

The MLP reads the prepared window and returns ten vocabulary scores. inference_mode disables gradient tracking for this forward pass, and no optimizer updates the parameters.

A frozen model scores the next tokenID / tokenLogit0 <PAD>-0.1381 <BOS>-0.0002 <EOS>-0.4323 <UNK>0.1244 .0.125ID / tokenLogit5 a-0.2956 ball0.1897 found0.1078 lily-0.0669 red0.333
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
with torch.inference_mode():
    next_logits = mlp(inference_X)[0]
assert next_logits.shape == (C,)
6. Inference · Step 81 / 88Open optional worked slide ↗

The generation map: choosing a token

Vocabulary softmax turns the scores into probabilities. Greedy decoding or sampling chooses the next token ID.

load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
6. Inference · Step 82 / 88Open optional worked slide ↗

Special input tokens are excluded from generation

Set the scores of PAD, BOS and UNK to negative infinity before softmax. Their probabilities become zero, while EOS remains available as a stopping token.

Special input tokens are excluded from generationToken<PAD><BOS><UNK><EOS>Masked score−∞−∞−∞-0.432Probability0.00%0.00%0.00%9.04%Softmax still uses all 10 vocabulary entries.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
blocked_ids = [vocab.pad_id, vocab.bos_id, vocab.unk_id]
with torch.inference_mode():
    next_logits[blocked_ids] = float('-inf')
    next_p = next_logits.softmax(-1)
assert next_p[blocked_ids].eq(0).all()
assert torch.allclose(next_p.sum(), torch.tensor(1.0))
6. Inference · Step 83 / 88Open optional worked slide ↗

Greedy decoding and sampling make different choices

Greedy chooses the ID with the highest probability, while sampling uses the cumulative probabilities (CDF). The fixed draw 0.8 selects the first CDF entry at least 0.8, making this example reproducible.

Greedy decoding and sampling make different choicesWordProbabilityCDF<PAD>0.00%0.000<BOS>0.00%0.000<EOS>9.04%0.090<UNK>0.00%0.090.15.78%0.248WordProbabilityCDFa10.37%0.352ball16.83%0.520found15.51%0.675lily13.04%0.806red19.43%1.000Greedy: red; fixed sampling draw 0.8: lily
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
greedy_id = int(next_p.argmax())
cdf = next_p.cumsum(0)
sample_id = int(torch.searchsorted(cdf, torch.tensor(0.8)))
6. Inference · Step 84 / 88Open optional worked slide ↗

The generation map: extending the history

EOS ends generation. Otherwise append the chosen ID and prepare the newest window for another forward pass.

load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
The highlighted boxes locate the next worked example. The layout stays the same as we move through the model.
6. Inference · Step 85 / 88Open optional worked slide ↗

A selected token extends the history unless it is EOS

Use the greedy ID for this demonstration, and check whether it is EOS before changing the history. If it is EOS, stop generation, otherwise append it and prepare the next window.

A selected token extends the history unless it is EOSChosen ID: 9 (red)Before: BOS lily foundAfter: BOS lily found redContinue with the newest four IDs as the next input.
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
history_before_choice = history.copy()
chosen = greedy_id
if chosen != vocab.eos_id:
    history.append(chosen)
assert history_before_choice == [1, 8, 7]
6. Inference · Step 86 / 88Open optional worked slide ↗

The chosen token becomes input on the next step

Repeat window preparation, scoring and token selection, stopping on EOS or after four steps. This table repeats generation from the original prompt, with all parameters frozen.

The chosen token becomes input on the next stepStepPrepared inputChosen token1PAD BOS lily foundred2BOS lily found redred3lily found red red.4found red red .found
Same figure as the lecture. Numeric values are computed from the code below.
Where are we on the full MLP map?
load weights for every learned layerPrompt textknown words onlySaved tokenizersame vocabulary + <BOS>Crop, then left-padlast w IDs → X [1,w]Saved parametersθ stays fixedToken lookupE [1,w,d]Join ordered rowsflatten [B,w·d]Affine + ReLUhidden [B,h]Vocabulary affinelogits z [B,C]Vocabulary softmaxz / temperatureChoose a tokensample or greedyIs it <EOS>?yes: finish · no: appendAppend to historyrepeat if budget remains
Python · run after the previous step
history = history_before_choice.copy()
frozen = {n:p.detach().clone() for n,p in mlp.named_parameters()}
generation_trace = []
with torch.inference_mode():
    for generation_step in range(4):
        kept = history[-w:]
        x_next = torch.tensor([[vocab.pad_id]*(w-len(kept)) + kept])
        z_next = mlp(x_next)[0]
        z_next[[0, 1, 3]] = float('-inf')
        chosen = int(z_next.argmax())
        generation_trace.append((x_next[0].tolist(), chosen))
        if chosen == vocab.eos_id: break
        history.append(chosen)
assert all(torch.equal(p, frozen[n]) for n,p in mlp.named_parameters())
7. The real experiment · Step 87 / 88Open optional worked slide ↗

The trained TinyStories comparison

The saved experiment uses the same documents, vocabulary, targets and 64-token context for both models. Each model has validation-tuned optimizer settings. These are three-seed test results, separate from our ten-token arithmetic example.

The trained TinyStories comparisonModelTest perplexityParametersmlp51.41 ± 0.442,332,832attention31.31 ± 0.301,321,12039.1% lower perplexity for attention in this saved experiment.
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
comparison = benchmark['aggregate']
for name in ['mlp', 'attention']:
    stats = comparison[name]['test_perplexity']
    print(name, stats['mean'], '+/-', stats['sample_std'])
Printed output
mlp 51.414434800867234 +/- 0.44307967119606395
attention 31.31463254211675 +/- 0.3047408118077636
7. The real experiment · Step 88 / 88Open optional worked slide ↗

The same steps in the runnable notebooks

This notebook executes the small calculation behind every figure. Notebook 1 loads the corpus and trains the MLP. Notebook 3 trains attention. Notebook 4 checks the saved comparison and studies representations.

The same steps in the runnable notebooksCompanionPurposeNotebook 1Corpus preparation and MLP trainingNotebook 3Attention operations and trainingNotebook 4Measured comparison and representationsNotebook 5Every calculation and figure in this walkthrough
Same figure as the lecture. Numeric values are computed from the code below.
Python · run after the previous step
assert B == 2 and w == 4 and C == 10
print('Toy example: 7 targets; worked batch: 2 examples.')
print('Real experiment: 969,160 training windows; context 64; vocabulary 4,000.')
Printed output
Toy example: 7 targets; worked batch: 2 examples.
Real experiment: 969,160 training windows; context 64; vocabulary 4,000.