Notation used on this pagesymbols, meanings, shapes
    01

    Language tasks ask for different outputs, but all use text to make predictions.

    What can a network do with text?

    Before choosing a network, decide what its output should mean. A spam detector assigns one label to an email. An entity tagger attaches labels to words. Translation, summarization, and question answering return text. The examples below illustrate these tasks; they are not outputs of the character model we build later.

    Email

    You have won a free phone! Click here to claim your prize.

    Label

    Spam

    The whole email gets one label: spam or not spam.

    Review

    I loved this film. I would watch it again.

    Sentiment label

    Positive

    For this task, choose among positive, negative, and neutral.

    Here, the labels belong to individual words.

    Maya
    Person
    joined
    Other
    IIT
    Organization
    Gandhinagar
    Organization

    The output has a label for each word, rather than one label for the sentence.

    English input

    The child is reading.

    One Hindi translation

    बच्चा पढ़ रहा है।

    The network must produce another sequence of words.

    Notice

    The library will close at 5 pm today for repairs. It will reopen at 9 am tomorrow.

    One summary

    Library closed from 5 pm today until 9 am tomorrow.

    The output is shorter text that preserves the main information.

    Context

    Maya left her notebook in the library.

    Question

    Where is Maya's notebook?

    Answer

    In the library.

    TaskWhat comes out?
    Spam detection, sentimentOne label for the input
    Finding entitiesLabels attached to words
    Translation, summarization, answeringA sequence of text

    How could a network produce text of any length?

    The student opened the laptop and …

    started coding.
    checked email.
    joined class.
    closed a browser tab.

    After “The student opened the laptop and”

    Illustrative probabilities. “Other” groups the remaining words.

    Choose a word, append it to the text, then predict again.

    GitHub preview of balasahebgulave's Dataset-indian-names repository: Indian_Names.csv, with entries aabid, aabida, aachal, and aadesh.
    Indian_Names.csv on GitHub

    We’ll learn from a list of Indian given names, with one name per row.

    aabid, aabida, aachal, …

    Then we’ll generate a name: choose the next letter, append it, and repeat until the name ends.

    We can produce a sequence by repeatedly predicting one next token. A token is a unit of text. We will start with characters, so every prediction chooses a letter or an ending marker. This keeps the vocabulary small enough to inspect each probability.

    02

    With a small character vocabulary, we can inspect every input, score, and probability.

    Generate a name

    We will train a network on Indian given names. One letter is one token. A boundary symbol fills the space before a name and marks where it ends, so our vocabulary is the alphabet plus this boundary. At each step, the network reads three character IDs and returns a probability for every possible next character. The bars for aab come from the trained model bundled with this page. The network does not need to be certain. It should assign useful probability to characters that often follow similar contexts.

    A name in our data: aabid

    Read a a b. Predict the next character: i

    During training, the name tells us what came next.

    During generation, the model chooses a character and continues from there.

    One boundary and the alphabet

    The vocabulary is $\mathcal V=\{\text{-},a,\ldots,z\}$, with one output class per token.

    Names in the source list

    aabidaaravavninaveenvanizeel

    Run the snippets in order. We use the same 27-token vocabulary throughout.

    import torch
    from torch import nn
    from torch.nn import functional as F
    vocab = ["-"] + list("abcdefghijklmnopqrstuvwxyz")
    stoi = {c: j for j, c in enumerate(vocab)}

    stoi maps a character to an integer: "-" → 0, "a" → 1, "b" → 2.

    These snippets construct a fresh model. The numerical worksheets show a saved trained model; creating a layer does not load those trained values.

    a a b → ?

    One probability per possible next character.

    One forward pass is a classification problem with one class for every possible next character.

    The boundary token does two jobs: it pads the opening window and tells us when a generated name has ended.

    03

    The MLP needs a fixed-width input. How many earlier characters should we include?

    Choose a context length

    A one-character window remembers very little. Using the whole prefix brings a different problem: its length changes from one training row to the next. A plain MLP needs another design choice to handle that changing input size. We start with three previous characters. Boundary padding gives even the first prediction a full three-character window. Sliding it across aabid makes six context-target pairs: five letters followed by the ending boundary. We split the names into training and held-out sets before making the windows, so examples from one name never cross the split.

    w = 1

    One previous character

    w = 3

    Three previous characters

    w = all

    A growing prefix

    For aabid, the window a a b is the input. The next character i is the training target.

    We have read a a b. How likely is i next?

    $$\pProb{p}_{\pParam{\theta}}(\pTarget{t_4=\mathtt{i}}\mid \pIn{t_1=\mathtt{a},\ t_2=\mathtt{a},\ t_3=\mathtt{b}})$$
    $\pProb{p}$: probability
    How likely the model thinks this next character is.
    $\pParam{\theta}$: learned parameters
    The model's current embedding table, weights, and biases.
    $\pTarget{t_4=\mathtt{i}}$: proposed next character
    Character 4 is “i”. Here, “i” is a letter, not an index.
    $\pIn{t_1,t_2,t_3}$: known context
    Characters 1, 2, and 3 are “a”, “a”, and “b”.

    Read $\mid$ as “given”: the probability of “i” next, given “a a b” already read.

    The same rule at any position:

    $$\pProb{p}_{\pParam{\theta}}(\pTarget{t_i}\mid\pIn{t_{i-3},t_{i-2},t_{i-1}})$$

    Probability under the learned parameters $\pParam{\theta}$ of the next token $\pTarget{t_i}$ at position $i$, given the three previous tokens. Here $i$ is a position index. The compact notation $\pIn{t_{i-3:i-1}}$ names those same three tokens, including both endpoints.

    ctx = torch.tensor([[stoi[c] for c in "aab"]])
    target = torch.tensor([stoi["i"]])
    print(ctx.shape, target.shape)

    ctx has shape [1, 3]: one example, three known character IDs.

    target has shape [1]: the next character ID for that example.

    Integer tensors select embedding rows. The extra outer brackets keep the batch dimension, even for one example.

    One window at a time

    w = 3
    padded = [stoi["-"]] * w + [stoi[c] for c in "aabid-"]
    X = torch.tensor([padded[j:j+w] for j in range(6)])
    y = torch.tensor(padded[w:])

    X has shape [6, 3]. Each row holds one context. y has shape [6]. Each entry is that row's observed target.

    The first three boundary IDs pad the opening. The final boundary is the target that teaches stopping.

    In Python, j:j+w includes position j and stops before j+w.

    Each name of length $L$ contributes $L+1$ context-target pairs.

    Rows made from the cleaned data

    The fixed window gives every MLP input the same size, including the first one.

    04

    Characters are categorical symbols. An MLP needs rows of numbers.

    Characters to numbers

    A one-hot vector identifies a character with one active coordinate. Any two different characters have zero dot product. A network can still learn useful features from those IDs. An embedding layer makes that learned mapping explicit: each ID selects a trainable row. We will first draw those rows as points, then see how the same idea appears in other domains, and return to our name model.

    Selected columns of a 27-wide one-hot row

    Different characters occupy different coordinates.
    one_hot = F.one_hot(ctx, num_classes=27)
    print(one_hot.shape)  # [1, 3, 27]

    Each of the three character IDs becomes a row with one 1 and twenty-six 0s.

    This representation contains no learned parameters. The ID simply chooses where the 1 goes.

    Compare a with b

    Every coordinate-wise product is zero.

    Compare b with i

    One-hot dot products cannot tell us which characters behave similarly.

    a and i are both vowels.

    b is a consonant.

    Could a learned representation make a useful distinction available to the next-character predictor?

    One-hot IDs give each character a separate identity. We will let training choose a short row of numbers for each character.

    A network can learn from one-hot inputs too. The limitation belongs to the raw one-hot geometry, not to every model that uses it.

    The vector for a is a row of two coordinates. The plot shows the same two numbers as a point.

    This row comes from our saved character model. Values here are rounded. Next we will ask how training chose them.

    Here a and i are closer to each other than to b.

    Our toy training constrains coordinate 1 to be positive for vowels and negative for consonants. This picture is not evidence that an unconstrained model discovered that rule.

    $$\pSymbol{c}=\pToken{\texttt{"a"}}$$

    c is a placeholder for whichever character we choose.
    "a" is the actual character chosen in this example.

    The quotation marks identify a character value. They are not part of the name.

    In another example, c could have the value "b" or "i".

    $$\pIndex{\operatorname{id}}(\pToken{\texttt{"a"}})=\pIndex{1}$$

    id is our character-to-integer mapping. It gives the character a the ID 1. The quoted "a" is the character we give to that mapping.

    IDs start at 0 in our code. 1 identifies a's table row, regardless of where a occurs in a name.

    $$\pIn{E_{\text{tok}}}$$

    Capital E names the whole embedding table. The subscript tok is short for token. Here, each token is a character.

    Three of 27 rows are shown. Each embedding contains two numbers. Character and ID just label its row.

    $$\pIn{E_{\text{tok}}}[\pIndex{1}]$$

    Take the embedding table. Select its row with ID 1. The square brackets [ ] mean “look up this row”.

    The two blue coordinates form the embedding. The purple 1 selects them and is not part of the vector. This is selection, not multiplication.

    $$\pIn{e}_{\pToken{\texttt{"a"}}}$$

    Lowercase e names one embedding vector. The subscript "a" tells us which character it represents.

    The two saved coordinates, rounded to two decimal places.

    Capital E with subscript tok is the whole table. Lowercase e with subscript "a" is one retrieved row. Reading a row leaves the table unchanged.

    $$\pIn{e}_{\pSymbol{c}}=\pIn{E_{\text{tok}}}[\pIndex{\operatorname{id}}(\pSymbol{c})]$$
    $\pSymbol{c}$: the character placeholder
    Our example gives it the value "a".
    $\pIndex{\operatorname{id}}(\pSymbol{c})$: its integer ID
    The mapping gives the character a the ID 1.
    $\pIn{E_{\text{tok}}}$: the whole table
    All 27 embedding rows.
    $\pIn{e}_{\pSymbol{c}}$: the retrieved vector
    The embedding row for that character.

    Read right to left: find the character's ID, look up that row in the table, and call the returned vector its embedding.

    [ ] selects the row. = says the expressions on both sides name the same vector.

    Start with initial coordinates. Make a prediction, measure its error, and update the table along with the predictor.

    The objective is better next-character prediction. We do not directly tell every pair of characters how far apart to sit.

    Our vowel-sign constraint is an extra teaching choice. Ordinary embedding layers need no named axes or such constraint.

    embedding = nn.Embedding(27, 2)
    char_id = torch.tensor([1])  # id("a") = 1
    e_a = embedding(char_id)
    print(e_a.shape)  # [1, 2]

    embedding.weight is the table $\pIn{E_{\text{tok}}}$, with 27 rows of two numbers.

    The character "a" has ID 1. embedding(char_id) selects that row.

    e_a holds one embedding with two numbers. The [1, 2] here describes its shape, not its coordinate values.

    This code creates a fresh table. Its coordinates differ from the saved trained table shown in the worksheet.

    stoi["a"] is our code for id("a"), which gives 1. The integer IDs stay fixed while training changes the embedding coordinates. In this example, nn.Embedding takes integer indices and returns one row per index. See the PyTorch Embedding documentation for the input and output shapes.

    ctx means context. Here the characters "a", "a", "b" have IDs 1, 1, 2.

    ctx = torch.tensor([[1, 1, 2]])  # one example: "a", "a", "b"
    e = embedding(ctx)
    print(e.shape)  # [1, 3, 2]

    ctx has shape [1, 3]: 1 example, 3 character IDs. The outer brackets keep the example dimension.

    e has shape [1, 3, 2]: 1 example, 3 positions, 2 embedding coordinates per position.

    We reuse the same embedding table. The two occurrences of ID 1 select the same row we just retrieved for "a".

    Construct the embedding layer once, not for each input. Calling it selects rows, while training changes their values. We leave the boundary row trainable, so we do not set padding_idx.

    a, b, and i are highlighted. These points come from the trained table; we have not invented semantic coordinates.

    The context a a b

    Each ID selects one trainable row. Three IDs return three embedding rows.

    An embedding is a numerical vector used to represent an object so that a model can work with it.

    A character or word ID

    Select a stored row from a learned table.

    One token ID gives one vector.

    A passage, image, or signal

    Run an encoder over the input to compute a vector.

    Many input values can give one vector.

    A representation is useful when it preserves information the task needs. Its coordinates usually have no simple English names.

    the king ruled the kingdom

    the queen ruled the kingdom

    King and queen can occur in similar surroundings. Across many sentences, these contexts supply a learning signal.

    Word2vec trains word vectors by predicting nearby words. Similar contexts can encourage related representations.

    These two sentences illustrate the idea; they are not enough to train useful word vectors.

    The observed neighbour provides the target. Prediction errors update both the selected embedding row and the predictor.

    This direction is called skip-gram. Its context can include words on either side, unlike left-to-right name generation.

    Each point represents one word vector: man, woman, king, queen.

    A hand-drawn two-dimensional illustration, not measured Word2vec vectors. Real vectors have many coordinates, usually without named meanings.

    The move from man to king resembles the move from woman to queen.

    In a learned space, some relationships can appear as similar vector offsets. The diagram makes that pattern exact just to show the idea.

    $$\pIn{e_{\text{woman}}}+\pParam{\big(e_{\text{king}}-e_{\text{man}}\big)}\approx\pTarget{e_{\text{queen}}}$$

    Start at the vector for woman. Add the displacement from man to king. Look for a word vector near the result: queen. The symbol ≈ means approximately.

    This is the familiar king − man + woman ≈ queen example, read as a move through the vector space.

    This is an analogy to test after training, not Word2vec's training rule. It is not guaranteed for every model or word pair.

    man, woman, king, queen = torch.tensor(
        [[1., 1.], [3., 1.], [1., 3.], [3., 3.]])
    candidate = king - man + woman
    print(candidate)  # tensor([3., 3.])

    The move is [0, 2]. Starting at [3, 1] reaches [3, 3]: the toy queen vector.

    These are the diagram's invented coordinates. We have not trained Word2vec here; its learned vectors only approximately support some such relationships.

    The original Word2vec paper discusses the king–man+woman analogy. We use a hand-chosen 2D parallelogram to explain an offset, not to claim that a real model uses a “gender coordinate” or a “royalty coordinate.” Learned associations depend on the corpus and can also reflect its biases. Skip-gram predicts surrounding words from a center word; the other Word2vec architecture, CBOW, predicts a center word from surrounding words. Neither is trained by directly imposing this analogy.

    She deposited money at the bank.

    They sat on the river bank.

    Ordinary Word2vec looks up the same stored bank vector in both sentences.

    Its training contexts influence that row, but the lookup does not read the current sentence. Part 2 will make a representation depend on the context of this occurrence.

    A search system might compare a query about borrowing money with a passage about loan rates.

    It needs an encoder trained for that comparison. The query and document vectors must live in a compatible space.

    doc_embedding = nn.Embedding(4, 2)
    doc_ids = torch.tensor([[1, 2, 3]])  # bank lends money
    doc_rows = doc_embedding(doc_ids)
    e_doc = doc_rows.mean(dim=1)

    One passage with three word vectors, [1, 3, 2], becomes one document vector, [1, 2].

    An average is a simple baseline. It gives “dog bites man” and “man bites dog” the same vector when they use the same lookup rows.

    This fresh, untrained table only shows the shapes. Learned document encoders can preserve more context and order; this mean is not a Doc2vec implementation.

    The image vector represents this image. A new image can receive a vector by running the same encoder.

    This is the encoder's role, not a measured model output for this scene. A vocabulary lookup alone cannot encode arbitrary new images.

    pixels = torch.tensor([[[[0., 1.], [1., 0.]]]])
    image_encoder = nn.Conv2d(1, 4, kernel_size=2)
    e_image = image_encoder(pixels).flatten(1)
    print(e_image.shape)  # [1, 4]

    The input is a separate 2 × 2 grayscale toy, shaped [batch, channels, height, width]. Four learned filters produce four numbers.

    These fresh filters are untrained. They do not encode the photograph above. The vision lessons explain how filters or attention learn useful image features.

    Count the mugs

    The representation must preserve how many mugs appear.

    Changing an orange mug to blue should not change the count.

    Identify the mug's colour

    The representation must preserve colour.

    Ignoring that change would lose the information the task needs.

    A useful similarity depends on the training objective. There is no single geometry that captures every possible task equally well.

    An encoder can represent a window of electricity use, movement, or temperature readings.

    The window vector can feed a classifier or a forecaster. For forecasting, encode only observations available at prediction time.

    For recognizing repeating behaviour, A and B share a useful pattern despite their different levels. The isolated spike is different.

    These are synthetic signals. For a task about absolute consumption, their levels also matter. Training determines which distinctions the representation keeps.

    signal = torch.tensor([[[0., 1., 0., -1., 0., 1.]]])
    time_encoder = nn.Conv1d(1, 3, kernel_size=3)
    e_time = time_encoder(signal).mean(dim=-1)
    print(e_time.shape)  # [1, 3]

    signal is [1, 1, 6]: one example, one channel, six readings. Three filters give [1, 3, 4]; averaging over time gives [1, 3].

    This fresh encoder only demonstrates the interface. Its numbers become useful after training for a suitable objective.

    A word vector

    Coordinates reflect its word model and training data.

    An image or signal vector

    Coordinates reflect that encoder and its task.

    We cannot compare unrelated embeddings just because both contain, say, 128 numbers.

    Image–text retrieval needs training that aligns the two spaces. We will return to that idea in the CLIP lesson.

    Our input is still a a b, and the observed next character is i. We learn the character table together with the MLP.

    The lookup gives three rows of two numbers. Next, we place them side by side to form the predictor's input.

    The name model learns its own character vectors. We are not inserting pretrained Word2vec vectors into it.

    Sources: Mikolov et al., Word2vec architectures (2013); negative sampling and phrases (2013); Le and Mikolov, Paragraph Vector (2014); Chen et al., SimCLR (2020); Yue et al., TS2Vec (2022). The domain diagrams illustrate the role of a representation, not results from those models. The recurring scene is the existing generated teaching illustration documented in the scene notes.

    A lookup selects a learned row. Keep the axis labels beside the numbers so we can read what each column means.

    05

    The MLP needs one fixed-width row. Concatenation keeps the three positions separate.

    Concatenate the embeddings

    We have looked up one embedding row for each character in the window. The MLP expects one row, so we put the three embeddings side by side. Three positions with two coordinates each give us six numbers. The first two always describe the oldest character, and the last two describe the newest. Reversing the characters changes where their numbers appear in this row. Adding the embeddings would lose that order because addition is commutative. The values below are the trained rows from the table we just inspected.

    The three rows for a a b

    $\pIn{e_1,e_2,e_3}$ are the embedding rows at the three window positions. $\pIn{a_0}$ is their concatenation, the MLP input. $\mathbb R^{1\times6}$ means one row of six real numbers.

    a0 = e.flatten(start_dim=1)
    print(a0.shape)  # [1, 6]

    e has three rows of two coordinates. a0 places those six numbers side by side, in position order.

    start_dim=1 preserves dimension 0, the batch dimension. For 6 examples, the same operation maps [6, 3, 2] to [6, 6].

    Concatenation keeps each position in its own part of the input row.
    The concatenations differ. The two sums are equal.

    With concatenation, each character position has a fixed slot. With addition, we can no longer recover the order.

    06

    The MLP combines the six input numbers to score each possible next character.

    Pass the vector through an MLP

    The concatenated row enters an MLP with one hidden layer. Each hidden unit combines all six input coordinates through the first weight matrix, adds a bias, and applies ReLU. ReLU keeps a positive result and replaces a negative result with zero. The second weight matrix combines the hidden activations to give one logit per vocabulary token. We draw only a few hidden nodes; the label beside them gives the trained layer's full width. The worksheets work through one hidden unit and the logit for i. Every other unit and output uses the same operations. Throughout this series, our vectors are rows and multiply matrices on the right.

    W: weights on connections. b: one bias per receiving unit.
    Dots mark omitted units. All six inputs come from a a b.

    $$\pAct{a_1}=\operatorname{ReLU}(\pIn{a_0}\pParam{W_1}+\pParam{b_1})$$
    $\pIn{a_0}$: input row
    The six numbers from our three embeddings.
    $\pParam{W_1}$: connection weights
    $6\times32$ learned numbers. Each hidden unit reads all six inputs.
    $\pParam{b_1}$: hidden biases
    32 learned numbers. Add one to each hidden unit's weighted sum.
    $\pAct{a_1}$: hidden activations
    The row of numbers produced by this layer.
    $\operatorname{ReLU}$: activation function
    Keep positive numbers. Replace negative numbers with zero.

    Each hidden unit can combine information from all six input coordinates.

    After adding the bias, apply the same rule to every hidden unit:

    $$\begin{aligned}-2 &\;\longrightarrow\; \pAct{0}\\0 &\;\longrightarrow\; \pAct{0}\\3 &\;\longrightarrow\; \pAct{3}\end{aligned}$$

    Negative results become zero. Positive results stay unchanged.

    before_relu = torch.tensor([-2., 0., 3.])
    after_relu = torch.relu(before_relu)
    print(after_relu)  # tensor([0., 0., 3.])

    These three numbers illustrate the rule. Our model applies it to all 32 hidden results.

    hidden = nn.Linear(6, 32)
    a1 = torch.relu(hidden(a0))
    print(a1.shape)  # [1, 32]

    nn.Linear(6, 32) creates the learned weights and bias. a1 contains 32 hidden activations for each example.

    hidden(a0) multiplies and adds the bias. torch.relu then acts on each result.

    W1 = hidden.weight.T
    b1 = hidden.bias
    check = torch.relu(a0 @ W1 + b1)
    torch.testing.assert_close(a1, check)

    Our equation uses $\pParam{W_1}$ with shape [6, 32]. PyTorch stores hidden.weight with shape [32, 6].

    .T swaps rows and columns. The bias has 32 entries and is added to every example. Both expressions calculate the same hidden row.

    $$\pScore{z}=\pAct{a_1}\pParam{W_2}+\pParam{b_2}$$
    $\pAct{a_1}$: hidden activations
    The same row we just calculated.
    $\pParam{W_2}$: output weights
    $32\times27$ learned numbers. Each logit reads all 32 hidden activations.
    $\pParam{b_2}$: output biases
    27 learned numbers. Add one to each vocabulary score.
    $\pScore{z}$: logits
    One score for each of the 27 vocabulary tokens.

    These scores become probabilities in the next step.

    output = nn.Linear(32, 27)
    z = output(a1)
    print(z.shape)  # [1, 27]

    z contains one logit for each vocabulary token. output.weight.T is $\pParam{W_2}$; output.bias is $\pParam{b_2}$.

    A column index selects a possible next character. For example, z[0, stoi["i"]] is the score for “i”.

    model_seq = nn.Sequential(
        embedding, nn.Flatten(1), hidden, nn.ReLU(), output
    )
    z_seq = model_seq(ctx)

    nn.Sequential runs the same layers in order, with the same parameters.

    z_seq: 27 logits per example. Softmax comes next.

    class NameMLP(nn.Module):
        def __init__(self, embedding, hidden, output):
            super().__init__()
            self.embedding = embedding
            self.hidden = hidden
            self.output = output
        def forward(self, ctx):
            a0 = self.embedding(ctx).flatten(1)
            a1 = torch.relu(self.hidden(a0))
            return self.output(a1)
    
    model = NameMLP(embedding, hidden, output)
    z = model(ctx)  # [1, 27], same logits as model_seq(ctx)

    __init__: store layers. forward: compute logits on each call.

    NameMLP inherits from nn.Module. self refers to this model object. super().__init__() initializes the PyTorch module machinery before we attach layers. Storing each layer on self registers it, so model.parameters() finds its learned weights. We pass in the existing layer objects, so this class and the nn.Sequential version share parameters and produce the same logits. Call model(ctx) rather than calling forward directly, so PyTorch also handles registered hooks. From here on, model means this NameMLP instance, including during training and generation. model_seq is the equivalent Sequential version.

    For $\pIn{a_0}$, $\pAct{a_1}$, and $\pScore{z}$: rows count examples; columns count numbers per example.

    Our whole window a a b is one example, so each array has one row.

    With B examples, the shapes become [B, 6], [B, 32], and [B, 27]. The feature and vocabulary widths stay the same.

    Weight-matrix axes count input and output features—not examples.

    A bias row's leading 1 does not count examples. Its offsets are shared.

    PyTorch stores biases as [32] and [27]. Broadcasting adds them to every row.

    Take the first four context windows from our training array X. Each window becomes one row.

    a0_batch = embedding(X[:4]).flatten(1)
    a1_batch = torch.relu(hidden(a0_batch))
    z_batch = output(a1_batch)
    print(a0_batch.shape, a1_batch.shape, z_batch.shape)
    # [4, 6], [4, 32], [4, 27]

    4 rows of 6 input features. 4 rows of 32 hidden activations. 4 rows of 27 vocabulary scores.

    Each row is processed independently using the same weights and biases. We changed the number of examples, not the network.

    Two entries from the forward pass for a a b:

    Hidden unit 1: sum six input-weight products, add its bias, then apply ReLU.

    Score for i: sum all 32 hidden-weight products, then add the output bias.

    The output layer repeats this calculation for every vocabulary token, giving 27 logits. Softmax will turn them into probabilities.

    Displayed values are rounded. Both sums use the full-precision products.

    For a closer look, these tables show the products inside the two sums above.

    Column 1 of $\pParam{W_1}$

    The logit for target i

    The network repeats the output calculation for every possible next character. To check the two shorter calculations shown in class, you can follow every scalar multiplication in these worksheets.

    Every hidden unit reads the full concatenated window. Every output then receives the full hidden row.

    07

    Logits can be any real numbers. We need probabilities that add up to one.

    27 scores to probabilities

    The output layer gives us one logit for each vocabulary token. A logit can be positive or negative; we cannot read it as a probability yet. Softmax exponentiates each score, then divides by the sum of all the exponentials. The classroom table uses that formula directly for aab, keeping vocabulary order and showing a few rows. Dots stand for the omitted tokens c through y. Every vocabulary entry still contributes to the denominator. PyTorch subtracts the largest logit internally for numerical stability, as explained below. This leaves the probabilities unchanged.

    $$\pProb{p_j}=\frac{\exp(\pScore{z_j})}{\sum_{r\in\mathcal V}\exp(\pScore{z_r})}$$
    $\pScore{z_j}$: logit for token $j$
    $\exp(z_j)=e^{z_j}$ makes it positive.
    $\pProb{p_j}$: probability for token $j$
    Divide by the total for all 27 tokens.

    $\sum_{r\in\mathcal V}$ adds over the vocabulary $\mathcal V$. The index $r$ visits each token.

    p = z.softmax(dim=-1)
    print(p.shape, p.sum(dim=-1))  # [1, 27], one total

    dim=-1 selects the vocabulary columns. Every example's probability row sums to one.

    The same softmax, with smaller intermediate numbers

    $$\pProb{p_j}=\frac{\exp(\pScore{z_j}-\pScore{z_{\max}})}{\sum_{r\in\mathcal V}\exp(\pScore{z_r}-\pScore{z_{\max}})}$$
    $\pScore{z_{\max}}$: the largest logit
    Subtract this same number from every score before exponentiating.
    $\pProb{p_j}$: the same probability
    Numerator and denominator both lose the factor $e^{z_{\max}}$, so it cancels.

    $j$ is the token we are scoring. The sum visits every token $r$ in the vocabulary $\mathcal V$.

    The largest exponential is now 1. This avoids very large intermediate numbers.

    shifted = z - z.max(dim=-1, keepdim=True).values
    weights = shifted.exp()
    p = weights / weights.sum(dim=-1, keepdim=True)

    Subtract the largest score in each row, then exponentiate. Divide by that row's total to get probabilities.

    keepdim=True keeps the maximum and total shaped [batch, 1], so each broadcasts across its own row.

    In normal model code, use z.softmax(dim=-1) for this operation.

    The boundary token is -, as in our vocabulary and code. Displayed values are rounded to three decimals. A positive probability below 0.001 is shown as <0.001. The probabilities use all full-precision logits. The next chart ranks the most likely outcomes, while this worksheet preserves vocabulary order.

    Next-character probabilities after a a b

    The observed next character is i. Its probability is . Next, we use that probability to measure the prediction's loss.

    Changing one vocabulary logit affects every probability, because all of them share the same softmax denominator.

    08

    How much probability did the model give the observed target? The loss turns that into one number.

    The loss

    Each training row tells us the next character. Cross-entropy takes the probability assigned to that target and applies the negative logarithm. A confident correct prediction has a small loss. If the model is confidently wrong, it gives the observed target very little probability and gets a large loss. The comparison below uses two simple probabilities from the lecture. The interactive worksheet instead reads all six aabid rows from the trained model. Choose a different row and its context, target probability, and loss change together. During training, we average this signal across many rows and update every parameter that contributed to the prediction.

    Loss for a a b → i

    $$\pLoss{\operatorname{loss}}=-\log\pProb{p}(\pTarget{\text{target}})$$
    Target: the observed next character
    For this training row, it is “i”.
    $\pProb{p}$: assigned probability
    Read the probability at the observed target.
    Loss: prediction penalty
    One number used to train the model.
    $-\log$: negative natural logarithm
    A lower target probability gives a larger penalty.
    loss = F.cross_entropy(z, target)
    print(loss.item())

    z: logits shaped [1, 27]. target: the observed token ID, shaped [1].

    loss is one scalar. .item() reads it as a Python number for display.

    Pass logits, not p. This function combines a numerically stable log-softmax with the target's negative log-probability.

    Confident and correct

    Confident and wrong

    $\pProb{p}(\text{target})$ is the probability of the observed next character. Loss measures the penalty for that prediction.

    Choose a context-target pair

    Giving more probability to the observed target makes its loss smaller.

    As the target probability approaches zero, the cross-entropy penalty grows sharply.

    09

    To improve future predictions, training must change the model's learned numbers.

    Learn the embeddings and the weights

    The optimizer updates the parameters that produced the loss: the embedding table, both weight matrices, and both bias rows. For a single context, only the embedding rows we looked up receive a direct embedding gradient. The shared MLP parameters receive gradients through the whole calculation. Repeating these updates over many shuffled training rows lowers the average loss. The curve records that loss at regular checkpoints. To see what changed for aab, compare the bars before and after training. They use two saved parameter sets with the same context, so a different input cannot explain the change.

    All learned parameters $\pParam{\theta}$

    $\pIn{E_{\text{tok}}}$ stores embedding rows. $\pParam{W_1,W_2}$ are weight matrices. $\pParam{b_1,b_2}$ are bias rows.

    Our model is a NameMLP, which inherits parameters() from nn.Module.

    for name, param in model.named_parameters():
        print(name, list(param.shape))

    Assigning layers to self registers them. parameters() yields their tensors without names.

    The names describe paths inside the model object: embedding.weight belongs to model.embedding. In NameMLP.__init__, self.embedding = embedding registers the embedding layer as a child module, and the other assignments register the two linear layers. Each layer already stores its learned tensors as nn.Parameter objects. The inherited parameters() method visits the registered children and yields their parameter tensors. It does not create weights or return the temporary outputs e, a0, a1, or z. The shapes here follow PyTorch storage: linear weights are transposed relative to our equations, and biases are one-dimensional.

    Continue with model, the NameMLP object we inspected.

    params = list(model.parameters())  # the five tensors above
    optimizer = torch.optim.SGD(params, lr=0.1)
    model.train()

    The optimizer keeps references to these same five tensors. lr=0.1 sets the learning rate.

    model.train() selects training mode. optimizer.step() will change the parameters after backpropagation.

    optimizer.zero_grad()
    z = model(X)
    loss = F.cross_entropy(z, y)
    loss.backward()
    optimizer.step()

    zero_grad() clears old gradients. The next two lines calculate fresh predictions and their mean loss.

    backward() computes gradients through the whole graph. step() uses them to change the stored parameters.

    This is one illustrative update. Training the name model repeats updates over many batches from the training names.

    print(model.embedding.weight.grad.shape)  # [27, 2]
    print(model.hidden.weight.grad.shape)     # [32, 6]

    A parameter's .grad has the same shape as that parameter. It describes how changing each entry would affect the loss.

    The optimizer updates the tables and weights. The next forward pass recomputes e and a0, a1, z, and p.

    For aab → i:

    The input looks up embedding rows a and b.

    The observed target i selects the cross-entropy loss.

    The i embedding is not looked up for this input.

    Cross-entropy every 50 steps

    Before training

    After training

    The same loss trains the embedding table and the MLP.

    Run the same input through two saved checkpoints to see how the changed parameters affect its prediction.

    10

    To generate a name, feed each chosen character back into the model.

    Sample, append, repeat

    We start generation with three boundary tokens. The model returns a distribution. We can choose the character with the largest probability or sample from the whole distribution. After appending that character, we shift the window right and run the same model again. Sampling allows different names because it can choose lower-probability characters too. Temperature controls how sharply the logits become probabilities: lower values concentrate probability near the largest logits, while higher values spread it out. The bars below always belong to the exact three tokens in the visible window. Choosing the boundary token ends generation.

    Next/Back pauses at each step. − − − marks the start.

    Each new input uses the newest three tokens. The output name keeps all the letters chosen so far.

    Keep the input − s a and the trained model fixed. Change only the temperature setting.

    The highest-scoring character stays n. Greedy still chooses it. Sampling uses the changed probabilities.

    $$\pProb{p}=\operatorname{softmax}\!\left(\frac{\pScore{z}}{\htmlClass{p1-temperature}{T}}\right)$$
    $\pScore{z}$: the 27 logits
    Scores from the unchanged model.
    $\htmlClass{p1-temperature}{T}$: temperature
    A positive setting we choose. Divide every logit by it.
    $\pProb{p}$: the 27 probabilities
    Softmax makes them positive and normalizes their sum to 1.

    Dividing by a number below 1 amplifies the differences between logits. Dividing by a number above 1 reduces those differences.

    The model's weights stay fixed. We change how its scores become sampling probabilities.

    After generating s and a, the input window is − s a. The plots above use all 27 logits; two of them are .

    The last column shows the probability of choosing n after softmax over all 27 divided logits, not just the two shown. Lower temperature gives this highest-scoring character more probability; higher temperature gives it less.

    @torch.no_grad()
    def sample_next(ctx, temperature=1.0):
        z = model(ctx)
        p = (z / temperature).softmax(dim=-1)
        return torch.multinomial(p, num_samples=1)

    multinomial samples one token ID from each probability row. For our single name, its result has shape [1, 1].

    Divide z by a positive temperature before softmax. no_grad() skips the gradient graph. For greedy choice, use argmax rather than dividing by zero.

    model.eval()
    ctx = torch.full((1, 3), stoi["-"], dtype=torch.long)
    name = []
    temperature = 1.0

    ctx is the opening window - - -. name will collect the chosen letters.

    eval() selects evaluation behaviour. It does not disable gradients; the decorator on sample_next does that. Our small model has no dropout or batch normalization, so its forward calculation is unchanged.

    for _ in range(18):
        next_id = sample_next(ctx, temperature)
        if next_id.item() == stoi["-"]: break
        name.append(vocab[next_id.item()])
        ctx = torch.cat([ctx[:, 1:], next_id], dim=1)

    Replay the earlier sample: the first two calls chose s and a. We now have name = ["s", "a"] and ctx = tensor([[0, 19, 1]]), representing − s a. The next draw in that run is m. Another sampling run can choose a different letter.

    ctx = torch.tensor([[stoi["-"], stoi["s"], stoi["a"]]])
    with torch.no_grad(): p = model(ctx).softmax(dim=-1)
    greedy_id = p.argmax(dim=-1, keepdim=True)
    sampled_id = torch.multinomial(p, num_samples=1)

    argmax takes the largest entry. multinomial draws from all 27 entries. Both return one ID, shape [1, 1].

    Blank starts at − − −. Only the last 3 letters feed the model.

    ?
    Name so far

    Saved outputs from the training script

    Generation runs the trained model repeatedly. Each next token comes from a choice, rather than an observation in the data. The worked animation replays sampling seed 3 in the browser at temperature 1. Its probabilities come from the saved trained model. PyTorch uses a different random generator, so its sampled letters can differ even when the probabilities agree. Greedy choice takes the largest probability at each call. Those local choices need not produce the most probable whole name.

    In the live generator, try starting text sa: its first input is − s a. A longer starting fragment stays in the name, but only its last three letters enter the model. Capitals become lowercase. The seed sets the sequence of random draws; it does not change the probabilities or retrain the model. Generate replays the entered settings from the start. New seed changes the seed and clears the run; Start over clears the run without changing any settings. Changing a setting also starts a fresh run.

    11

    Training and generation use the same forward pass. What changes is where the next token comes from.

    Training and generation

    During training, the data gives us both the context and the observed next character. We run the model to get probabilities, use cross-entropy to score the observed target, and update the parameters. Since all windows have the same shape, we can process many in a batch. During generation, we do not know the next character. We run the same forward pass, then choose a character from its distribution using a rule. Appending that choice makes it part of the next input. We have no observed target and make no parameter update. The calculation inside the forward pass is unchanged; what we do with its output is different.

    The full name aabid is in the data. We can prepare every input and target before running the model.

    A wrong prediction for a a b does not change the next training input a b i. The i comes from the data.

    batch_z = model(X)
    batch_loss = F.cross_entropy(batch_z, y)

    X holds the six windows we just saw: [6, 3] means 6 examples, 3 token IDs each. The output [6, 27] has 27 scores for each example.

    y is [6]: one observed target ID per example. Cross-entropy pairs each score row with its target and averages the six losses.

    Computing a loss alone does not update the model. That still requires backward() and optimizer.step().

    No next letter is supplied. Continue from aab by sampling from the frozen model.

    Within one name, each step waits for the previous choice. We can still generate several names in a batch.

    Training uses observed targets to change the model. Generation keeps the model fixed and extends the sequence. The lock and snowflakes refer to the stored parameters: the embedding table, both weight matrices, and both bias vectors. The activations and probabilities are still recomputed on each call. In the code we already used, torch.no_grad() skips gradient tracking and there is no optimizer step. model.eval() alone does not freeze parameters.

    12

    A token need not be a character. We still ask the model to predict the next one.

    What should a token be?

    A character vocabulary can represent unfamiliar words when it includes all their letters, but it turns ordinary sentences into long sequences. Whole words give shorter sequences. Their vocabulary grows quickly, though, and a fixed word vocabulary can map unseen words to one unknown token. Subwords offer a compromise: common pieces can be reused inside both frequent and rare words. Try the three simple tokenizers below on the same text. Each reports the sequence length and the number of distinct tokens in that sample. The subword merges illustrate the idea; they do not come from a trained tokenizer. This first demo shows splitting only. The examples below also show what happens when a fixed vocabulary does not contain a word.

    Character

    One letter or mark per token

    Word

    One space-separated word per token

    Subword

    Reusable pieces inside words

    One sentence, three token choices

    text = "deep learning is amazing"
    char_tokens = list(text)
    word_tokens = text.split()

    These Python operations choose the units first. Character tokens include spaces; split() produces four word tokens here.

    Next, a vocabulary maps tokens to IDs, and torch.tensor stores those IDs. A trained subword tokenizer replaces this simple splitting step when we want subword tokens.

    The trade-off

    None of these whole words is in our small, fixed word vocabulary.

    All four use the same <UNK> ID. The model cannot recover which spelling appeared from that ID alone.

    The same words, split into letters. All their letters are in a–z.

    The spellings remain distinct. We need an entry for each letter, not each whole word.

    This works because every letter is covered. An unsupported character would still need handling.

    The same words, using our demo's fixed merge rules and base letters a–z.

    unbelievable uses three pieces. nanobotany stays as letters because none of our rules merges them.

    These illustrative merges preserve each spelling. Subwords still need a vocabulary that covers the input.

    An unseen whole word can still consist entirely of known letters or subword pieces. That is why the character and subword versions above preserve the four different spellings while the word version maps all four to <UNK>. Every box represents one token. Repeated pieces reuse the same vocabulary ID and embedding row. Splitting into more pieces means more tokens, but does not require a new embedding row for the entire word.

    The word example turns on our demo's fixed-vocabulary lookup. Calling split() alone only separates words; it does not decide whether a word is known. The character and subword demos illustrate splitting and merges, not a complete production tokenizer. These examples use lowercase a–z, so every base letter is available. For other characters, coverage depends on the vocabulary and its fallback policy. Being able to spell a new word does not by itself mean the model understands it.

    13

    Word tokens change the vocabulary and data. The prediction steps stay the same.

    Same model, word tokens

    With word tokens, we still look up embeddings, combine the context, produce logits, and apply softmax. We have changed the tokenizer, vocabulary, and training text. Several continuations can fit a prefix, so we still want a distribution, even though training gives us just one observed target at each position. The chain rule writes the probability of a complete sequence as a product of conditional next-token probabilities. It tells us which probabilities the model must supply as the context grows. It does not tell us which network to use to represent that context.

    Character model

    a a b → i

    embedding → concatenate → MLP → softmax

    Word model

    The cat sat on → the

    embedding → concatenate → MLP → softmax

    We changed the vocabulary, tokenizer, and training data.
    word_vocab = ["<unk>", "The", "cat", "sat", "on", "the"]
    word_stoi = {t: j for j, t in enumerate(word_vocab)}
    word_ctx = torch.tensor([[2, 3, 4]])  # cat sat on
    word_embedding = nn.Embedding(len(word_vocab), 2)
    word_e = word_embedding(word_ctx)

    word_e is still [1, 3, 2]: one example, three tokens, two coordinates.

    Keep the latest three words of “The cat sat on” to predict what follows. This is a separate word vocabulary and a fresh table, which must be trained on word sequences.

    word_hidden = nn.Linear(6, 32)
    word_output = nn.Linear(32, len(word_vocab))
    word_a1 = torch.relu(word_hidden(word_e.flatten(1)))
    word_z = word_output(word_a1)

    word_z has shape [1, 6], matching this six-word vocabulary.

    The operations stay the same. The learned parameters belong to this new word model; they are not the trained character model's weights.

    The cat sat on the ___

    Illustrative word probabilities.

    Whether our tokens are characters or words, the next-token objective and the conditional-probability chain stay the same.

    14

    Window size and embedding width determine the MLP input width.

    The context keeps growing

    Our trained name model uses three token slots, with two numbers per embedding. Concatenation gives six inputs. Here we explore other architecture choices. Increasing the window adds more embeddings. Increasing the embedding width adds more numbers to every embedding. Both enlarge the first weight matrix. Only embedding width changes the size of the embedding table, because that table has one shared row per vocabulary token, not one row per position. None of these controls change the trained model used earlier.

    Shapes show one example. Hidden width stays 32, vocabulary size stays 27.

    Larger w changes W₁. Larger d changes $E_{\mathrm{tok}}$ and W₁. Token positions share the embedding table.

    new_w, new_d = 5, 4
    size_model = nn.Sequential(
        nn.Embedding(27, new_d), nn.Flatten(1),
        nn.Linear(new_w * new_d, 32), nn.ReLU(), nn.Linear(32, 27))
    print(sum(p.numel() for p in size_model.parameters()))  # 1671

    5 embeddings × 4 numbers = 20 inputs per example. The table is [27, 4]. W₁ is [20, 32].

    PyTorch stores the first linear weight as [32, 20], the transpose of our W₁. Each linear layer includes a bias.

    This example creates a new, untrained model. The earlier trained model still uses w = 3 and d = 2.

    w = 100 token slots; d = 256 numbers per embedding.
    Hidden width $d_h$ = 1,024 units.

    $$\underbrace{\pIn{(w d)}}_{\text{input features}}\times\underbrace{\pAct{d_h}}_{\text{hidden units}}=\pIn{(100\times256)}\times\pAct{1{,}024}=26{,}214{,}400$$

    w d is the length of the concatenated input. $d_h$ is the number of hidden units, each connected to every input.

    That is about 26 million weights in W₁ alone. Its 1,024 biases and the other parameters are extra.

    The model still expects exactly 100 input slots.

    Can we read more context without changing the learned matrix shapes?

    A fixed concatenation needs a wider first layer whenever we add input slots.

    Continue to Part 2: Self-attention

    A larger window lets us look further back and requires a larger matrix. It still imposes a fixed limit on the context. Attention lets the number of source positions vary while keeping the learned projection shapes fixed, though computation, memory, and positional schemes still impose practical limits.

    15

    This model is small enough to check the reason for each design choice.

    Pause and think

    Try reconstructing the model from these questions without rereading the page in order. For each design choice, explain what it does. One-hot rows identify tokens but give us no compact learned geometry. Concatenation preserves position, and the output needs one logit per vocabulary token. The negative logarithm rewards probability on the observed target and strongly penalizes probabilities near zero. A different window changes the first matrix because its input width is $wd$. For generation, we rely on training to make this same forward pass produce useful next-token distributions.

    Each component has a job: preserve information, help produce a distribution, or supply a training signal. Can you explain which job each one does?

    16

    Connect the token IDs to the loss, then to generation and the fixed-context limit.

    Three summaries

    We have built a complete character language model. Three token IDs select embedding rows. Concatenation keeps their positions separate, the MLP produces one logit per vocabulary token, and softmax gives us probabilities. Cross-entropy scores the observed next character; its gradients train every learned table and matrix. To generate, we run the same model, sample a character, append it, shift the window, and repeat. This works, but $W_1$ fixes the window width. Part 2 keeps the next-token objective and changes how a token gathers information from its preceding context.

    Intuitive

    Use the last few characters to guess what character tends to come next.

    Operational

    Look up embeddings, concatenate the window, run the MLP, apply softmax, then sample or score the observed target.

    $$\pIn{a_0}=[\pIn{e_1},\ldots,\pIn{e_w}]$$

    Place the $w$ embedding rows side by side to form the input row.

    $$\pAct{a_1}=\operatorname{ReLU}(\pIn{a_0}\pParam{W_1}+\pParam{b_1})$$

    Hidden activations combine the input row using learned weights and bias, then apply ReLU.

    $$\pProb{p}=\operatorname{softmax}(\pAct{a_1}\pParam{W_2}+\pParam{b_2})$$

    Next-token probabilities come from transforming the hidden row with the output weights and bias, then applying softmax.

    e = embedding(ctx)
    a0 = e.flatten(1)
    a1 = torch.relu(hidden(a0))
    z = output(a1)
    p = z.softmax(dim=-1)

    Lookup and concatenate. Calculate the hidden activations. Produce vocabulary scores. Normalize them into probabilities.

    During training, send z and an observed target to cross-entropy. During generation, sample from p and form the next context.

    Download the code in lecture order

    We still want to predict the next token. We need a different way to represent its context.

    Next: Part 2, Self-attention

    Attention will replace the fixed concatenation. We will keep using embeddings, logits, softmax, loss, and generation.