Next-token prediction
Give three trained models the same opening. Compare what they write and how long they take.
The prompt stays on this device. Initial downloads include the models and ONNX Runtime. Timings below each output exclude downloads and warm-up. Short continuations are examples, not a quality score.
Held-out results
All three models use the same stories, vocabulary, 64-token window and maximum training budget. Lower cross-entropy and perplexity are better.
| Model | Parameters | Test cross-entropy | Test perplexity | Training time |
|---|
Cross-entropy and perplexity
For each test example, take −ln of the probability assigned to the observed next token. The average is cross-entropy, in nats per token. Perplexity is exp(cross-entropy): a loss of ln(4) gives perplexity 4, equivalent to assigning probability 1/4 to every target. It is not an accuracy percentage. Compare these numbers only with the same tokenizer and test targets.
What differs between the models?
- The MLP concatenates the 64 token embeddings. Slot order is built into the concatenation.
- Both attention models add learned absolute position embeddings before computing queries, keys and values. Positions refer to slots 0–63 in the current cropped window.
- One head uses 64 matching coordinates. Four heads use 16 each, form separate weight rows, then concatenate their messages. The total width and parameter count stay fixed.
- Each attention model has one attention block, an output projection and residual, then the same prediction MLP. These are teaching models without a full Transformer’s normalization, block FFN or stack of blocks.
Four heads reuse the one-head optimizer settings. We did not tune them against test results. More heads need not improve every run or prompt.