Writing 4 min read
How a large language model works: a film, 30 labs and a workbook
A 79-minute animated film that follows one sentence through a language model, from the first token to the last, with hands-on labs and a workbook to go with it.
Ch. 1 · The Mystery 0:00 / 1:19:12
You type a question into a chat window. A moment later an answer starts to appear, piece by piece, as if someone were writing it. Freeze it mid-sentence and pick one piece of the text that just appeared. What actually happened, inside the computer, that produced that piece?
That question opens the film above, and the rest of it is a single long answer. I wanted an explanation that starts from nothing, doesn’t skip steps, and never uses a word before it has shown you the idea behind it. I couldn’t find one I liked, so I made one.
What the film is
How a Large Language Model Works is a narrated, fully animated film: 79 minutes, 59 chapters, 8 parts. It follows one running example, “The canoe finally reached the bank.”, through the whole machine: from raw text to tokens, vectors, attention, transformer layers, probabilities and the chosen next token, and then back through training to show where all those numbers came from.
One rule drove the whole script: no jargon before its idea. Every term is shown first, then explained, and only then named. “Softmax” doesn’t appear until you’ve watched a handful of scores turn into shares of 100%. “Attention head” doesn’t appear until you’ve followed one complete attention calculation by hand. Six quick-check pauses along the way ask you to predict what happens next before the film shows you.
| Part | What it covers | Jump in |
|---|---|---|
| I · From words to numbers | tokens, token IDs, vectors, embeddings, position | ▶ 2:09 |
| II · Attention | queries, keys, values, dot products, softmax, the causal mask, heads | ▶ 12:31 |
| III · The transformer | the MLP, residual connections, normalization, stacking layers | ▶ 32:12 |
| IV · Numbers back to words | logits, probabilities, temperature, top-k/top-p, the KV cache | ▶ 44:15 |
| V · Training | the target, loss, gradients, backpropagation, the optimizer | ▶ 57:32 |
| VI · What it all means | post-training, training vs. inference, why models state falsehoods fluently | ▶ 1:08:14 |
| VII · Beyond the basics | mixture of experts, quantization, speculative decoding | ▶ 1:14:58 |
The timestamps jump the player at the top of this page. The film page has all 59 chapters.
The whole machine, in ten lines
If the film had to fit on an index card, it would be this loop. Every line is a chapter or two; the film spends 79 minutes making each one feel obvious.
tokens = tokenize("The canoe finally reached the")
while not done:
x = embed(tokens) + position(tokens) # a vector per token
for layer in layers: # dozens of layers
x = x + attention(norm(x)) # gather context
x = x + mlp(norm(x)) # process each position
logits = unembed(norm(x[-1])) # a score per vocab token
probs = softmax(logits / temperature) # scores -> probabilities
next_token = sample(probs) # "bank" is likely, not certain
tokens.append(next_token) # ...and run it all again
Training explains where the numbers inside embed, attention, mlp and unembed came from. You compare the predicted
probabilities with the token that actually came next, measure the error as one number (the loss), and nudge every
parameter a little in whichever direction lowers it. Then you repeat that across trillions of tokens of text. Part V builds each of those
steps from scratch.
Touch the machinery: the labs
Watching is not the same as understanding, so the film has a companion: The LLM Lab, 30 interactive experiments that run entirely in your browser. You can type into a real tokenizer, push attention scores around and watch softmax respond, toggle the causal mask, and train a tiny language model while you watch the loss fall. A few good places to start:
- Lab 01 Text → Tokens Type anything and watch a real tokenizer cut it into pieces.
- Lab 09 Softmax Push raw scores around and watch them become shares of 100%.
- Lab 11 Causal Masking Toggle the mask and see why the triangle exists.
- Lab 20 Temperature One slider that reshapes the whole distribution.
- Lab 28 A Tiny Training Loop Train a real (tiny) language model in your browser.
- Lab 30 Mini Transformer Workbench Step a real, trained, tiny transformer through every stage - every number inspectable.
The labs follow the same order as the film, and the film page links each part to its labs.
Check yourself: the workbook
The companion workbook is for after (or during) the film. It has a watch guide where every chapter’s timestamp jumps straight into the film, diagrams in the film’s colours, seven more small in-page labs, 36 quiz questions, 50 explain-it-back prompts, and a 70-term glossary. It keeps your progress in your browser, so you can come back to it.
How it was made
Everything on screen is code. The film is built with Remotion, which lets you write video as React components. The visuals are SVG, DOM and Three.js, and every worked example is computed rather than hand-typed, so the numbers on screen are the real results of the maths they illustrate. The music and sound effects are generated procedurally. The narration is synthesized speech, and the narration drives the timing: change a line of the script, and the timeline re-flows around the new audio.
The master is 2560×1440 at 60 fps, with loudness-mastered audio and soft subtitles. On this site it streams as adaptive HLS in three renditions (1440p, 1080p and 720p), served from Cloudflare’s edge, so it starts quickly on a phone and still looks sharp on a large monitor.
What’s next
This is the first film in what I’m calling The Personal Learning Series. More are coming, each with the same shape: a film that builds intuition, labs that let you poke at the mechanism, and a workbook to check what stuck. Subscribe via RSS if you want to know when the next one lands.