Companion workbook · 79-minute film · 59 chapters
How a Large Language Model Works
This workbook sits beside the film. The film shows each mechanism moving; the workbook slows it down so you can touch it: a real tokenizer to type into, a softmax you can push around, a sampler you can run a hundred times, a gradient you can watch overshoot. Each part ends with links for deeper study and questions to check yourself.
How to use this workbook
Each part below lists its chapters with their start times in the film.
Short summaries and diagrams, in the film’s colours, so the pictures connect.
Change the numbers and watch the mechanism respond. Every lab starts with the film’s own example.
Answer the quiz, then explain the ideas back in your own words before revealing the model answer.
The visual language
The film gives every concept one colour and one shape, and keeps it for 79 minutes. The workbook uses the same key, so a diagram here reads the same way as a frame there.
Orientation
Watch guide
Every chapter with its start time in the film and the terms it introduces. Use it to jump back to a chapter before a quiz, or to rewatch the one moment that didn’t land.
| Ch. | Starts | Chapter | New terms |
|---|
Orientation · film ch. 39 “The Whole Machine” · 56:20
The whole machine at a glance
Keep this map in mind while reading. Every part of the workbook zooms into one stretch of it. Text goes in at the top left, a single next token comes out at the bottom left, and the loop repeats.
Part I · film 2:02 – 12:24
From words to numbers
A computer only calculates. Before anything clever can happen, text has to become numbers that are worth calculating with.
- Tokens are pieces, not words. A tokenizer cuts text into pieces from a fixed vocabulary. Common words are one piece, rarer words split into fragments, punctuation is often its own piece, and the space before a word usually travels with it.
- A token ID is a row number. It says where to look in the vocabulary, nothing about meaning. Neighbouring IDs are unrelated, and arithmetic on IDs is nonsense.
- A vector is an ordered list of numbers; each position is a dimension. Lists can describe many aspects at once and can be compared, added and scaled.
- An embedding is a learned row of a big table: the starting vector for a token. The table is a parameter; the vector it produces for your text is an activation.
- Order matters, so position is added. “Dog bites man” and “man bites dog” contain the same embeddings; positional information is what tells them apart. Implementations vary.
o200k_base tokenizer; the vector values are illustrative.Lab 1A real tokenizer, in your browser
film ch. 2–3 · 2:02These are real, publicly available tokenizers used by well-known model families, loaded from the open-source gpt-tokenizer library. The film used illustrative IDs; here you see the real ones.
r50k_base and watch non-English text use many more tokens; newer tokenizers with bigger vocabularies cover more languages efficiently. A ␣ marks a leading space inside a token.Go deeper
- InteractiveTiktokenizertiktokenizer.vercel.app
Pick a model family and see how chat-formatted text is tokenized, including the special tokens that mark roles.
- Interactive
- VideoLet’s build the GPT TokenizerAndrej Karpathy · 2 h
Builds byte-pair encoding from scratch and explains many odd model behaviours (spelling, arithmetic, trailing spaces) as tokenization effects.
- Codeminbpegithub.com/karpathy
A tiny, readable implementation of the BPE algorithm used by modern tokenizers.
- Codetiktokengithub.com/openai
The fast production tokenizer library behind the encodings in Lab 1 (Python).
- DocsHugging Face Tokenizershuggingface.co
Train and use BPE, WordPiece and Unigram tokenizers; good explanations of each algorithm.
- PaperNeural Machine Translation of Rare Words with Subword UnitsSennrich et al., 2016
The paper that brought byte-pair encoding to neural language models.
- PaperSentencePieceKudo & Richardson, 2018
A language-independent tokenizer that treats text as raw characters, spaces included.
- InteractiveEmbedding ProjectorTensorFlow
Explore real word embeddings in 3-D projections; search a word and see its neighbours. Remember film ch. 4: every such picture is a shadow of many more dimensions.
- ArticleThe Illustrated Word2vecJay Alammar
A visual introduction to learned word vectors, the ancestors of today’s embeddings.
Part II · film 12:24 – 32:05
Attention
“Bank” starts with the same vector in “the canoe finally reached the bank” and “she deposited the check at the bank”. Attention is how the surrounding words change that.
- Each position makes a Query (what am I looking for?), a Key (how might I match a search?) and a Value (what can I contribute?), each by multiplying its representation with its own learned matrix. The phrases are intuition aids: all three are vectors.
- Scores: the dot product of one Query with each Key. Large when they point the same way.
- Scale, mask, softmax: divide scores by √dₖ, set future positions to −∞ (the causal mask), and turn the scores into positive weights that sum to 100%.
- Mix: the output is the weighted sum of the Values. In the canoe sentence, 62% of “bank”’s mixture comes from “canoe”.
- One such calculation is an attention head. Models run many heads in parallel and blend their outputs with one more matrix, the attention output projection: multi-head attention.
Lab 2Dot products, scaling and softmax
film ch. 13–17 · 18:42A 3-dimensional Query for “bank” meets four Keys. Edit any number. Scores are dot products; softmax turns them into weights. Values start as in film ch. 13.
Lab 3Why the causal mask is a triangle
film ch. 16 · 23:54Rows are the token doing the looking; columns are the token being looked at. Select a row to see what that position may use. Hatched cells are future positions, set to −∞ before softmax, which makes their weight exactly zero.
Go deeper
- PaperAttention Is All You NeedVaswani et al., 2017
The transformer paper. Section 3.2 is exactly the calculation in Fig 2.1, including the footnote on why scaling by √dₖ helps.
- InteractiveTransformer ExplainerGeorgia Tech Polo Club
Runs a small GPT-2 in your browser and lets you inspect the Q, K, V and attention weights for your own sentence.
- InteractiveLLM VisualizationBrendan Bycroft
A 3-D walk-through of every matrix multiplication in a small GPT, step by step.
- ArticleThe Illustrated TransformerJay Alammar
A classic visual explanation of self-attention and multi-head attention.
- VideoAttention in transformers, visually explained3Blue1Brown
A second visual angle on Queries, Keys and Values; good right after this part.
- PaperRoFormer: Rotary Position EmbeddingSu et al., 2021
The “build position into the calculation” approach mentioned in film ch. 8, used by many modern models.
- PaperGQA: Grouped-Query AttentionAinslie et al., 2023
Heads share Keys and Values to shrink the KV cache; a common modern variant of multi-head attention.
- ArticleIn-context learning and induction headsOlsson et al., 2022
One of the clearest cases of an attention head with an interpretable job, and a reminder that most heads are harder to describe.
Part III · film 32:05 – 44:08
The transformer
Attention moves information between positions. The rest of a layer works on each position by itself, and residual connections keep every update.
- The MLP (feed-forward network) expands each position’s vector (often ×4), bends every number with a nonlinear activation function, and compresses it back. Same weights at every position; no mixing between positions.
- Why the bend: two matrices in a row equal one matrix. Without a nonlinearity, depth adds nothing.
- Residual connections add each block’s output to the running representation instead of replacing it. Think of each position as a stream that blocks read from and write into.
- Normalization rescales vectors before each block (RMSNorm divides by the root mean square) so values stay in a stable range.
- A transformer stacks dozens of these layers. Parameters stay fixed between requests; activations exist only for the current text.
Lab 4Where the billions of parameters live
film ch. 22, 28, 55Estimate a transformer’s parameter count from its shape. Per layer: attention holds 4·d² (Query, Key, Value and output matrices); the MLP holds 2·r·d². Embeddings add V·d, and a separate output projection adds another V·d. Biases and norms are ignored. Presets reproduce published totals closely.
A parameter is one learned number. Reading 6.7 billion of them aloud at one per second would take about 212 years.
Go deeper
- VideoLet’s build GPT: from scratch, in codeAndrej Karpathy · 2 h
Writes every block of Fig 3.1 in PyTorch and trains it on Shakespeare. The best next step if you can read a little Python.
- CodenanoGPTgithub.com/karpathy
About 300 lines of model code that reproduce GPT-2 (124M). Read
model.pynext to this part. - ArticleA Mathematical Framework for Transformer CircuitsElhage et al., 2021
Where the “residual stream” picture comes from, and how to reason about what layers read and write.
- PaperTransformer Feed-Forward Layers Are Key-Value MemoriesGeva et al., 2021
Evidence behind the film’s “research suggests MLPs store and apply associations”.
- PaperDeep Residual LearningHe et al., 2015
Why adding instead of replacing made very deep networks trainable.
- PaperRoot Mean Square Layer NormalizationZhang & Sennrich, 2019
RMSNorm, the normalization worked by hand in film ch. 25.
- PaperGLU Variants Improve TransformerShazeer, 2020
SwiGLU and the other gated activation functions named in film ch. 23.
Part IV · film 44:08 – 57:25
From numbers back to words
The last position’s vector becomes one score per vocabulary token, then probabilities, then one chosen token that is appended before everything runs again.
- The output projection turns the final representation into a logit for every token in the vocabulary: raw scores that can be negative.
- Softmax, the same operation used inside attention, turns logits into a probability distribution. A token’s probability measures how well it fits the text, not whether a statement is true.
- Decoding: greedy picks the top token; sampling picks in proportion to probability. Temperature reshapes the distribution; top-k and top-p cut the long tail.
- Autoregressive: each chosen token is appended and the model runs again. The KV cache keeps earlier Keys and Values so only the new position is computed; it holds activations, not parameters, and is discarded afterwards.
Lab 5Temperature, top-k, top-p and sampling
film ch. 32–35 · 46:02The film’s illustrative next-token distribution after “The canoe finally reached the”, over a 50,000-token vocabulary. At temperature 1 with no filters it reproduces the film: bank 58.5%, shore 29.0%, river 3.2%.
Go deeper
- PaperThe Curious Case of Neural Text DegenerationHoltzman et al., 2019
Introduces nucleus (top-p) sampling and shows why greedy text becomes repetitive.
- ArticleHow to generate textHugging Face blog
Greedy, beam search, top-k and top-p side by side with runnable examples.
- DocsGeneration strategiesTransformers docs
The real knobs (temperature, top_k, top_p, repetition penalty) and what each does.
- PaperPagedAttention (vLLM)Kwon et al., 2023
How serving systems manage KV-cache memory, the cost the film points out in ch. 37.
Part V · film 57:25 – 68:07
Training
Nobody types in the billions of numbers. They start random and are nudged, one small step at a time, toward predicting text better.
- Targets are free. In “The largest planet in our solar system is Jupiter”, every token is the target for the text before it. The causal mask lets one forward pass train every position at once.
- Loss = −log(probability given to the correct token). 100% → 0; 1 in 50,000 → 10.8. Averaged over many predictions this is cross-entropy.
- The gradient gives, for every parameter, how the loss would change under a tiny nudge. Backpropagation computes all of them in one backward sweep using the chain rule.
- The optimizer steps each parameter against its slope (learning rate × gradient; Adam/AdamW refine this). Parameters change only at this step, and only during training.
Lab 6Loss as a penalty curve
film ch. 43 · 59:46Slide the probability the model gave to the correct next token. The penalty climbs steeply as the probability approaches zero.
Lab 7Gradient descent on one parameter
film ch. 45–47 · 61:35The film’s one-parameter loss curve. Each step moves w by −(learning rate × slope). Try a small rate, a good one and one that is too large.
Go deeper
- VideoThe spelled-out intro to neural networks and backpropagationAndrej Karpathy · micrograd
Builds backpropagation by hand on tiny expressions until the chain rule feels obvious.
- ArticleHow the backpropagation algorithm worksMichael Nielsen
A patient, free book chapter deriving backpropagation step by step.
- PaperAdam: A Method for Stochastic OptimizationKingma & Ba, 2014
Running averages of gradients and per-parameter step sizes (film ch. 47).
- PaperDecoupled Weight Decay Regularization (AdamW)Loshchilov & Hutter, 2017
The variant most large models are trained with.
- PaperScaling Laws for Neural Language ModelsKaplan et al., 2020
How loss falls with more parameters, data and compute; also the source of “training costs about 6 × parameters FLOPs per token”.
- PaperTraining Compute-Optimal LLMs (Chinchilla)Hoffmann et al., 2022
Why a smaller model trained on more data can beat a larger one (film ch. 55).
Part VI · film 68:07 – 74:51
What it all means
Training builds the model. Inference uses the model. Keeping those apart explains assistants, memory, hallucinations and scale.
- Post-training (fine-tuning on example conversations, preference training, reinforcement learning, safety and tool-use training) turns a raw text continuer into an assistant. Recipes differ between developers.
- Hallucination is a fluent, plausible continuation that is false or unsupported. Nothing in the generation loop checks claims against the world.
- Grounding puts outside information (search results, retrieved documents, tool outputs) into the context window. It reduces errors but does not eliminate them.
- Knowledge is distributed across many parameters, not stored as rows or single neurons. Interpretability research finds some features and circuits; most of the picture is still open.
Go deeper
- PaperTraining language models to follow instructions with human feedbackOuyang et al., 2022
The instruction-tuning plus preference-training recipe behind many assistants (film ch. 51).
- PaperDirect Preference OptimizationRafailov et al., 2023
A simpler way to train on preference comparisons without a separate reward model.
- Paper
- PaperA Survey on Hallucination in Large Language ModelsHuang et al., 2023
Causes, detection and mitigation of fluent-but-false output.
- ArticleScaling MonosemanticityTempleton et al., 2024
Interpretable features found inside a production-scale model: the “real progress” of film ch. 54.
- PaperExtracting Training Data from Large Language ModelsCarlini et al., 2020
Evidence for the “some memorization does happen” part of film ch. 50.
Part VII · film 74:51 – 77:28
Beyond the basics
Three ideas you will hear about constantly, each a small change to something you already understand.
Many smaller MLPs (“experts”) plus a learned router that sends each token to a few of them. Large total capacity, a fraction of the compute per token; all experts still sit in memory.
Storing each parameter with fewer bits (16 → 8 → 4) by rounding to fewer allowed values. Smaller and often faster, at some cost in precision.
A small draft model guesses several tokens; the large model checks them in one pass and keeps the ones it agrees with. Done correctly, the output distribution is unchanged, only faster.
Go deeper
- PaperOutrageously Large Neural Networks: the Sparsely-Gated MoE LayerShazeer et al., 2017
The router-and-experts design.
- Paper
- Paper
- Paper
- PaperFast Inference from Transformers via Speculative DecodingLeviathan et al., 2022
Includes the proof that the output distribution is preserved.
Self-test · after the whole film
Explain it back
The film was designed so that a beginner could answer every one of these. Say your answer out loud or write it down first, then open the model answer. Tick the box when you could explain it without looking.
Reference
Glossary
Every term the film teaches, with its on-screen definition and the chapter that earns it.
| Term | Definition | Ch. |
|---|
Continued study
Where to go next
Three paths, depending on what you want to do with this understanding.
Understand more deeply
No code required.
- 3Blue1Brown · Neural networks series
- Transformer Explainer with your own sentences
- The Illustrated Transformer
- LLM Visualization for the full scale picture
Build one yourself
Some Python helps.
- Neural Networks: Zero to Hero (micrograd → GPT)
- nanoGPT: train a small model
- Hugging Face LLM Course
- Dive into Deep Learning as a reference textbook
Follow the research
Papers, in reading order.
- Attention Is All You Need
- Scaling Laws and Chinchilla
- InstructGPT for post-training
- Transformer Circuits Thread for interpretability