LLM Companion Workbook
LLM Companion WorkbookFor the film · How a Large Language Model Works

Companion workbook · 79-minute film · 59 chapters

How a Large Language Model Works

This workbook sits beside the film. The film shows each mechanism moving; the workbook slows it down so you can touch it: a real tokenizer to type into, a softmax you can push around, a sampler you can run a hundred times, a gradient you can watch overshoot. Each part ends with links for deeper study and questions to check yourself.

7 parts, matching the film8 diagrams7 hands-on labs36 quiz questions50 explain-it-back prompts70 glossary terms

How to use this workbook

Watch a part

Each part below lists its chapters with their start times in the film.

Read the core ideas

Short summaries and diagrams, in the film’s colours, so the pictures connect.

Try the lab

Change the numbers and watch the mechanism respond. Every lab starts with the film’s own example.

Check yourself

Answer the quiz, then explain the ideas back in your own words before revealing the model answer.

The visual language

The film gives every concept one colour and one shape, and keeps it for 79 minutes. The workbook uses the same key, so a diagram here reads the same way as a frame there.

Tokena paper “type slug”
Token IDsteel number tag
Activation / vectorsky-blue bars; flows
Parameter / matrixamber lattice; stored
Querycoral probe ▶
Keyviolet notched tab
Valuegreen capsule ●
Attention weightline thickness + %
Logitslate bars, can go below 0
Probabilityrose bars, never below 0
Gradientred dashes pointing back
Inference / Trainingcool world / warm world

Orientation

Watch guide

Every chapter with its start time in the film and the terms it introduces. Use it to jump back to a chapter before a quiz, or to rewatch the one moment that didn’t land.

Ch.StartsChapterNew terms

Orientation · film ch. 39 “The Whole Machine” · 56:20

The whole machine at a glance

Keep this map in mind while reading. Every part of the workbook zooms into one stretch of it. Text goes in at the top left, a single next token comes out at the bottom left, and the loop repeats.

Inference pipeline: text, tokenizer, tokens, IDs, embeddings plus position, transformer layers, final representation, output projection, logits, softmax, probabilities, selection, append and repeat. Text “…mostly an” Tokenizer ch. 2 ␣an … TOKENS #448 TOKEN IDS · ch. 3 Embedding + position CH. 4–8 TRANSFORMER LAYER × N · CH. 9–29 Q·K·V attention MLP norm → attention → + → norm → MLP → + Final representation last position · ch. 29 Output projection ch. 31 Logits one per token softmax Probabilities ch. 32 Select a token “ illusion” · ch. 33–35 APPEND ↻ AUTOREGRESSIVE LOOP · CH. 36 — EACH NEW TOKEN IS APPENDED AND THE WHOLE COMPUTATION RUNS AGAIN (WITH A KV CACHE, CH. 37)
FIG 0The inference pipeline as the film draws it in chapter 39. Amber boxes hold learned parameters; blue boxes are activations computed for this particular text. The token ID “#448” is illustrative.

Part I · film 2:02 – 12:24

From words to numbers

A computer only calculates. Before anything clever can happen, text has to become numbers that are worth calculating with.

  • Tokens are pieces, not words. A tokenizer cuts text into pieces from a fixed vocabulary. Common words are one piece, rarer words split into fragments, punctuation is often its own piece, and the space before a word usually travels with it.
  • A token ID is a row number. It says where to look in the vocabulary, nothing about meaning. Neighbouring IDs are unrelated, and arithmetic on IDs is nonsense.
  • A vector is an ordered list of numbers; each position is a dimension. Lists can describe many aspects at once and can be compared, added and scaled.
  • An embedding is a learned row of a big table: the starting vector for a token. The table is a parameter; the vector it produces for your text is an activation.
  • Order matters, so position is added. “Dog bites man” and “man bites dog” contain the same embeddings; positional information is what tells them apart. Implementations vary.
Embedding lookup: the token " bank" has ID 6922, which selects row 6922 of the embedding table, producing a vector. ␣bank TOKEN #6,922 TOKEN ID selects row EMBEDDING TABLE · ONE ROW PER VOCABULARY ENTRY (≈200,000 × thousands) EMBEDDING VECTOR (ACTIVATION)
FIG 1.1The lookup behind every token. The ID is a row number; the numbers in the row are learned during training. Real IDs from the o200k_base tokenizer; the vector values are illustrative.

Lab 1A real tokenizer, in your browser

film ch. 2–3 · 2:02

These are real, publicly available tokenizers used by well-known model families, loaded from the open-source gpt-tokenizer library. The film used illustrative IDs; here you see the real ones.

loading tokenizer…
0characters
0tokens
0chars / token
—vocabulary size
Try this. Compare “bank”, “ bank” and “Bank” (film ch. 3): three unrelated IDs. Paste a long number and see how it splits. Switch to r50k_base and watch non-English text use many more tokens; newer tokenizers with bigger vocabularies cover more languages efficiently. A ␣ marks a leading space inside a token.

Go deeper

Part II · film 12:24 – 32:05

Attention

“Bank” starts with the same vector in “the canoe finally reached the bank” and “she deposited the check at the bank”. Attention is how the surrounding words change that.

  • Each position makes a Query (what am I looking for?), a Key (how might I match a search?) and a Value (what can I contribute?), each by multiplying its representation with its own learned matrix. The phrases are intuition aids: all three are vectors.
  • Scores: the dot product of one Query with each Key. Large when they point the same way.
  • Scale, mask, softmax: divide scores by √dₖ, set future positions to −∞ (the causal mask), and turn the scores into positive weights that sum to 100%.
  • Mix: the output is the weighted sum of the Values. In the canoe sentence, 62% of “bank”’s mixture comes from “canoe”.
  • One such calculation is an attention head. Models run many heads in parallel and blend their outputs with one more matrix, the attention output projection: multi-head attention.
One attention head: the representation passes through W_Q, W_K and W_V to produce Query, Key and Value; Query-Key dot products are scaled, masked and softmaxed into weights that mix the Values. Representation each position W_Q W_K W_V Query Key Value compareq · k scale÷ √dₖ maskfuture → −∞ softmax→ weights mixΣ weight × V Updatefor this position SOFTMAX( Q·Kᵀ / √dₖ ) · V ONE ATTENTION HEAD
FIG 2.1One attention head (film ch. 11–18). Amber: learned matrices shared by every position. The dashed frame is what the film finally names an attention head.

Lab 2Dot products, scaling and softmax

film ch. 13–17 · 18:42

A 3-dimensional Query for “bank” meets four Keys. Edit any number. Scores are dot products; softmax turns them into weights. Values start as in film ch. 13.

Try this. Make the Key for “the” point the same way as the Query and watch it take weight from “canoe”. Set dₖ to 64 with scaling off, multiply every Query entry by 8, and see softmax go all-or-nothing; then turn scaling on (film ch. 15).

Lab 3Why the causal mask is a triangle

film ch. 16 · 23:54

Rows are the token doing the looking; columns are the token being looked at. Select a row to see what that position may use. Hatched cells are future positions, set to −∞ before softmax, which makes their weight exactly zero.

Go deeper

Part III · film 32:05 – 44:08

The transformer

Attention moves information between positions. The rest of a layer works on each position by itself, and residual connections keep every update.

  • The MLP (feed-forward network) expands each position’s vector (often ×4), bends every number with a nonlinear activation function, and compresses it back. Same weights at every position; no mixing between positions.
  • Why the bend: two matrices in a row equal one matrix. Without a nonlinearity, depth adds nothing.
  • Residual connections add each block’s output to the running representation instead of replacing it. Think of each position as a stream that blocks read from and write into.
  • Normalization rescales vectors before each block (RMSNorm divides by the root mean square) so values stay in a stable range.
  • A transformer stacks dozens of these layers. Parameters stay fixed between requests; activations exist only for the current text.
A pre-norm transformer layer: residual stream flows upward; normalization, multi-head attention and an addition; then normalization, MLP and another addition. IN OUT RESIDUAL STREAM Normalize Multi-head attentionbetween positions + Normalize MLPwithin each position + ONE LAYER · ×N
FIG 3.1A common “pre-norm” layer (film ch. 26). Exact designs and orderings vary between model families.

Lab 4Where the billions of parameters live

film ch. 22, 28, 55

Estimate a transformer’s parameter count from its shape. Per layer: attention holds 4·d² (Query, Key, Value and output matrices); the MLP holds 2·r·d². Embeddings add V·d, and a separate output projection adds another V·d. Biases and norms are ignored. Presets reproduce published totals closely.

A parameter is one learned number. Reading 6.7 billion of them aloud at one per second would take about 212 years.

Go deeper

Part IV · film 44:08 – 57:25

From numbers back to words

The last position’s vector becomes one score per vocabulary token, then probabilities, then one chosen token that is appended before everything runs again.

  • The output projection turns the final representation into a logit for every token in the vocabulary: raw scores that can be negative.
  • Softmax, the same operation used inside attention, turns logits into a probability distribution. A token’s probability measures how well it fits the text, not whether a statement is true.
  • Decoding: greedy picks the top token; sampling picks in proportion to probability. Temperature reshapes the distribution; top-k and top-p cut the long tail.
  • Autoregressive: each chosen token is appended and the model runs again. The KV cache keeps earlier Keys and Values so only the new position is computed; it holds activations, not parameters, and is discarded afterwards.

Lab 5Temperature, top-k, top-p and sampling

film ch. 32–35 · 46:02

The film’s illustrative next-token distribution after “The canoe finally reached the”, over a 50,000-token vocabulary. At temperature 1 with no filters it reproduces the film: bank 58.5%, shore 29.0%, river 3.2%.

no samples yet
Try this. Push temperature to 0.1: sampling becomes greedy. Push it to 2.5: the tens of thousands of long-shot tokens together take most of the probability, which is why high temperatures produce nonsense. Then set top-p to 0.90 and see only bank, shore and river survive (film ch. 35).
Sequence diagram of generation with a KV cache. Application Tokenizer Transformer KV cache Sampler 1 · prompt text 2 · token IDs (all n positions) 3 · forward pass overevery position (“prefill”) 4 · store K, V for n positions 5 · logits at the last position 6 · chosen token appended (softmax → sample) LOOP · ONE ITERATION PER NEW TOKEN 7 · only the newest token’s ID 8 · Query reads cached Keys 9 · weights mix cached Values 10 · append its own K, V (cache grows) 11 · logits for the new position 12 · next token → repeat until a stop token or the length limit PARAMETERS NEVER CHANGE HERE · THE CACHE IS DISCARDED WHEN THE TEXT IS FINISHED
FIG 4.1Generation as a sequence diagram (film ch. 36–38). Causal masking is what makes caching possible: earlier positions never look at later ones, so their Keys and Values never change.

Go deeper

Part V · film 57:25 – 68:07

Training

Nobody types in the billions of numbers. They start random and are nudged, one small step at a time, toward predicting text better.

  • Targets are free. In “The largest planet in our solar system is Jupiter”, every token is the target for the text before it. The causal mask lets one forward pass train every position at once.
  • Loss = −log(probability given to the correct token). 100% → 0; 1 in 50,000 → 10.8. Averaged over many predictions this is cross-entropy.
  • The gradient gives, for every parameter, how the loss would change under a tiny nudge. Backpropagation computes all of them in one backward sweep using the chain rule.
  • The optimizer steps each parameter against its slope (learning rate × gradient; Adam/AdamW refine this). Parameters change only at this step, and only during training.
The training loop: input text, forward pass, prediction, compare with target, loss, backpropagation, gradients, optimizer updates parameters, repeat. Repeat millions of steps trillions of tokens
FIG 5.1The training loop (film ch. 48). Only the optimizer step changes parameters; everything else computes activations.

Lab 6Loss as a penalty curve

film ch. 43 · 59:46

Slide the probability the model gave to the correct next token. The penalty climbs steeply as the probability approaches zero.

4.61loss = −ln p

Lab 7Gradient descent on one parameter

film ch. 45–47 · 61:35

The film’s one-parameter loss curve. Each step moves w by −(learning rate × slope). Try a small rate, a good one and one that is too large.

0steps
2.60w
—loss
—slope
Try this. At learning rate 0.3 the ball settles in a valley. At 2.0 it overshoots and bounces (film ch. 47). Start at w = −1.5 and it may settle in a different, higher valley: the gradient is only a local guide.

Go deeper

Part VI · film 68:07 – 74:51

What it all means

Training builds the model. Inference uses the model. Keeping those apart explains assistants, memory, hallucinations and scale.

Training versus inference side by side. ⟲ TRAINING · BUILDS THE MODEL ▸ INFERENCE · USES THE MODEL
FIG 6.1Training versus inference (film ch. 52). A chat does not rewire the model; what it “knows” about your conversation is the text in its context window.
  • Post-training (fine-tuning on example conversations, preference training, reinforcement learning, safety and tool-use training) turns a raw text continuer into an assistant. Recipes differ between developers.
  • Hallucination is a fluent, plausible continuation that is false or unsupported. Nothing in the generation loop checks claims against the world.
  • Grounding puts outside information (search results, retrieved documents, tool outputs) into the context window. It reduces errors but does not eliminate them.
  • Knowledge is distributed across many parameters, not stored as rows or single neurons. Interpretability research finds some features and circuits; most of the picture is still open.
Sequence diagram of retrieval-augmented generation. Person Application Search / retrieval Language model 1 · question 2 · search for relevant sources 3 · passages + where they came from 4 · context window = instructions + passages + question 5 · generate, token by token (Fig 4.1) 6 · answer that can cite its sources 7 · answer (still worth checking)
FIG 6.2Grounding with retrieval (film ch. 53). The model is unchanged; better inputs make plausible and true line up more often.

Go deeper

Part VII · film 74:51 – 77:28

Beyond the basics

Three ideas you will hear about constantly, each a small change to something you already understand.

Mixture of Experts · ch. 56

Many smaller MLPs (“experts”) plus a learned router that sends each token to a few of them. Large total capacity, a fraction of the compute per token; all experts still sit in memory.

Quantization · ch. 57

Storing each parameter with fewer bits (16 → 8 → 4) by rounding to fewer allowed values. Smaller and often faster, at some cost in precision.

Speculative decoding · ch. 58

A small draft model guesses several tokens; the large model checks them in one pass and keeps the ones it agrees with. Done correctly, the output distribution is unchanged, only faster.

Go deeper

Self-test · after the whole film

Explain it back

The film was designed so that a beginner could answer every one of these. Say your answer out loud or write it down first, then open the model answer. Tick the box when you could explain it without looking.

Reference

Glossary

Every term the film teaches, with its on-screen definition and the chapter that earns it.

TermDefinitionCh.

Continued study

Where to go next

Three paths, depending on what you want to do with this understanding.

Understand more deeply

No code required.

  1. 3Blue1Brown · Neural networks series
  2. Transformer Explainer with your own sentences
  3. The Illustrated Transformer
  4. LLM Visualization for the full scale picture

Build one yourself

Some Python helps.

  1. Neural Networks: Zero to Hero (micrograd → GPT)
  2. nanoGPT: train a small model
  3. Hugging Face LLM Course
  4. Dive into Deep Learning as a reference textbook

Follow the research

Papers, in reading order.

  1. Attention Is All You Need
  2. Scaling Laws and Chinchilla
  3. InstructGPT for post-training
  4. Transformer Circuits Thread for interpretability