A model is just two files
Take Llama 2 70B, Meta's open-weights model. It's a folder with a parameters file (the weights) and a run file (~500 lines of C, no dependencies). Copy both to a laptop, offline, and you can talk to it.
This file is only the weights — a giant list of numbers. The 500-line run.c never changes size; only the parameters file grows with the model.
Getting those numbers is the hard part
Running the model is cheap. Training it — finding the right 140GB of numbers — means compressing a huge chunk of the internet into that one file. For Llama 2 70B: ~10TB of text, ~6,000 GPUs, ~12 days, ~$2M. Today's frontier models are 10×+ bigger than this.
What actually happens between the two files
The talk skips this on purpose — it's the part that turns "two files" into a working brain. Four pieces, in the order data actually moves: how it's built at a glance, tokenization (text → integers), embeddings (integers → vectors), and position + self‑attention (vectors → a prediction).
How an LLM gets built, at a glance
Everything below is one of those four stages, zoomed in.
Tokenization — text becomes integers, once, up front
A neural net only takes numbers. A tokenizer is a fixed, frozen program — not learned during training — that turns text into a list of integer IDs, built by repeatedly merging the most frequent adjacent pair (byte‑pair encoding).
training text: "the cat in the hat" — every character starts as its own token.
| model | vocabulary size |
|---|---|
| GPT‑2 | 50,257 |
| Llama 2 | 32,000 |
| GPT‑4 | ~100,000 |
| Llama 3 | 128,256 |
Tokenization always succeeds — the vocabulary's first 256 entries are literally every byte value, so any text can be spelled out letter by letter in the worst case. Whether the model has any learned meaning for a chunk is a separate question.
| id | token |
|---|---|
| 0 | ! |
| 1 | " |
| 2 | # |
| 3 | $ |
| 4 | % |
| ⋯ every printable byte gets its own id ⋯ | |
| 255 | � ← single‑character / raw‑byte tokens end here |
| 256 | ␣t |
| 257 | ␣a |
| 258 | he |
| 259 | in |
| 260 | re |
| ⋯ thousands more merges, each built from earlier ones ⋯ | |
| 298 | ent |
| 299 | ␣n |
| 300 | ␣the |
| ⋯ continues to the vocabulary's target size ⋯ | |
| 2420 | ␣text |
| 50256 | <|endoftext|> ← special token, GPT‑2's last id |
this is an illustrative excerpt, not the full list — Llama 2's vocabulary has 32,000 entries, Llama 3 expanded that to 128,256. Every one is built the same way: merge the most frequent adjacent pair, repeat.
for a from‑scratch walkthrough of this exact algorithm: BPE from scratch ↗
Embeddings — integers become vectors
Each token ID looks up one row in a learned table E: 32,000 rows (one per Llama 2 token), 8,192 columns wide. Stack the rows for your input, in order, and you get the matrix X the network actually runs on.
Every column is a learned feature — nobody names them.
Real embeddings have 4,096 columns (Llama 2) and every one of them is just some direction gradient descent found useful for predicting the next word — not something a human labeled. Rows start as random numbers; every time a token shows up in a training batch, backprop nudges its row a little, so tokens used in similar contexts drift toward similar rows.
Nobody can point at column #2,301 and say what it means. But to build intuition, imagine we could — say a handful of the 4,096 columns happened to behave like these recognizable dials. This table is illustrative, not real output from a model.
| token | royalty | gender (+male/−female) | animate | size | noun-ness |
|---|---|---|---|---|---|
| "king" | 0.90 | 0.80 | 0.90 | 0.10 | 0.90 |
| "queen" | 0.90 | −0.80 | 0.90 | 0.00 | 0.90 |
| "man" | 0.10 | 0.80 | 0.90 | 0.20 | 0.90 |
| "woman" | 0.10 | −0.80 | 0.90 | 0.10 | 0.90 |
| "cat" | −0.60 | 0.00 | 0.90 | −0.60 | 0.90 |
| "the" | 0.00 | 0.00 | −0.90 | 0.00 | −0.90 |
| "sat" | 0.00 | 0.00 | 0.00 | 0.00 | −0.90 |
real columns: 4,096 · all uninterpretable · found automatically, not designed
Slice out just two of those illustrative columns — royalty and gender — and plot each token as an arrow from the origin. The pattern that falls out is exactly why king − man + woman ≈ queen works: "royalty" shifts king away from man by the same offset it shifts queen away from woman, and "gender" shifts king to queen by the same offset it shifts man to woman. Same shape, different starting point.
This lookup alone is not "understanding" anything yet — it has no notion of order (the rows for "cat sat" and "sat cat" are identical, just stacked differently) and no notion of context (the row for "bank" is the same in "river bank" and "bank account"). Both of those get fixed next.
Position + self‑attention — order and context, added back in
Transformers process every token in parallel, not one at a time like older RNNs — which is fast, but means the network has no built‑in sense of order. So position is added back in explicitly: a second vector P, one per position, is added to each token's embedding: X = embedding + P.
X still doesn't know about other tokens yet — each row is still just "this token, at this position," computed alone. Self‑attention is the step that lets every row read every earlier row and mix in what's relevant, producing the final, context‑aware vectors y₀ … yₙ that actually flow into the rest of the network.
for a deeper walkthrough of this exact mechanism: Contextual Transformer Embeddings using Self‑Attention ↗
a token's query is compared against every earlier token's key; softmax turns those scores into weights that sum to 1; the token's new vector is that weighted mix of everyone's value. Future tokens are masked out — at prediction time, they don't exist yet. (Real models do this 64 times in parallel, one "head" per kind of relationship — this shows one.)
"Predict the next word" is a trick
The training task sounds trivial — guess the next word. But to guess well on real text, the model is forced to absorb facts, not just grammar.
Ruth Handler invented the ▁▁▁▁▁
Grammar alone can't answer this — only knowing who Ruth Handler was can.
Training and chatting run different loops
Let’s say a user types "Summarize this document." and presses send. To the user, the reply just starts appearing a moment later. On the assistant side, two things are working together. The external code (the server around the model) prepares the text and runs the generation loop. The model (a transformer with fixed, already-trained weights) does the math. This diagram follows that one short message through every step, and marks which work happens in parallel (all tokens at once) and which happens sequentially (one step after another).
We built it, we don't fully get it
We know every math operation inside the network. What we can't explain is how ~100 billion numbers cooperate to know things. We can only measure whether it works — so behavior is judged empirically, by testing it, not by reading the weights.
— mostly inscrutable —
Same fact, asked backwards. Knowledge learned in one direction doesn't automatically work in the other — a strange, still not fully understood limit of how these models store facts.
Same training, a different dataset
Stage two keeps the exact same next-word training — it just swaps the data. Instead of raw internet, the model trains on ~100,000 human-written Q&A conversations. That's the entire trick behind turning a document-completer into an assistant.
The whole recipe, end to end
Click a stage to see what happens in it.
Stage 3, briefly: instead of writing answers, labelers just compare a few model-generated answers and pick the best — easier than authoring one from scratch. This comparison data trains the model further; OpenAI calls this RLHF.
it's much easier to spot the better haiku than to write one — that gap is what stage 3 exploits