Two Files, One Brain
Part 1 of 3 · How LLMs work

Two files, one brain

A large language model is, physically, just two files on a disk. Everything else — the knowledge, the voice, the "thinking" — is what happens when you run them. This walks through how, one live piece at a time.

based on Andrej Karpathy's "Intro to LLMs", Nov 2023
🗂️
parameters
140 GB
the weights — everything the model "knows," as one long list of numbers
📄
run.c
~500 lines
the code that runs those numbers — no internet required
01Inference

A model is just two files

Take Llama 2 70B, Meta's open-weights model. It's a folder with a parameters file (the weights) and a run file (~500 lines of C, no dependencies). Copy both to a laptop, offline, and you can talk to it.

A MacBook running a language model locally
the whole model, running on a laptop — no network needed
Terminal chatting with the model
"chat with a large language model" — plain text in, plain text out
drag: how big is the parameters file?70B
each weight stored as 2 bytes 140 GB

This file is only the weights — a giant list of numbers. The 500-line run.c never changes size; only the parameters file grows with the model.

02Training

Getting those numbers is the hard part

Running the model is cheap. Training it — finding the right 140GB of numbers — means compressing a huge chunk of the internet into that one file. For Llama 2 70B: ~10TB of text, ~6,000 GPUs, ~12 days, ~$2M. Today's frontier models are 10×+ bigger than this.

A visualization of a large web crawl
~10TB of crawled text
A GPU data center
~6,000 GPUs, ~12 days
press play: compress the internet
🌐
raw text
10 TB
→
🖥️
GPU cluster
12 days
→
🗜️
parameters.zip
140 GB
~71×
smaller — but lossy. Not a real zip: it keeps the gist, not an exact copy.
03Inside the network

What actually happens between the two files

The talk skips this on purpose — it's the part that turns "two files" into a working brain. Four pieces, in the order data actually moves: how it's built at a glance, tokenization (text → integers), embeddings (integers → vectors), and position + self‑attention (vectors → a prediction).

03·A

How an LLM gets built, at a glance

raw data→ learn structure by predicting→ refine behavior→ serve it

Everything below is one of those four stages, zoomed in.

AI TRAINING PIPELINE How a Large Language Model Gets Built 01 Pre-Training COLLECT RAW DATA Web Books Wikipedia Code repos Datasets raw dataset Clean & filter remove spam, duplicates, unsafe text, PII Tokenize — BPE vocab, 32K–128K entries ID TOKEN 1212This 318␣is 617␣some 2420␣text Llama 2: 32,000 tokens Llama 3: 128,256 tokens Embeddings id 1212 → + = input X token id token vector position vector to model embedding matrix shape: vocab_size × d_model e.g. 32,000 × 4,096 (Llama 2 7B) learned: similar tokens → similar vectors 02 Training predict the next token, billions of times Tokens input X Network Predict Loss Update next batch of tokens, weights slightly changed what actually gets multiplied y = x · W + b x — this layer's input vectors W — learned weight matrix b — learned bias (optional) weights, before → after one update 0.42 -1.1 0.08 → 0.39 -1.0 0.11 a nudge in the direction that reduces error Scale — billions of prediction steps Hardware — clusters of GPUs / TPUs Duration — weeks to months 03 Fine-Tuning Supervised fine-tuning (SFT) curated instruction → response pairs "Explain photosynthesis simply" → "Plants use sunlight to…" RLHF / RLAIF (alignment) Generate multiple replies Human ranks Reward model Policy update PPO / DPO updated model generates the next round what changes here same weight-update math as pre-training — just trained on much smaller, curated data, optimizing for "helpful and safe" instead of "predicts internet text well." 04 Deployment frozen weights, compressed, served via API Your prompt Tokenize Model Generate one token at a time Detokenize Response loops per token until the reply is complete what's added at serving time Quantize smaller, faster RAG fetch live facts Guardrails safety filters nothing trains here weights are frozen — this is the same Q/K/V + softmax math from training, just run forward, once per new token.
One pass, four stages: raw text becomes IDs and vectors (01), the network learns by repeatedly predicting and correcting (02), curated feedback aligns it to be helpful and safe (03), then the frozen model answers your prompts one token at a time (04).
03·B

Tokenization — text becomes integers, once, up front

A neural net only takes numbers. A tokenizer is a fixed, frozen program — not learned during training — that turns text into a list of integer IDs, built by repeatedly merging the most frequent adjacent pair (byte‑pair encoding).

press merge: build a tiny vocabulary

training text: "the cat in the hat" — every character starts as its own token.

how real words get chopped
click a word

modelvocabulary size
GPT‑250,257
Llama 232,000
GPT‑4~100,000
Llama 3128,256

Tokenization always succeeds — the vocabulary's first 256 entries are literally every byte value, so any text can be spelled out letter by letter in the worst case. Whether the model has any learned meaning for a chunk is a separate question.

a real vocabulary, excerpted (GPT‑2 style)0 – 50,256
idtoken
0!
1"
2#
3$
4%
⋯ every printable byte gets its own id ⋯
255�  ← single‑character / raw‑byte tokens end here
256␣t
257␣a
258he
259in
260re
⋯ thousands more merges, each built from earlier ones ⋯
298ent
299␣n
300␣the
⋯ continues to the vocabulary's target size ⋯
2420␣text
50256<|endoftext|>  ← special token, GPT‑2's last id

this is an illustrative excerpt, not the full list — Llama 2's vocabulary has 32,000 entries, Llama 3 expanded that to 128,256. Every one is built the same way: merge the most frequent adjacent pair, repeat.

for a from‑scratch walkthrough of this exact algorithm: BPE from scratch ↗

03·C

Embeddings — integers become vectors

Each token ID looks up one row in a learned table E: 32,000 rows (one per Llama 2 token), 8,192 columns wide. Stack the rows for your input, in order, and you get the matrix X the network actually runs on.

Every column is a learned feature — nobody names them.

Real embeddings have 4,096 columns (Llama 2) and every one of them is just some direction gradient descent found useful for predicting the next word — not something a human labeled. Rows start as random numbers; every time a token shows up in a training batch, backprop nudges its row a little, so tokens used in similar contexts drift toward similar rows.

Nobody can point at column #2,301 and say what it means. But to build intuition, imagine we could — say a handful of the 4,096 columns happened to behave like these recognizable dials. This table is illustrative, not real output from a model.

tokenroyaltygender (+male/−female)animatesizenoun-ness
"king"0.900.800.900.100.90
"queen"0.90−0.800.900.000.90
"man"0.100.800.900.200.90
"woman"0.10−0.800.900.100.90
"cat"−0.600.000.90−0.600.90
"the"0.000.00−0.900.00−0.90
"sat"0.000.000.000.00−0.90

real columns: 4,096 · all uninterpretable · found automatically, not designed

Slice out just two of those illustrative columns — royalty and gender — and plot each token as an arrow from the origin. The pattern that falls out is exactly why king − man + woman ≈ queen works: "royalty" shifts king away from man by the same offset it shifts queen away from woman, and "gender" shifts king to queen by the same offset it shifts man to woman. Same shape, different starting point.

king − man + woman ≈ queen, geometrically
Royalty Royalty Gender Gender KING QUEEN MAN WOMAN

This lookup alone is not "understanding" anything yet — it has no notion of order (the rows for "cat sat" and "sat cat" are identical, just stacked differently) and no notion of context (the row for "bank" is the same in "river bank" and "bank account"). Both of those get fixed next.

03·D

Position + self‑attention — order and context, added back in

Transformers process every token in parallel, not one at a time like older RNNs — which is fast, but means the network has no built‑in sense of order. So position is added back in explicitly: a second vector P, one per position, is added to each token's embedding: X = embedding + P.

X still doesn't know about other tokens yet — each row is still just "this token, at this position," computed alone. Self‑attention is the step that lets every row read every earlier row and mix in what's relevant, producing the final, context‑aware vectors y₀ … yₙ that actually flow into the rest of the network.

for a deeper walkthrough of this exact mechanism: Contextual Transformer Embeddings using Self‑Attention ↗

"the cat sat" the cat sat tokenize → ids 1820 3797 7731 map to learned embeddings v₁₈₂₀ v₃₇₉₇ v₇₇₃₁ + + + P₀ P₁ P₂ = = = X₀ X₁ X₂ self‑attention (tokens read each other) every token attends to every earlier token y₀ y₁ y₂
self‑attention, hands‑on — pick a token to see what it reads

a token's query is compared against every earlier token's key; softmax turns those scores into weights that sum to 1; the token's new vector is that weighted mix of everyone's value. Future tokens are masked out — at prediction time, they don't exist yet. (Real models do this 64 times in parallel, one "head" per kind of relationship — this shows one.)

04Why it learns

"Predict the next word" is a trick

The training task sounds trivial — guess the next word. But to guess well on real text, the model is forced to absorb facts, not just grammar.

A Wikipedia article about Ruth Handler used as training text
a random Wikipedia page — dates, relationships, inventions. To predict what comes next here, the model has to know this, not just pattern-match sentences.
try predicting the next word yourself

Ruth Handler invented the ▁▁▁▁▁

05The core idea

Training and chatting run different loops

Let’s say a user types "Summarize this document." and presses send. To the user, the reply just starts appearing a moment later. On the assistant side, two things are working together. The external code (the server around the model) prepares the text and runs the generation loop. The model (a transformer with fixed, already-trained weights) does the math. This diagram follows that one short message through every step, and marks which work happens in parallel (all tokens at once) and which happens sequentially (one step after another).

Diagram tracing a user message through tokenization, parallel prefill through the transformer, and sequential one-token-at-a-time generation, showing which steps run in parallel versus sequentially.
06The black box

We built it, we don't fully get it

We know every math operation inside the network. What we can't explain is how ~100 billion numbers cooperate to know things. We can only measure whether it works — so behavior is judged empirically, by testing it, not by reading the weights.

"who is Tom Cruise's mother?"
100B paramsnext-word machinery
— mostly inscrutable —
"Mary Lee Pfeiffer"
the "reversal curse" — click both
"Who is Tom Cruise's mother?"
"Who is Mary Lee Pfeiffer's son?"

Same fact, asked backwards. Knowledge learned in one direction doesn't automatically work in the other — a strange, still not fully understood limit of how these models store facts.

07Finetuning

Same training, a different dataset

Stage two keeps the exact same next-word training — it just swaps the data. Instead of raw internet, the model trains on ~100,000 human-written Q&A conversations. That's the entire trick behind turning a document-completer into an assistant.

toggle: base model vs. assistant model
An example labeled Q&A conversation used for finetuning
one of ~100,000 hand-written conversations a human labeler produces for finetuning — quality over quantity, the opposite of pretraining's data
08Summary so far

The whole recipe, end to end

Click a stage to see what happens in it.

1
Pretrain
2
Finetune
3
RLHF (optional)
4
Deploy
5
Monitor ↺

Stage 3, briefly: instead of writing answers, labelers just compare a few model-generated answers and pick the best — easier than authoring one from scratch. This comparison data trains the model further; OpenAI calls this RLHF.

Two candidate haikus, one marked as picked Two candidate haikus, one marked as picked

it's much easier to spot the better haiku than to write one — that gap is what stage 3 exploits

Excerpt of labeling instructions from the InstructGPT paper
an excerpt of the instructions given to human labelers — helpful, truthful, harmless
A leaderboard of chat models ranked by Elo rating
models ranked like chess players — win a head-to-head, gain rating