How a picture becomes a prediction

Scroll. Every stage is computed live. Hover, drag the sliders, press play.

00The whole workflow at a glance

in plain words

This is the whole journey of one picture: split into channels, filtered many times, shrunk, flattened, and turned into a probability for each class. This is also the single network every later chapter refers to — each one shows this same strip at the top with only its own layer lit, so you always know where you are. Play with the knobs below and watch the workflow, the neuron view, and every number recompute together, live.

every box is computed live on the 12×12 heart · click a box to jump to the chapter that explains it

compare with label → loss → backprop → update every weight — kernels + FC → repeat ↺
input → prediction · loops conv1 ⇄ conv2 twice before flatten

the same network drawn as neurons — every column responds to the knobs above · scroll sideways if it is wide

live results — every number recomputes from the knobs
notation

Both cᵢ and cₒ are channel counts, not layer numbers — that little subscript is a letter, in or out, never a digit.

cᵢ  “c-in”  = how many channels go in — decided for you by the previous layer
cₒ  “c-out”  = how many channels come out — your choice, and always equal to the number of filters
k  = kernel size (3 here)  ·  hᵢ, hₒ  = height going in, height coming out

So one filter is always k × k × cᵢ, and stacking cₒ of them gives the next layer its cᵢ. That hand-off is what the arrows across this map are tracing.

01A picture is a grid of numbers

in plain words

A picture is just a table of numbers. Each cell is one pixel; the number is its brightness, from 0 (black) to 255 (white).

image → matrix of pixels · shape (H, W) = (12, 12) · hover a pixel

what you see
→
what the computer sees (brightness 0–255)

02Color = 3 channels stacked

in plain words

A colour pixel is three numbers, not one: how much red, how much green, how much blue. So a colour image is three tables stacked, written as (C, H, W) = (3, 12, 12) — 3 channels, 12 rows, 12 columns.

tensor shape (C, H, W) = (3, 12, 12) · one 12×12 matrix for R, one for G, one for B

One pixel = 3 numbers. A red heart pixel is high in R, low in G and B; a baby‑blue pixel is high in B and G.

03Convolution: a small filter slides over the input

in plain words

A kernel is a tiny table of weights — 3×3 here. It slides over the input; at each position it multiplies its 9 weights with the 9 numbers underneath and adds them up, and that single sum is one output number. Doing this at every position builds a new, slightly smaller table called a feature map.

kernel vs. kernel window — two levels

A kernel window (or "slice") is a single k×k grid of weights that slides over one channel. In our case each window is 3×3 = 9 weights.

A kernel (also called a filter) is the full stack of windows needed to cover all input channels — one window per channel, glued together. With 3 input channels, one kernel = 3 windows = 3×3×3 = 27 weights. The kernel produces one output feature map.

input 3 × 12 × 12 (cᵢ × H × W) — the real heart from 01–02, real pixel values · kernel 3 × 3 × 3 (cᵢ × k × k), one 3×3 slice per input channel · output 1 × 10 × 10

the 3 kernel boxes = the weights

We have 3 kernel boxes, one per input channel. The numbers inside them are the weights — numbers we choose (later: numbers training chooses), not the result of applying them to the picture.

Together that is 3 × 3 × 3 = 27 weights + 1 bias. What a weight really is, and how training picks its value, comes in 03·R and chapter 10 — for now just keep the idea: kernel boxes = weights; the coloured sums below = what happens when we apply them.

out[y,x] = Σc Σi Σj  x[c][y+i][x+j] · w[c][i][j]  + b   — one 3×3 window per channel, 27 multiply-adds, summed across channels
output (feature map) 1×10×10 · hover any cell
what is a MAC?

MAC = Multiply‑ACcumulate: one multiplication followed by adding the result onto a running total, a ← a + w·x. It is the basic unit of work in a neural network, and the lecture counts cost in MACs.

In the demo above, one output cell costs 27 MACs: 0×204, 0×204, … , 2×204 … each product is added onto the sum → 334. The whole 10×10 map is 100 cells × 27 = 2,700 MACs, and with cₒ kernels it is 2,700·cₒ.

MACsconv = ci · kh · kw · ho · wo · co

Our case: 3 · 3 · 3 · 10 · 10 · 1 = 2,700. One MAC = 2 FLOPs (a multiply and an add), so hardware speeds quoted in FLOPS or OPS (operations per second) tell you how many of these a chip can do each second — that is how latency is estimated: #MACs ÷ chip speed. Note the two counts are different things: #parameters = numbers stored (27 here), #MACs = work done (2,700 here).

what the letters mean

hᵢ = height of the input (12 rows of the heart) · hₒ = height of the output feature map · k = kernel size (3) · the same letters with w are for width.
p = padding, a ring of extra cells added around the input — 0 here, explained in step 04 · s = stride, how far the window jumps each step — 1 here, explained in step 05.

Now, how do we know hₒ? We know k = 3, hᵢ = 12, p = 0, s = 1. There is one general equation that takes all of this into account:

ho = ⌊ (hi + 2p − k) / s ⌋ + 1

So in our case: hₒ = ⌊(12 + 2·0 − 3) / 1⌋ + 1 = 10. With no padding and stride 1 this is simply hᵢ − k + 1: the window fits 12 − 3 + 1 = 10 times across, so the map shrinks from 12 to 10.

cᵢ = "c‑in", the number of channels coming in to the layer — 3 here (R, G, B). cₒ = "c‑out", the number of channels coming out — equal to how many kernels you use (1 in this demo, so the output is 1 × 10 × 10). One kernel must cover all cᵢ channels, so it stores cᵢ·k·k = 3·3·3 = 27 weights (its own numbers) + 1 bias. At every window position each weight is multiplied by the pixel under it — 9 products per channel, 27 in total — and all 27 products plus the bias are added into one output cell. Same 27 weights, reused at all 100 positions of the 10×10 map: 27 weights, but 100 × 27 = 2,700 multiply‑adds.

in plain words

One kernel makes one feature map. If you want to detect several things — vertical edges, horizontal edges, blur — you use several kernels, and each produces its own output channel. Below we use 4 kernels (red detector, vertical edge, horizontal edge, blur) on the 3 input channels, so 4 feature maps come out. That is why the weights have shape (cₒ, cᵢ, k, k) = (4, 3, 3, 3): cₒ = 4 kernels, each looking at all cᵢ = 3 input channels through a 3×3 window.

Both subscripts are letters, not numbers: cₒ = “c-out”, how many channels come out (= how many kernels you use); cᵢ = “c-in”, how many channels go in.

many kernels → many output channels · here cₒ = 4 kernels, cᵢ = 3 channels, k = 3 · weights shape (cₒ, cᵢ, k, k) = (4, 3, 3, 3)

in plain words · the loop

The output of one conv layer is just another image with more channels. So we do it again: the next layer takes the (cₒ, 10, 10) maps as its input, applies new kernels, and gives even smaller maps with even more channels. Each round the picture gets spatially smaller (fewer rows/columns) and deeper (more channels), because we trade "where" for "what". Pooling and stride (06, 05) make the shrinking faster.

⚠ what if we only have one kernel window?

A kernel is a detector. "No channels" isn't quite right: one kernel → one feature map, and that map is one channel, so layer 2 would receive cᵢ = 1. Not "no channels" — only one. But the point stands: you've collapsed everything the layer could notice into a single number per position.

Why one kernel is a bad idea: the red‑detector kernel answers only "is this pixel red?". If that's your single output channel, layer 2 has nothing else to work with — no edges, no corners, no texture, no green/blue information. Everything but one feature is thrown away. To recognise a heart you need many detectors at once (edges of several orientations, colour blobs, curves), so you use many kernels → many channels, and each later layer combines them into more complex features.

layer L reads the (cₒ, h, w) of layer L−1 · every layer: hₒ = hᵢ − k + 1 · #weights = cₒ·cᵢ·k·k · computed live on the heart

03·RQuick revision: what is a neural network?

in plain words

A neuron takes some inputs, multiplies each by its own weight, adds them up, adds a bias, and squashes the sum through an activation function. That is all a neuron ever does. A network is many neurons in layers, each layer feeding the next.
In our case: the inputs are the 27 pixel values under the 3×3×3 window, the weights are the 27 kernel numbers, and one output cell of the feature map is exactly one neuron. Parameters = weights + biases = the numbers training will change.

one neuron = the top-left cell of the feature map · net = Σ wᵢ·xᵢ + b · o = φ(net) · the neuron follows the sliding window of the demo above

or hover / play in the demo above — this neuron follows the window
how one neuron becomes an activation

Every output cell is one such neuron. Its net (before φ) goes into the left map; its activation o = φ(net) goes into the right map — that right map is what the next layer receives. Details in Activation: a non-linear squeeze after every linear layer.

in plain words · the whole network

Stack neurons into layers. Input layer = the raw numbers (our 3·12·12 = 432 pixel values). Hidden layers = neurons whose outputs nobody sees directly, they only feed the next layer. Output layer = one neuron per class (heart, leaf, ball). When every neuron connects to every neuron before it, the layer is fully connected. A conv layer is the same picture, except each neuron only connects to a small window and all neurons of one map share the same 27 weights — that is why it needs far fewer parameters.

The network, drawn as neurons — reusing the same live diagram from 00 · the whole workflow at a glance, so counts and connections always match the knobs there.

how many neurons in a conv layer?

Kernels set the depth (how many feature maps); the input size, kernel size, padding and stride set each map's height × width. The neuron count is just their product:

#neurons = co · ho · wo,   where  ho = ⌊ (hi + 2p − k) / s ⌋ + 1

Our heart, layer 1: cₒ = 4 kernels, hᵢ = 12, k = 3, p = 0, s = 1 → hₒ = 12 − 3 + 1 = 10 → #neurons = 4 · 10 · 10 = 400.

Another example — a 10×10×3 image with 2 kernels: hₒ = 10 − 3 + 1 = 8 → #neurons = 2 · 8 · 8 = 128 (weights = 2·27 = 54, biases = 2). The 2 kernels give 2 maps; each map is 8×8 no matter how many kernels there are.

04Padding: add a border so the output does not shrink

in plain words

Every convolution eats k−1 pixels from the border, so the map shrinks. Padding adds a ring of extra cells — usually zeros — around the input, so the output keeps its size and edge pixels get looked at as often as centre ones.

input with padding p
⊛ 3×3 →
output
does padding change the receptive field?

Padding — keeps the map size and stops edge pixels being eaten. Preserves data, not RF.

What does change the receptive field is covered in 07.

05Stride: how far the window jumps

in plain words

Stride is how far the window jumps after each step. Stride 1 visits every position; stride 2 skips every other one and produces a map about half the size — cheaper, and it lets later layers see a bigger region.

input 7×7 (+ padding)
→
output
(try s = 2 → "stride‑2", the downsampling from step 07)

06Pooling: shrink the map, keep the strongest signal

in plain words

Pooling shrinks a feature map with no weights to learn: look at a 2×2 window and keep only the biggest value (max) or the average. It keeps the strong signals and throws away exact position.

2×2 window, stride 2 · no weights to learn · each channel pooled on its own · hover the output

4×4
→
2×2
click a 2×2 output cell, or edit any input number
does the max change the receptive field?

Pooling's max — keeping the biggest value preserves the strongest feature; that is separate from the downsampling that helps RF.

What does change the receptive field is covered in 07.

the same network, now with a pooling layer

Here is our network again, but with a pool layer inserted after the conv feature maps. Pooling has no neurons of its own to learn — it just shrinks each of the 4 maps from 10×10 to 5×5, so the 4 feature maps become 4×5×5 = 100 neurons before the dense layer. One node = one neuron, dashed pointer marks it.

The workflow again, greyed out except pooling — same input → … → softmax pipeline as chapter 00, here dense = fully‑connected layer (has weights); it is not the activation.

07Receptive field: how much input one output neuron sees

in plain words

The receptive field is how many input pixels one output number depends on. Each 3×3 conv layer (a full conv layer whose kernels use a 3×3 window) adds 2 more pixels, so stacking layers — or using stride — lets a single neuron eventually see the whole heart.

why it matters

A neuron can only use what is inside its receptive field — everything outside it is invisible to that neuron. So the receptive field is literally how much of the image this neuron is allowed to know about.

To recognise a whole heart, some neuron must see the whole heart at once. If a neuron's field is only 5×5 pixels it can spot a tiny edge but never "heart", because a heart does not fit in its view. Recognising a shape needs a receptive field at least as big as the shape.

It decides how deep / how much pooling you need. One 3×3 conv layer sees only 3×3; covering a 224×224 image with 3×3 conv layers alone would take ~111 layers. Each stride‑2 (or pool) step roughly doubles how fast the field grows, so you reach "sees everything" in a handful of layers instead of a hundred — that trade‑off is the whole reason downsampling exists.

It explains the early‑to‑deep story. Early layers (small field) can only see edges, corners, colour; deep layers (large field) can see objects and layout. The growing receptive field is exactly why a CNN builds from simple features to complex ones — and if your object is bigger than the final field, accuracy silently caps out no matter how long you train.

so how many layers do we need?

Not a random choice. More layers = more understanding of the picture, but there is a minimum to start from: one output must be able to see the whole object. That minimum is exactly what the receptive field tells us. Setting RF ≥ input size and solving RF = L·(k−1)+1 for L:

L ≥ (input − 1) / (k − 1)

A 150×150 image with 3×3 kernels: L ≥ (150−1)/(3−1) = 74.5 → 75 layers. Bigger kernels help — with 10×10: 149/9 ≈ 16.6 → 17 layers — but a big kernel costs about (k/3)² more per layer, so stacking small kernels is usually cheaper.

Small‑and‑deep beats big‑and‑shallow. The receptive field after L layers is RF = L·(k−1)+1, so two 3×3 layers reach the same RF as one 5×5 layer:

two 3×3: RF = 2·(3−1)+1 = 5  =  one 5×5: RF = 1·(5−1)+1 = 5

Same reach, but the weights differ: two 3×3 layers use 2·(3·3) = 18 weights, one 5×5 uses 5·5 = 25. The stacked pair is cheaper and adds an extra activation (more non‑linearity), which is why modern CNNs stack 3×3 convs instead of using large kernels.

The real fix is downsampling. With a stride‑2 or 2×2 pool every stage, each stage halves the map and roughly doubles the reach, so you only need about log₂(input) stages: 150 → 75 → 38 → 19 → 10 → 5 → 3 → 1, i.e. ~7 stages, not 75 layers. That is why real CNNs are ~10–50 layers, not hundreds — and why the lecture says: for large images, downsample inside the network. Extra layers beyond this minimum are added for richer features, then capped by compute and memory.

which knobs change the receptive field

Only three things change the receptive field: kernel size k, number of layers L, and stride/pooling (downsampling). Everything else is useful for other reasons.

Kernel k — bigger k adds k−1 to RF each layer, but ~(k/3)² more memory.

Layers L — RF = L·(k−1)+1, and deeper = richer features; costs more compute. (More kernels ≠ more layers: extra kernels add feature maps in one layer, they do not grow RF.)

Stride / pooling — downsample: RF of later layers grows twice as fast. Too aggressive → detail lost.

the recipe: 3 knobs, and how each moves the receptive field

The rule of thumb is small kernel + moderate depth + downsampling — not "medium everything". Three separate knobs:

1 · Kernel size k → small (3×3). Bigger k grows RF faster per layer (RF adds k−1 each layer), but costs ~(k/3)² more weights and blurs detail. Good: 3×3 stacked reaches any RF cheaply. Bad: one 7×7 reaches RF 7 in a layer but for 49 weights vs 3 stacked 3×3 layers reaching RF 7 for 27 — wasteful.

2 · Depth L → moderate, and it's a result not a target. RF = L·(k−1)+1, so more layers = more RF and richer features. Good: enough layers that RF ≥ object size. Bad: too few → RF never covers the heart, it literally cannot be recognised; too many → huge cost and, past ~20 plain layers, training degrades.

3 · Downsampling (stride‑2 / pool) → yes. This is the knob people forget. Each halving stage makes RF grow geometrically, not by a flat k−1. Good: 150×150 needs ~7 stages instead of 75 layers. Bad: overdo it and the map becomes 1×1 too early, throwing away spatial detail before features are built.

On our 12×12 heart: plain 3×3, no downsampling → RF ≥ 12 needs L ≥ 6 layers. Add one 2×2 pool after 2 layers → RF passes 12 in ~2 conv layers (what chapter 11 actually does). Same k, far fewer layers — because downsampling did the heavy lifting.

the base design: how deep must the net be for a meaningful RF?

Getting a receptive field that covers the whole heart is a design decision. Here is the base design with the minimum depth needed. With plain 3×3 convs (no downsampling) RF = L·(k−1)+1 = 2L+1, and we need it ≥ the picture size (12), because the heart is 12×12:

2L + 1 ≥ 12 → L ≥ 5.5 → L = 6 conv layers

Six layers on a 12×12 image is wasteful, so we add padding + pooling and reach RF ≥ 12 in far fewer layers. This base design (3×3 convs, padding 1, one 2×2 pool every couple of layers) is the minimum that gives a meaningful RF:

stageopmap sizeRF on input
input—12×121
conv1 (4 kernels, 3×3, p=1)conv12×12×43
conv2 (3×3, p=1)conv12×125
pool 2×2 s=2pool6×66
conv3 (3×3, p=1)conv6×610
pool 2×2 s=2pool3×312 ✅

08Activation: a non-linear squeeze after every linear layer

in plain words

After every linear operation we squash the result with a non-linear function. Without it, stacking layers would still add up to one big linear equation and could never learn shapes. ReLU is the simplest: negatives become 0, positives stay as they are.

y = f(Σ wᵢxᵢ + b) · the map is the heart's vertical-edge output, every 2nd cell (÷50) · click a function · red = negative, teal = positive

why activation = "non‑linearity"

Conv and fully‑connected layers are linear (output = Σ weight×input + b). Stacking linear layers collapses into one linear layer — 50 layers with no activation = the power of 1, only straight‑line boundaries, never a curved shape like a heart. A non‑linear activation (ReLU bends everything below 0 to 0) is the hinge between layers that stops the collapse, so each new layer actually adds power. Without it, depth would be pointless.

before f
→ f →
after f
does activation change the receptive field?

Activation — adds non‑linearity. Zero effect on RF.

What does change the receptive field is covered in 07.

09Flatten & dense: from feature maps to class scores

in plain words

After the last conv/pool, the network holds a stack of small feature maps — a 3‑D block (C × H × W). But the answer we want is just 3 numbers (heart / leaf / ball). Flatten unrolls that 3‑D block into one long 1‑D vector, and the dense (fully‑connected) layer turns that vector into the final class scores — every output neuron connects to every input number.

09·AFlatten: from feature maps to one vector

from the base design in 06 — how many feature maps just before flatten?

Channels grow as the map shrinks (trade "where" for "what"). Doubling the kernels each conv, the last conv outputs 16 feature maps, each 3×3:

stagemap size#feature maps (channels)
input12×123 (R,G,B)
conv1 (4 kernels)12×124
conv2 (8 kernels)12×128
pool6×68
conv3 (16 kernels)6×616
pool3×316

So just before flattening we have 16 feature maps, each 3×3 — a 3‑D block of shape (C×H×W) = 16×3×3.

the 3‑D block = C matrices stacked · here C=16, each 3×3 · scroll sideways

now we need one representation of that block

A dense layer needs a 1‑D list, not a 3‑D block. One representation is to flatten: unroll the block into a single long vector. That is what we chose here — 16 × 3 × 3 = 144 numbers.

flatten — just a reshape, no maths

Flatten reads the block in a fixed order (channel, then row, then column) and lays every number end‑to‑end into one column. Nothing is computed, nothing is learned — same numbers, new shape. Our 16×3×3 block becomes a vector of 16·3·3 = 144 numbers. Hover a cell to see where it lands.

16 feature maps · 3×3 each
→ flatten →
1 vector · 144×1
hover a map cell (or a vector cell) to link the two

09·BDense: from the vector to class scores

what dense does

A dense layer has one neuron per class. Each neuron sees all the flattened numbers, multiplies each one by its own weight (a synapse), adds them up, and adds a bias. Dense doesn't pick a class. It produces one new number per class.

What a dense layer does Four of the flattened numbers each connect to three neurons, one per class: heart, leaf and ball. Every connection is a synapse with its own weight. Each neuron sums its weighted inputs, adds a bias, and outputs one new number: 9.0 for heart, 2.0 for leaf, 1.0 for ball. No class is chosen. Flattened numbers 144 in the real model Neurons sum, then + bias New numbers one per class Synapse: each line has its own weight 3 1 2 4 Σ + b Σ + b Σ + b 9.0 2.0 1.0 heart leaf ball Dense doesn't pick a class: it only turns the inputs into one number per class

Synapse = one weight on one connection. Neuron = the sum of all its weighted inputs plus a bias. (The f in the drawing is skipped in this layer: the output is just the sum plus bias.)

how each score is computed
How the flattened vector becomes the z scores For each class, every input in the flattened vector x = 3, 1, 2, 4 is multiplied by that class's own weight, the products are summed, the bias is added, and the result is the score z. Heart gives 9.0, leaf gives 2.0, ball gives 1.0. x 3 1 2 4 144 numbers in the real model weight × input sum bias b z heart leaf ball 1.0 × 3 3.0 0.5 × 1 0.5 1.5 × 2 3.0 0.5 × 4 2.0 8.5 + 0.5 = 9.0 0.2 × 3 0.6 −0.3 × 1 −0.3 0.4 × 2 0.8 0.1 × 4 0.4 1.5 + 0.5 = 2.0 −0.1 × 3 −0.3 0.4 × 1 0.4 0.2 × 2 0.4 0.2 × 4 0.8 1.3 + −0.3 = 1.0 z = W·x + b: each class has its own row of weights and its own bias

z = W·x + b. Every class gets the same x but its own row of weights and its own bias. The drawing uses 4 inputs. The real vector has 144, so W is 3 × 144 = 432 weights, plus 3 biases.

what z is

z is the raw score of each class: one number per class (here heart 9.0, leaf 2.0, ball 1.0). It says how strongly the input matches what that class's weights look for, and bigger means more evidence. It is not a probability: it can be negative or bigger than 1, and the three numbers don't add up to anything. Only which one is largest, and by how much, matters. Which output is 'heart' is just our labelling of position 0. Training is what makes row 0 respond to hearts.

Next: turning the scores into a decision.

why the dense layer is expensive

Because it connects everything to everything, its weight count is cₒ·cᵢ — and cᵢ is the whole flattened block, which can be huge. In AlexNet the first dense layer alone holds ~38 million weights (9216 × 4096), far more than all the conv layers combined. That is why modern networks pool aggressively before flattening (smaller block → smaller vector → cheaper dense) or replace flatten+dense with global average pooling (average each map to one number, so C×H×W → C directly).

10Classification: turning one vector into one of three classes

in plain words After flatten the network holds one list of D numbers. Classifying means: from that list, decide heart, leaf or ball. The dense layer does it with three weighted sums — one per class — and the biggest sum wins. The rest of this chapter is about how the network learns the weights of those three sums: what a label must look like (one‑hot), how scores become probabilities (softmax), how "wrong" is measured (cross‑entropy), and how every weight is nudged (gradient descent) — image after image, epoch after epoch.

A · one vector in, one class out

the question Suppose the vector has only D = 6 numbers — pretend the conv layers measured six things about the picture. To classify it we need 3 scores, one per class. Each score is a weighted sum of the same 6 numbers with its own row of weights: zc = Σi W[c][i]·xi + bc. That is the dense layer of chapter 09 — a 3×6 matrix. The class is the row with the largest score (argmax). Drag the x sliders and watch the three sums; then press random weights to see what an untrained network does: it still answers — just wrongly. Training (part G) is how the matrix gets its right values.
x = a vector of D = 6 numbers (drag) · W = 3 rows × 6 weights, one row per class · cell colour = the product W[c][i]·xi (teal +, red −) · z = row sum · class = argmax z
← this is the whole problem of training: find the 3×D numbers in W that make the right row win for every image.

B · why the label is not heart = 1, ball = 2, leaf = 3

in plain words If you used a single number as the label, the network would treat classes as having an order and distance — it would learn that "ball" (2) is between heart (1) and leaf (3), and that leaf is further from heart than ball is. That's meaningless here — these are three unrelated categories, not points on a scale. Integer labels are fine when there is genuine order (mild / moderate / severe), but wrong for unrelated classes.
one output number on a line · drag the network's output · then press swap the order

C · what's used instead: one‑hot encoding

in plain words Each label becomes a vector of length 3 — one slot per class, a 1 in the slot of the true class, 0 elsewhere. Slot order is fixed once: [heart, leaf, ball]. This has exactly the shape of the network's output (3 scores), so label and prediction live in the same 3‑D space and can be compared number by number. And every pair of labels is the same distance apart — nobody is "between".
click an image → its label vector · slot order [heart, leaf, ball]
click a card…
heart [1,0,0] leaf [0,1,0] ball [0,0,1] √2√2√2
The three one‑hot vectors are the corners of an equilateral triangle: every pair differs in exactly two slots, distance √(1²+1²) = √2. No class is closer to any other — which is precisely what "three unrelated categories" should look like. Compare with the line in B, where the middle class was 1 away from both neighbours and the outer ones 2 apart.

D · softmax: 3 raw scores → 3 probabilities that sum to 1

in plain words The dense layer's three sums z are just numbers — could be −4 or 37. The label is [1,0,0]. To compare them we rescale z into three numbers that behave like the label: all between 0 and 1, summing to 1. That is softmax: pc = ezc / Σj ezj. The exponential makes every entry positive and stretches gaps (a lead of a few points becomes most of the probability); the division makes them sum to 1.
is softmax "a classification algorithm"? Not by itself. Softmax has no weights and changes nothing about which score is largest — the argmax of z and of p is the same row. The classifier is dense layer + softmax together, a model called softmax regression (multinomial logistic regression). The dense layer does the deciding; softmax's job is to turn the decision into probabilities so that it can be compared with a one‑hot label and a loss can be computed (part G). So: flatten → dense → softmax is exactly "a softmax classifier on top of the conv features".
drag the three raw scores · watch ez, the division, and that the probabilities always sum to 1
+2 to all → p unchanged (softmax only cares about differences). ×2 → same winner, but sharper: bigger gaps = more confident.
raw scores z (any real number)
probabilities p = softmax(z) — sum to 1

E · the dataset: 300 images, each with its one‑hot label

in plain words Training needs examples with answers. Here is a real dataset generated in your browser: 100 hearts, 100 leaves, 100 balls, each an 8×8×3 picture with a random shift, size, tilt, colour jitter and pixel noise — so no two are identical. Every image is flattened to D = 8·8·3 = 192 numbers (in the full network this vector would be the 144‑vector out of conv→pool→flatten, chapter 09; here we skip the conv part so the dense part is fully visible). We split the data: 80 of each class (240) for training, 20 of each (60, dashed border) for testing — the test images are never used to update weights; they only measure whether what was learned generalises.
all 300 images · solid border = train (240) · dashed = test (60) · click any image to make it the current example for part G
the dataset as a table · x = the flattened 192 numbers (three rows of 64 = the R, G, B channels, drawn as grey levels; hover a cell) · y = one‑hot label

F · nudging one weight: what a gradient and η are

in plain words The loss is one number for how wrong the network is. It depends on the weights, so plot it against a single weight and you get a curve. The gradient is the slope of that curve at the weight's current value (the red dashed line): it says which way is uphill, and how steeply. Each step moves the weight downhill, against the slope: w ← w − η · slope. The learning rate η sets the step size. Try a tiny η, then a big one. A real network does exactly this for every weight at once (part G).
  learning rate η
left: loss against one weight, the ball is the current w · right: loss after each step
the ball only moves downhill and never backtracks: once a step would send it back the way it came, it stops there for good. A small η settles exactly at the bottom of the valley; too big an η can overshoot past the bottom and freeze partway up the far slope instead.

G · training the classifier: forward → loss → backward → nudge → repeat

in plain words W starts random (part A, "untrained"). For one training image: run it forward to get p (the predicted class probabilities, from softmax); compare with the one‑hot y using the cross‑entropy loss L = −log ptrue (only the probability given to the correct class matters: ptrue = 1 → L = 0; ptrue = 0.7 → L = 0.36; ptrue = 0.1 → L = 2.3); go backward: for softmax + cross‑entropy the gradient on the scores is simply g = p − y, and the gradient on each weight is ∂L/∂W[c][i] = gc·xi; update every weight a small step against its gradient, W ← W − η·g·xᵀ (η = learning rate). In the full network this same g keeps flowing back through flatten into the conv kernels, so kernels and dense weights are all updated by the same loss. Then take the next image — or better, a shuffled batch of 16 and average their gradients — and repeat; one pass over all 240 training images is one epoch.
Training loop through the whole network Forward runs down the layers from input through conv, ReLU, pool, flatten, dense and softmax to the loss. Backward sends the gradient up from the loss through every layer to conv1. Only the layers with weights, conv1, conv2, conv3 and the two dense layers, are nudged. Then the loop repeats with the next image. Forward Backward Nudge Repeat with the next image input ReLU ReLU pool ReLU pool flatten ReLU softmax conv1 conv2 conv3 dense dense loss filter weights change filter weights change filter weights change W and b change W and b change Purple layers have weights, so they get nudged. Gray layers have none. They only pass the gradient on. The error starts at the loss and flows back up.
when the error propagates, does it hit the nearest layer first, or all the purple layers at once?

The gradient goes to the nearest layer first and moves backward one layer at a time, but no weights change during that trip. All the purple layers are updated together at the end.

Order.

  1. The loss and softmax produce p − y, which goes to the last dense layer.
  2. That layer computes the gradients for its own W and b, and also the gradient to pass further back.
  3. The next layer, ReLU, passes it on. Then comes dense, then flatten, then pool, and so on down to conv1.
  4. Once the gradient reaches conv1, an update step changes the weights of all five purple layers at once.
1 · training image
a real example from the training set (click any thumbnail in E to change it)
2 · flatten → xD = 192 numbers in [0,1] · row 1 = R, row 2 = G, row 3 = B · hover
3 · dense: z = W·x + bthree weighted sums, W is 3 × 192 = 576 weights + 3 biases
4 · softmax → pprobabilities, sum = 1 · outlined = argmax = the prediction
5 · compare with the true label yone‑hot, from the dataset table
6 · loss = −log ptruecross‑entropy: how far the probability of the correct class is from 1
7 · backward: g = p − ygradient on the scores · negative = "push this score up", positive = "push it down"
8 · update: W ← W − η · g · xᵀeach row c of W moves by −η·gc·x — the true class's row moves toward this image, the others away · teal +, red −
  batch size   learning rate η
"1 step on this image" updates with this image only (batch of 1) — watch ptrue rise and the loss fall for it. "1 batch" averages the gradient over the next shuffled batch. Test accuracy uses the 60 held‑out images, which never touch W.
loss per update step (train)
accuracy per epoch · train / test
confusion matrix on the 60 test images · rows = true, cols = predicted
one epoch = the 240 training images, shuffled, cut into batches (rows of 16) · colour = class · faded + red outline = currently misclassified · dark outline = the batch just used · right: the 60 test images, never used for updates
after training · inference Once the loss is small, W has learned which combinations of x push the "heart" score up, which push "leaf" up, and so on. To classify a new image you run only the forward half — flatten → W·x + b → softmax — and take the argmax: argmax([0.7, 0.2, 0.1]) = slot 0 = heart. No labels, no loss, no update. Click any dashed (test) thumbnail in E: the stages above show the forward pass and whether the prediction matches its label.

11The whole network, computed for real

in plain words

This is the real thing running on our heart: conv → ReLU → pool, repeated, then flatten into one long vector, multiply by a weight matrix, and softmax turns the scores into probabilities. Watch each stage light up in order.

16×16×3 → conv 3×3 (4 filters) → ReLU → maxpool 2 → conv 3×3 (4 filters) → ReLU → maxpool 2 → flatten → linear → softmax