Scroll. Every stage is computed live. Hover, drag the sliders, press play.
This is the whole journey of one picture: split into channels, filtered many times, shrunk, flattened, and turned into a probability for each class. This is also the single network every later chapter refers to — each one shows this same strip at the top with only its own layer lit, so you always know where you are. Play with the knobs below and watch the workflow, the neuron view, and every number recompute together, live.
every box is computed live on the 12×12 heart · click a box to jump to the chapter that explains it
the same network drawn as neurons — every column responds to the knobs above · scroll sideways if it is wide
Both cᵢ and cₒ are channel counts, not layer numbers — that little subscript is a letter, in or out, never a digit.
cᵢ “c-in” = how many channels go in — decided for you by the previous layer
cₒ “c-out” = how many channels come out — your choice, and always equal to the number of filters
k = kernel size (3 here) · hᵢ, hₒ = height going in, height coming out
So one filter is always k × k × cᵢ, and stacking cₒ of them gives the next layer its cᵢ. That hand-off is what the arrows across this map are tracing.
A picture is just a table of numbers. Each cell is one pixel; the number is its brightness, from 0 (black) to 255 (white).
image → matrix of pixels · shape (H, W) = (12, 12) · hover a pixel
A colour pixel is three numbers, not one: how much red, how much green, how much blue. So a colour image is three tables stacked, written as (C, H, W) = (3, 12, 12) — 3 channels, 12 rows, 12 columns.
tensor shape (C, H, W) = (3, 12, 12) · one 12×12 matrix for R, one for G, one for B
A kernel is a tiny table of weights — 3×3 here. It slides over the input; at each position it multiplies its 9 weights with the 9 numbers underneath and adds them up, and that single sum is one output number. Doing this at every position builds a new, slightly smaller table called a feature map.
A kernel window (or "slice") is a single k×k grid of weights that slides over one channel. In our case each window is 3×3 = 9 weights.
A kernel (also called a filter) is the full stack of windows needed to cover all input channels — one window per channel, glued together. With 3 input channels, one kernel = 3 windows = 3×3×3 = 27 weights. The kernel produces one output feature map.
input 3 × 12 × 12 (cᵢ × H × W) — the real heart from 01–02, real pixel values · kernel 3 × 3 × 3 (cᵢ × k × k), one 3×3 slice per input channel · output 1 × 10 × 10
We have 3 kernel boxes, one per input channel. The numbers inside them are the weights — numbers we choose (later: numbers training chooses), not the result of applying them to the picture.
Together that is 3 × 3 × 3 = 27 weights + 1 bias. What a weight really is, and how training picks its value, comes in 03·R and chapter 10 — for now just keep the idea: kernel boxes = weights; the coloured sums below = what happens when we apply them.
MAC = Multiply‑ACcumulate: one multiplication followed by adding the result onto a running total, a ← a + w·x. It is the basic unit of work in a neural network, and the lecture counts cost in MACs.
In the demo above, one output cell costs 27 MACs: 0×204, 0×204, … , 2×204 … each product is added onto the sum → 334. The whole 10×10 map is 100 cells × 27 = 2,700 MACs, and with cₒ kernels it is 2,700·cₒ.
MACsconv = ci · kh · kw · ho · wo · co
Our case: 3 · 3 · 3 · 10 · 10 · 1 = 2,700. One MAC = 2 FLOPs (a multiply and an add), so hardware speeds quoted in FLOPS or OPS (operations per second) tell you how many of these a chip can do each second — that is how latency is estimated: #MACs ÷ chip speed. Note the two counts are different things: #parameters = numbers stored (27 here), #MACs = work done (2,700 here).
hᵢ = height of the input (12 rows of the heart) · hₒ = height of the output feature map · k = kernel size (3) · the same letters with w are for width.
p = padding, a ring of extra cells added around the input — 0 here, explained in step 04 · s = stride, how far the window jumps each step — 1 here, explained in step 05.
Now, how do we know hₒ? We know k = 3, hᵢ = 12, p = 0, s = 1. There is one general equation that takes all of this into account:
ho = ⌊ (hi + 2p − k) / s ⌋ + 1
So in our case: hₒ = ⌊(12 + 2·0 − 3) / 1⌋ + 1 = 10. With no padding and stride 1 this is simply hᵢ − k + 1: the window fits 12 − 3 + 1 = 10 times across, so the map shrinks from 12 to 10.
cᵢ = "c‑in", the number of channels coming in to the layer — 3 here (R, G, B). cₒ = "c‑out", the number of channels coming out — equal to how many kernels you use (1 in this demo, so the output is 1 × 10 × 10). One kernel must cover all cᵢ channels, so it stores cᵢ·k·k = 3·3·3 = 27 weights (its own numbers) + 1 bias. At every window position each weight is multiplied by the pixel under it — 9 products per channel, 27 in total — and all 27 products plus the bias are added into one output cell. Same 27 weights, reused at all 100 positions of the 10×10 map: 27 weights, but 100 × 27 = 2,700 multiply‑adds.
One kernel makes one feature map. If you want to detect several things — vertical edges, horizontal edges, blur — you use several kernels, and each produces its own output channel. Below we use 4 kernels (red detector, vertical edge, horizontal edge, blur) on the 3 input channels, so 4 feature maps come out. That is why the weights have shape (cₒ, cᵢ, k, k) = (4, 3, 3, 3): cₒ = 4 kernels, each looking at all cᵢ = 3 input channels through a 3×3 window.
Both subscripts are letters, not numbers: cₒ = “c-out”, how many channels come out (= how many kernels you use); cᵢ = “c-in”, how many channels go in.
many kernels → many output channels · here cₒ = 4 kernels, cᵢ = 3 channels, k = 3 · weights shape (cₒ, cᵢ, k, k) = (4, 3, 3, 3)
The output of one conv layer is just another image with more channels. So we do it again: the next layer takes the (cₒ, 10, 10) maps as its input, applies new kernels, and gives even smaller maps with even more channels. Each round the picture gets spatially smaller (fewer rows/columns) and deeper (more channels), because we trade "where" for "what". Pooling and stride (06, 05) make the shrinking faster.
A kernel is a detector. "No channels" isn't quite right: one kernel → one feature map, and that map is one channel, so layer 2 would receive cᵢ = 1. Not "no channels" — only one. But the point stands: you've collapsed everything the layer could notice into a single number per position.
Why one kernel is a bad idea: the red‑detector kernel answers only "is this pixel red?". If that's your single output channel, layer 2 has nothing else to work with — no edges, no corners, no texture, no green/blue information. Everything but one feature is thrown away. To recognise a heart you need many detectors at once (edges of several orientations, colour blobs, curves), so you use many kernels → many channels, and each later layer combines them into more complex features.
layer L reads the (cₒ, h, w) of layer L−1 · every layer: hₒ = hᵢ − k + 1 · #weights = cₒ·cᵢ·k·k · computed live on the heart
A neuron takes some inputs, multiplies each by its own weight, adds them up, adds a bias, and squashes the sum through an activation function. That is all a neuron ever does. A network is many neurons in layers, each layer feeding the next.
In our case: the inputs are the 27 pixel values under the 3×3×3 window, the weights are the 27 kernel numbers, and one output cell of the feature map is exactly one neuron. Parameters = weights + biases = the numbers training will change.
one neuron = the top-left cell of the feature map · net = Σ wᵢ·xᵢ + b · o = φ(net) · the neuron follows the sliding window of the demo above
Every output cell is one such neuron. Its net (before φ) goes into the left map; its activation o = φ(net) goes into the right map — that right map is what the next layer receives. Details in Activation: a non-linear squeeze after every linear layer.
Stack neurons into layers. Input layer = the raw numbers (our 3·12·12 = 432 pixel values). Hidden layers = neurons whose outputs nobody sees directly, they only feed the next layer. Output layer = one neuron per class (heart, leaf, ball). When every neuron connects to every neuron before it, the layer is fully connected. A conv layer is the same picture, except each neuron only connects to a small window and all neurons of one map share the same 27 weights — that is why it needs far fewer parameters.
The network, drawn as neurons — reusing the same live diagram from 00 · the whole workflow at a glance, so counts and connections always match the knobs there.
Kernels set the depth (how many feature maps); the input size, kernel size, padding and stride set each map's height × width. The neuron count is just their product:
#neurons = co · ho · wo, where ho = ⌊ (hi + 2p − k) / s ⌋ + 1
Our heart, layer 1: cₒ = 4 kernels, hᵢ = 12, k = 3, p = 0, s = 1 → hₒ = 12 − 3 + 1 = 10 → #neurons = 4 · 10 · 10 = 400.
Another example — a 10×10×3 image with 2 kernels: hₒ = 10 − 3 + 1 = 8 → #neurons = 2 · 8 · 8 = 128 (weights = 2·27 = 54, biases = 2). The 2 kernels give 2 maps; each map is 8×8 no matter how many kernels there are.
Every convolution eats k−1 pixels from the border, so the map shrinks. Padding adds a ring of extra cells — usually zeros — around the input, so the output keeps its size and edge pixels get looked at as often as centre ones.
Padding — keeps the map size and stops edge pixels being eaten. Preserves data, not RF.
What does change the receptive field is covered in 07.
Stride is how far the window jumps after each step. Stride 1 visits every position; stride 2 skips every other one and produces a map about half the size — cheaper, and it lets later layers see a bigger region.
Pooling shrinks a feature map with no weights to learn: look at a 2×2 window and keep only the biggest value (max) or the average. It keeps the strong signals and throws away exact position.
2×2 window, stride 2 · no weights to learn · each channel pooled on its own · hover the output
Pooling's max — keeping the biggest value preserves the strongest feature; that is separate from the downsampling that helps RF.
What does change the receptive field is covered in 07.
Here is our network again, but with a pool layer inserted after the conv feature maps. Pooling has no neurons of its own to learn — it just shrinks each of the 4 maps from 10×10 to 5×5, so the 4 feature maps become 4×5×5 = 100 neurons before the dense layer. One node = one neuron, dashed pointer marks it.
The workflow again, greyed out except pooling — same input → … → softmax pipeline as chapter 00, here dense = fully‑connected layer (has weights); it is not the activation.
The receptive field is how many input pixels one output number depends on. Each 3×3 conv layer (a full conv layer whose kernels use a 3×3 window) adds 2 more pixels, so stacking layers — or using stride — lets a single neuron eventually see the whole heart.
A neuron can only use what is inside its receptive field — everything outside it is invisible to that neuron. So the receptive field is literally how much of the image this neuron is allowed to know about.
To recognise a whole heart, some neuron must see the whole heart at once. If a neuron's field is only 5×5 pixels it can spot a tiny edge but never "heart", because a heart does not fit in its view. Recognising a shape needs a receptive field at least as big as the shape.
It decides how deep / how much pooling you need. One 3×3 conv layer sees only 3×3; covering a 224×224 image with 3×3 conv layers alone would take ~111 layers. Each stride‑2 (or pool) step roughly doubles how fast the field grows, so you reach "sees everything" in a handful of layers instead of a hundred — that trade‑off is the whole reason downsampling exists.
It explains the early‑to‑deep story. Early layers (small field) can only see edges, corners, colour; deep layers (large field) can see objects and layout. The growing receptive field is exactly why a CNN builds from simple features to complex ones — and if your object is bigger than the final field, accuracy silently caps out no matter how long you train.
Not a random choice. More layers = more understanding of the picture, but there is a minimum to start from: one output must be able to see the whole object. That minimum is exactly what the receptive field tells us. Setting RF ≥ input size and solving RF = L·(k−1)+1 for L:
L ≥ (input − 1) / (k − 1)
A 150×150 image with 3×3 kernels: L ≥ (150−1)/(3−1) = 74.5 → 75 layers. Bigger kernels help — with 10×10: 149/9 ≈ 16.6 → 17 layers — but a big kernel costs about (k/3)² more per layer, so stacking small kernels is usually cheaper.
Small‑and‑deep beats big‑and‑shallow. The receptive field after L layers is RF = L·(k−1)+1, so two 3×3 layers reach the same RF as one 5×5 layer:
two 3×3: RF = 2·(3−1)+1 = 5 = one 5×5: RF = 1·(5−1)+1 = 5
Same reach, but the weights differ: two 3×3 layers use 2·(3·3) = 18 weights, one 5×5 uses 5·5 = 25. The stacked pair is cheaper and adds an extra activation (more non‑linearity), which is why modern CNNs stack 3×3 convs instead of using large kernels.
The real fix is downsampling. With a stride‑2 or 2×2 pool every stage, each stage halves the map and roughly doubles the reach, so you only need about log₂(input) stages: 150 → 75 → 38 → 19 → 10 → 5 → 3 → 1, i.e. ~7 stages, not 75 layers. That is why real CNNs are ~10–50 layers, not hundreds — and why the lecture says: for large images, downsample inside the network. Extra layers beyond this minimum are added for richer features, then capped by compute and memory.
Only three things change the receptive field: kernel size k, number of layers L, and stride/pooling (downsampling). Everything else is useful for other reasons.
Kernel k — bigger k adds k−1 to RF each layer, but ~(k/3)² more memory.
Layers L — RF = L·(k−1)+1, and deeper = richer features; costs more compute. (More kernels ≠ more layers: extra kernels add feature maps in one layer, they do not grow RF.)
Stride / pooling — downsample: RF of later layers grows twice as fast. Too aggressive → detail lost.
The rule of thumb is small kernel + moderate depth + downsampling — not "medium everything". Three separate knobs:
1 · Kernel size k → small (3×3). Bigger k grows RF faster per layer (RF adds k−1 each layer), but costs ~(k/3)² more weights and blurs detail. Good: 3×3 stacked reaches any RF cheaply. Bad: one 7×7 reaches RF 7 in a layer but for 49 weights vs 3 stacked 3×3 layers reaching RF 7 for 27 — wasteful.
2 · Depth L → moderate, and it's a result not a target. RF = L·(k−1)+1, so more layers = more RF and richer features. Good: enough layers that RF ≥ object size. Bad: too few → RF never covers the heart, it literally cannot be recognised; too many → huge cost and, past ~20 plain layers, training degrades.
3 · Downsampling (stride‑2 / pool) → yes. This is the knob people forget. Each halving stage makes RF grow geometrically, not by a flat k−1. Good: 150×150 needs ~7 stages instead of 75 layers. Bad: overdo it and the map becomes 1×1 too early, throwing away spatial detail before features are built.
On our 12×12 heart: plain 3×3, no downsampling → RF ≥ 12 needs L ≥ 6 layers. Add one 2×2 pool after 2 layers → RF passes 12 in ~2 conv layers (what chapter 11 actually does). Same k, far fewer layers — because downsampling did the heavy lifting.
Getting a receptive field that covers the whole heart is a design decision. Here is the base design with the minimum depth needed. With plain 3×3 convs (no downsampling) RF = L·(k−1)+1 = 2L+1, and we need it ≥ the picture size (12), because the heart is 12×12:
2L + 1 ≥ 12 → L ≥ 5.5 → L = 6 conv layers
Six layers on a 12×12 image is wasteful, so we add padding + pooling and reach RF ≥ 12 in far fewer layers. This base design (3×3 convs, padding 1, one 2×2 pool every couple of layers) is the minimum that gives a meaningful RF:
| stage | op | map size | RF on input |
|---|---|---|---|
| input | — | 12×12 | 1 |
| conv1 (4 kernels, 3×3, p=1) | conv | 12×12×4 | 3 |
| conv2 (3×3, p=1) | conv | 12×12 | 5 |
| pool 2×2 s=2 | pool | 6×6 | 6 |
| conv3 (3×3, p=1) | conv | 6×6 | 10 |
| pool 2×2 s=2 | pool | 3×3 | 12 ✅ |
After every linear operation we squash the result with a non-linear function. Without it, stacking layers would still add up to one big linear equation and could never learn shapes. ReLU is the simplest: negatives become 0, positives stay as they are.
y = f(Σ wᵢxᵢ + b) · the map is the heart's vertical-edge output, every 2nd cell (÷50) · click a function · red = negative, teal = positive
Conv and fully‑connected layers are linear (output = Σ weight×input + b). Stacking linear layers collapses into one linear layer — 50 layers with no activation = the power of 1, only straight‑line boundaries, never a curved shape like a heart. A non‑linear activation (ReLU bends everything below 0 to 0) is the hinge between layers that stops the collapse, so each new layer actually adds power. Without it, depth would be pointless.
Activation — adds non‑linearity. Zero effect on RF.
What does change the receptive field is covered in 07.
After the last conv/pool, the network holds a stack of small feature maps — a 3‑D block (C × H × W). But the answer we want is just 3 numbers (heart / leaf / ball). Flatten unrolls that 3‑D block into one long 1‑D vector, and the dense (fully‑connected) layer turns that vector into the final class scores — every output neuron connects to every input number.
Channels grow as the map shrinks (trade "where" for "what"). Doubling the kernels each conv, the last conv outputs 16 feature maps, each 3×3:
| stage | map size | #feature maps (channels) |
|---|---|---|
| input | 12×12 | 3 (R,G,B) |
| conv1 (4 kernels) | 12×12 | 4 |
| conv2 (8 kernels) | 12×12 | 8 |
| pool | 6×6 | 8 |
| conv3 (16 kernels) | 6×6 | 16 |
| pool | 3×3 | 16 |
So just before flattening we have 16 feature maps, each 3×3 — a 3‑D block of shape (C×H×W) = 16×3×3.
the 3‑D block = C matrices stacked · here C=16, each 3×3 · scroll sideways
A dense layer needs a 1‑D list, not a 3‑D block. One representation is to flatten: unroll the block into a single long vector. That is what we chose here — 16 × 3 × 3 = 144 numbers.
Flatten reads the block in a fixed order (channel, then row, then column) and lays every number end‑to‑end into one column. Nothing is computed, nothing is learned — same numbers, new shape. Our 16×3×3 block becomes a vector of 16·3·3 = 144 numbers. Hover a cell to see where it lands.
A dense layer has one neuron per class. Each neuron sees all the flattened numbers, multiplies each one by its own weight (a synapse), adds them up, and adds a bias. Dense doesn't pick a class. It produces one new number per class.
Synapse = one weight on one connection. Neuron = the sum of all its weighted inputs plus a bias. (The f in the drawing is skipped in this layer: the output is just the sum plus bias.)
z = W·x + b. Every class gets the same x but its own row of weights and its own bias. The drawing uses 4 inputs. The real vector has 144, so W is 3 × 144 = 432 weights, plus 3 biases.
z is the raw score of each class: one number per class (here heart 9.0, leaf 2.0, ball 1.0). It says how strongly the input matches what that class's weights look for, and bigger means more evidence. It is not a probability: it can be negative or bigger than 1, and the three numbers don't add up to anything. Only which one is largest, and by how much, matters. Which output is 'heart' is just our labelling of position 0. Training is what makes row 0 respond to hearts.
Next: turning the scores into a decision.
Because it connects everything to everything, its weight count is cₒ·cᵢ — and cᵢ is the whole flattened block, which can be huge. In AlexNet the first dense layer alone holds ~38 million weights (9216 × 4096), far more than all the conv layers combined. That is why modern networks pool aggressively before flattening (smaller block → smaller vector → cheaper dense) or replace flatten+dense with global average pooling (average each map to one number, so C×H×W → C directly).
The gradient goes to the nearest layer first and moves backward one layer at a time, but no weights change during that trip. All the purple layers are updated together at the end.
Order.
This is the real thing running on our heart: conv → ReLU → pool, repeated, then flatten into one long vector, multiply by a weight matrix, and softmax turns the scores into probabilities. Watch each stage light up in order.
16×16×3 → conv 3×3 (4 filters) → ReLU → maxpool 2 → conv 3×3 (4 filters) → ReLU → maxpool 2 → flatten → linear → softmax