Teach a computer to listen to music and guess its genre — using a Convolutional Neural Network trained on the GTZAN dataset.
A beginner's visual guide →
A program that records 15–30 seconds of a song and prints something like “HIPHOP — 84% confident.” Here's the whole journey of one sound clip:
GTZAN: 1,000 songs across 10 genres. The “textbook” dataset for this task.
A CNN — the same family of network that recognises faces and cats in photos.
We classify pictures of sound, so an image network works on audio.
It's a math function with adjustable knobs, loosely inspired by the brain. It takes numbers in, passes them through layers of simple units called neurons, and produces an answer out.
Numbers flow left → right. Every line is a weight the network can tune.
output = activation( x₁·w₁ + x₂·w₂ + x₃·w₃ + b )
A network is just millions of these wired together. Training finds the right w's and b's.
Learning is repetition with feedback — like practising an instrument. We loop over the examples again and again (each full pass is an epoch):
Do this thousands of times and the loss shrinks — the network gets good.
Each step moves the “ball” toward the lowest point — the best weights.
The big project uses PyTorch, which hides the maths. So we also built a tiny network in pure NumPy — about 30 lines — that learns XOR: output 1 when the two inputs differ, else 0. No single straight line can separate it, so you need a hidden layer. src/simple_nn/
# forward pass: inputs -> hidden -> output z1 = X @ W1 + b1 # weighted sums a1 = np.tanh(z1) # the "bend" z2 = a1 @ W2 + b2 pred = sigmoid(z2) # probability in (0,1) for epoch in range(5000): # the learning loop pred = forward(X) # 1 guess loss = cross_entropy(pred, y) # 2 score grads = backward(pred, y) # 3 blame W1 -= lr * grads.W1 # 4 adjust W2 -= lr * grads.W2 # (gradient descent)
| a | b | answer |
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
It's the same Guess → Score → Blame → Adjust loop as the CNN — just small enough to read every line.
# train it and watch loss fall to ~0 uv run simple-nn train # then predict XOR(1, 0) -> 1 uv run simple-nn predict 1 0
After the network guesses, we know how wrong it was (the loss). But which weights caused the error? Backprop answers that by working backwards from the output to the input, one layer at a time.
In network.py the whole thing is the backward() method — a few lines of multiplying matrices.
Without a non-linear bend, stacking layers would just be one big straight line. Activations let the network curve its decision boundary into any shape.
sigmoid
0 → 1, great for a probability
tanh
−1 → 1, centred on zero
ReLU
off below 0, linear above
Our tiny net uses tanh in the hidden layer and sigmoid at the output. The CNN uses ReLU — fast and the default for deep nets.
A Convolutional Neural Network is a neural network specialised for grids of pixels (images). Instead of connecting every pixel to every neuron (wasteful!), it slides small filters across the image to detect local patterns.
The image gets smaller while the stacks get deeper — the network trades “where” for “what”, then votes on the genre.
A tiny window (e.g. 3×3) slides over the image. At each spot it checks “does my pattern appear here?” producing a feature map.
ReLU keeps the strong responses and discards the rest — adding the non-linearity that lets layers stack meaningfully.
Shrink the map by keeping the strongest value in each little region. Smaller data, keeps what matters, resists tiny shifts.
Stack these blocks and the network builds a hierarchy: edges → textures → shapes → objects (for music: beats → rhythms → timbres → genres).
Input “image” — the pink box is the filter scanning it.
Many filters run in parallel — one might find vertical edges, another a drum hit.
Yes — it is a neural network. A conv filter does the same math as a neuron (multiply by weights, add a bias, apply ReLU), but its weights are a tiny 3×3 grid reused everywhere in the image. “conv · 16” = 16 different filters, each producing its own feature map.
The outlined cell = one number: the energy at that pitch & moment.
Each block keeps the strongest signals (pooling halves the width & height) until the grid is tiny but “deep”. Flatten just lines those 64×4×4 = 1024 numbers into a single list — ready for the ordinary fully-connected layers to vote on the genre.
A CNN wants an image — but a song is a 1-D wiggle (a waveform). The bridge is a mel-spectrogram:
Now genre becomes visible: metal is a dense bright wall, classical is sparse and flowing, hip-hop shows a strong steady low-end beat.
3-sec clip → 128 × 130 pixel image → fed to the CNN.
the raw audio wiggle ↑
These are actual mel-spectrograms of two GTZAN clips. Same kind of picture, totally different fingerprints — exactly what the CNN learns to tell apart.
Clean, separated notes with breathing room — you can almost see the melody lines.
Energy smeared across every frequency at once — loud, distorted, relentless.
The classic music-genre benchmark, free on Kaggle. We learn from real examples — the model is only as good as the music it hears.
Even people disagree! Rock vs. metal, disco vs. pop — the boundaries are blurry, so ~75% accuracy is genuinely good.
Four convolution blocks extract patterns, then two fully-connected layers vote on the genre. Notice the image gets smaller while the number of channels grows — the network trades “where” for “what”.
# model.py — one convolution block def conv_block(in_ch, out_ch): return Sequential( Conv2d(in_ch, out_ch, kernel_size=3, padding=1), BatchNorm2d(out_ch), # keep numbers stable ReLU(), # add non-linearity MaxPool2d(2), # shrink the image )
Training is what turns a network of random weights into one that actually recognises genres. The whole thing is four stages — and one command runs them all.
For every batch of spectrograms, repeat 4 steps:
↻ Do this for every batch, then repeat the whole pass 30 times (30 epochs). Each pass, the guesses get a little better.
So model.pt always holds the best version seen, not just the last.
one command does all of this → uv run genre train
The genre with the highest average wins, and we show the confidence.
# $ uv run genre record Genre: HIPHOP (84% confident) Analysed 10 segment(s): hiphop 84.1% █████████████████ pop 7.3% ██ reggae 4.0% █ rock 2.1% ▌
Confidence matters: 84% is a strong call; 30% means “honestly, not sure”.
# 1. install everything uv sync # 2. (download GTZAN into data/ first) # 3. train the CNN uv run genre train # 4. classify a file... uv run genre predict song.wav # 5. ...or record live! uv run genre record --seconds 30
# save the "picture of sound" as a PNG
uv run genre demo song.wav
This is the whole model in ~30 lines of PyTorch. It has two halves: a feature extractor (conv blocks that find patterns) and a classifier (dense layers that vote on the genre).
class GenreCNN(nn.Module): # our model = a CNN def __init__(self): # 1) FEATURES — 4 conv blocks find patterns self.features = nn.Sequential( conv_block(1, 16), # 1 channel -> 16 maps conv_block(16, 32), # deeper & richer... conv_block(32, 64), conv_block(64, 64), ) # 2) squash ANY size to a fixed 4x4 grid self.pool = nn.AdaptiveAvgPool2d((4, 4)) # 3) CLASSIFIER — fully-connected layers vote self.classifier = nn.Sequential( nn.Flatten(), # 64x4x4 -> 1024 nn.Linear(1024, 128), nn.ReLU(), nn.Dropout(0.3), # anti-overfit nn.Linear(128, 10), # 10 genre scores ) def forward(self, x): # x: (N, 1, 128, 130) x = self.features(x) # -> (N, 64, 8, 8) x = self.pool(x) # -> (N, 64, 4, 4) return self.classifier(x) # -> (N, 10)
# one repeatable conv block def conv_block(in_ch, out_ch): return nn.Sequential( nn.Conv2d(in_ch, out_ch, 3, padding=1), nn.BatchNorm2d(out_ch), nn.ReLU(), nn.MaxPool2d(2), )
Here's the model's very first block — self._conv_block(in_ch=1, out_ch=16). The single grayscale spectrogram channel fans out into 16 as 16 different 3×3 filters each make their own feature map:
Only Conv2d grows the channels (1 → 16, the fan-out). BatchNorm and ReLU keep 16, and MaxPool keeps 16 but halves the width & height.
# the model's first block — turns 1 channel into 16 feature maps self._conv_block(in_ch=1, out_ch=16) def _conv_block(in_ch, out_ch): return nn.Sequential( nn.Conv2d(in_ch, out_ch, kernel_size=3, padding=1), # 1 → 16 nn.BatchNorm2d(out_ch), # keep numbers stable nn.ReLU(), # add non-linearity nn.MaxPool2d(2), # shrink the image )
The whole loop is just a handful of lines. The four inner steps are exactly the Guess → Score → Blame → Adjust cycle from before.
# set up once model = GenreCNN().to(device) criterion = nn.CrossEntropyLoss() # the loss optimizer = torch.optim.Adam(model.parameters(), lr=1e-3) for epoch in range(EPOCHS): # repeat 30x for inputs, targets in train_loader: # one batch optimizer.zero_grad() # reset outputs = model(inputs) # 1 guess loss = criterion(outputs, targets) # 2 score loss.backward() # 3 blame optimizer.step() # 4 adjust
After each epoch we check accuracy on the held-out set and save the best model. That's the entire idea of training.
Sound becomes an image, a CNN learns the patterns, and the model names the genre. Now go train it, break it, and improve it.
README.md → setup & commands src/ → read the code
Happy learning! Press ← to review any slide.
A unit that weighs inputs, sums them, and fires through an activation.
The tunable knobs. Training adjusts them to reduce errors.
The squiggle that lets networks learn curves, not just lines.
A number measuring how wrong the model is. Lower is better.
Rolling downhill on the loss to find good weights.
One full pass over all the training data.
Sliding a small learned filter to detect local patterns.
Shrinking a feature map while keeping the strongest signals.
Memorising the training data instead of truly learning.
A picture of sound: time × pitch × energy.
Turns raw scores into probabilities that sum to 100%.
Held-back data for an honest accuracy check.