Use ← → arrows · or click the dots
🎵🧠

Song Genre Classifier

Teach a computer to listen to music and guess its genre — using a Convolutional Neural Network trained on the GTZAN dataset.

Neural networks CNNs Audio → images Python · PyTorch

A beginner's visual guide →

The big picture

What are we building?

A program that records 15–30 seconds of a song and prints something like “HIPHOP — 84% confident.” Here's the whole journey of one sound clip:

🎤
Record
grab audio from the mic
→
📊
Spectrogram
turn sound into a picture
→
🧠
CNN
find patterns in the picture
→
🏷️
Genre
jazz? metal? pop?

The data

GTZAN: 1,000 songs across 10 genres. The “textbook” dataset for this task.

The brain

A CNN — the same family of network that recognises faces and cats in photos.

The trick

We classify pictures of sound, so an image network works on audio.

Foundations

What is a neural network?

It's a math function with adjustable knobs, loosely inspired by the brain. It takes numbers in, passes them through layers of simple units called neurons, and produces an answer out.

  • Each neuron multiplies its inputs by weights, adds them up, and applies a squiggle (an activation) so it can learn curves, not just straight lines.
  • Layers stack neurons. Early layers learn simple things; later layers combine them into complex ideas.
  • “Learning” = automatically tuning the weights until the answers are good.
inputs hidden output

Numbers flow left → right. Every line is a weight the network can tune.

Foundations

Inside one neuron

x₁x₂x₃ ×w₁×w₂×w₃ Σ+b activation output

The recipe

output = activation( x₁·w₁ + x₂·w₂ + x₃·w₃ + b )

  • Weights (w): how much each input matters.
  • Bias (b): a nudge that shifts the result up or down.
  • Activation: e.g. ReLU (“keep positives, zero the rest”). Without it, stacking layers would still only draw straight lines.

A network is just millions of these wired together. Training finds the right w's and b's.

Foundations

How does it learn?

Learning is repetition with feedback — like practising an instrument. We loop over the examples again and again (each full pass is an epoch):

  • 1 · Guess (forward pass): feed in a spectrogram, get a genre prediction.
  • 2 · Score (loss): measure how wrong the guess was vs. the true label. Big mistake → big loss.
  • 3 · Blame (backpropagation): compute how each weight contributed to the error.
  • 4 · Adjust (gradient descent): nudge every weight a little in the direction that lowers the loss.

Do this thousands of times and the loss shrinks — the network gets good.

Gradient descent = rolling downhill

start (high loss) minimum (low loss)

Each step moves the “ball” toward the lowest point — the best weights.

See it in code

A neural network from scratch (no PyTorch)

The big project uses PyTorch, which hides the maths. So we also built a tiny network in pure NumPy — about 30 lines — that learns XOR: output 1 when the two inputs differ, else 0. No single straight line can separate it, so you need a hidden layer. src/simple_nn/

# forward pass: inputs -> hidden -> output
z1 = X @ W1 + b1        # weighted sums
a1 = np.tanh(z1)        # the "bend"
z2 = a1 @ W2 + b2
pred = sigmoid(z2)      # probability in (0,1)

for epoch in range(5000):    # the learning loop
    pred = forward(X)              # 1 guess
    loss = cross_entropy(pred, y)  # 2 score
    grads = backward(pred, y)      # 3 blame
    W1 -= lr * grads.W1            # 4 adjust
    W2 -= lr * grads.W2            #   (gradient descent)

The XOR truth table

abanswer
000
011
101
110

It's the same Guess → Score → Blame → Adjust loop as the CNN — just small enough to read every line.

# train it and watch loss fall to ~0
uv run simple-nn train
# then predict XOR(1, 0) -> 1
uv run simple-nn predict 1 0
The two big ideas

Backpropagation & activation functions

🔁 Backpropagation — assigning blame

After the network guesses, we know how wrong it was (the loss). But which weights caused the error? Backprop answers that by working backwards from the output to the input, one layer at a time.

  • Chain rule: it multiplies the slopes of each step together to get, for every weight, “if I nudge this, how much does the loss change?” — that number is the gradient.
  • Gradient descent: we then step each weight a little in the direction that lowers the loss: W -= lr * gradient.
  • learning rate (lr): the step size. Too big overshoots; too small crawls.

In network.py the whole thing is the backward() method — a few lines of multiplying matrices.

⚡ Activation functions — the “bend”

Without a non-linear bend, stacking layers would just be one big straight line. Activations let the network curve its decision boundary into any shape.

sigmoid
0 → 1, great for a probability

tanh
−1 → 1, centred on zero

ReLU
off below 0, linear above

Our tiny net uses tanh in the hidden layer and sigmoid at the output. The CNN uses ReLU — fast and the default for deep nets.

The star of the show

What is a CNN?

A Convolutional Neural Network is a neural network specialised for grids of pixels (images). Instead of connecting every pixel to every neuron (wasteful!), it slides small filters across the image to detect local patterns.

The shape of a CNN — picture → patterns → genre

spectrogram 1×128×130 conv · 16 conv · 32 conv · 64 flatten dense · 128 10 genres jazz ✓

The image gets smaller while the stacks get deeper — the network trades “where” for “what”, then votes on the genre.

🔍 Convolution

A tiny window (e.g. 3×3) slides over the image. At each spot it checks “does my pattern appear here?” producing a feature map.

⚡ Activation

ReLU keeps the strong responses and discards the rest — adding the non-linearity that lets layers stack meaningfully.

🔽 Pooling

Shrink the map by keeping the strongest value in each little region. Smaller data, keeps what matters, resists tiny shifts.

Stack these blocks and the network builds a hierarchy: edges → textures → shapes → objects (for music: beats → rhythms → timbres → genres).

CNN mechanics

A filter sliding over an image

Input “image” — the pink box is the filter scanning it.

→

What's happening

  • The 3×3 filter overlaps a patch of the image.
  • It multiplies + sums to get one number: how strongly its pattern is present there.
  • It slides to the next spot and repeats, building a feature map.
  • The filter's 9 numbers are learned during training — the network decides what's worth detecting.

Many filters run in parallel — one might find vertical edges, another a drum hit.

Zoom in

Inside a conv layer — what is “conv · 16”?

Yes — it is a neural network. A conv filter does the same math as a neuron (multiply by weights, add a bias, apply ReLU), but its weights are a tiny 3×3 grid reused everywhere in the image. “conv · 16” = 16 different filters, each producing its own feature map.

The input: a spectrogram tensor 1×128×130

128 mel bands · pitch ↑
130 time frames · ~3 s →

The outlined cell = one number: the energy at that pitch & moment.

  • 1 = channels — grayscale, one value per cell (a colour photo has 3: R, G, B).
  • 128 = height — mel frequency bands, low pitch (bottom) → high (top).
  • 130 = width — time steps, ~3 seconds, left → right.
  • Altogether 16,640 numbers describing the sound.

conv · 16 → 16 feature maps

1×128×130 3×3 filter ×16 16×128×130
  • Each filter = 3×3 = 9 weights + 1 bias, shared across the whole image.
  • 16 filters → just 160 weights. A dense layer here would need ~2 million!
  • Each map is one filtered “view” — one may light up on beats, another on cymbals.

How the shape shrinks, block by block → then flatten

Input
1×128×130
→
conv·16
16×64×65
→
conv·32
32×32×32
→
conv·64
64×16×16
→
conv·64 + pool
64×4×4
→
flatten
1024 numbers

Each block keeps the strongest signals (pooling halves the width & height) until the grid is tiny but “deep”. Flatten just lines those 64×4×4 = 1024 numbers into a single list — ready for the ordinary fully-connected layers to vote on the genre.

The clever bit

Turning sound into a picture

A CNN wants an image — but a song is a 1-D wiggle (a waveform). The bridge is a mel-spectrogram:

  • X-axis = time (left → right as the song plays).
  • Y-axis = pitch / frequency, on the mel scale that matches how humans hear.
  • Brightness = energy at that pitch and moment.

Now genre becomes visible: metal is a dense bright wall, classical is sparse and flowing, hip-hop shows a strong steady low-end beat.

3-sec clip → 128 × 130 pixel image → fed to the CNN.

Waveform → Spectrogram

the raw audio wiggle ↑

low energy high

Real songs, real pictures — press ▶ and compare

These are actual mel-spectrograms of two GTZAN clips. Same kind of picture, totally different fingerprints — exactly what the CNN learns to tell apart.

🎻 Classical — sparse & flowing

Mel-spectrogram of a classical music clip

Clean, separated notes with breathing room — you can almost see the melody lines.

🤘 Metal — a dense bright wall

Mel-spectrogram of a metal music clip

Energy smeared across every frequency at once — loud, distorted, relentless.

The dataset

Meet GTZAN

The classic music-genre benchmark, free on Kaggle. We learn from real examples — the model is only as good as the music it hears.

  • 1,000 audio clips, each 30 seconds long.
  • 10 genres · 100 clips per genre (nicely balanced).
  • We split each clip into 3-sec chunks → ~10,000 training images. More examples = better learning.
  • 80% to train, 20% held back to honestly test the model on music it never saw.

The 10 genres

🎸 blues🎻 classical 🤠 country🕺 disco 🎤 hiphop🎷 jazz 🤘 metal✨ pop 🌴 reggae🎸 rock

Why genre is hard

Even people disagree! Rock vs. metal, disco vs. pop — the boundaries are blurry, so ~75% accuracy is genuinely good.

Our model

The CNN we built

Four convolution blocks extract patterns, then two fully-connected layers vote on the genre. Notice the image gets smaller while the number of channels grows — the network trades “where” for “what”.

Input
1 × 128 × 130
spectrogram
→
Conv ×4
16→32→64→64
filters + pooling
→
Flatten
grid → list
of numbers
→
Dense
1024 → 128
+ dropout
→
Output
10 scores
→ probabilities
# model.py — one convolution block
def conv_block(in_ch, out_ch):
    return Sequential(
        Conv2d(in_ch, out_ch, kernel_size=3, padding=1),
        BatchNorm2d(out_ch),   # keep numbers stable
        ReLU(),                # add non-linearity
        MaxPool2d(2),          # shrink the image
    )

Two anti-cheating tricks

  • Dropout: randomly ignore some neurons each step so the net can't over-rely on any one — fights overfitting (memorising instead of learning).
  • Validation set: the 20% held-back music gives an honest score. If train accuracy ≫ validation accuracy, it's memorising.
Putting it together

How training works — start to finish

Training is what turns a network of random weights into one that actually recognises genres. The whole thing is four stages — and one command runs them all.

📊
1 · Prepare
GTZAN audio → spectrograms (cached once)
→
✂️
2 · Split
80% train · 20% validation
→
🔁
3 · Train loop
repeat for 30 epochs
→
💾
4 · Save best
keep the most accurate model

Inside one epoch — the core loop

For every batch of spectrograms, repeat 4 steps:

  • 1 · Guess (forward pass) — predict the genres.
  • 2 · Score (loss) — measure how wrong the guesses are.
  • 3 · Blame (backprop) — find each weight's contribution to the error.
  • 4 · Adjust (optimizer step) — nudge every weight to lower the loss.

↻ Do this for every batch, then repeat the whole pass 30 times (30 epochs). Each pass, the guesses get a little better.

How we know it's really learning

  • After each epoch we test on the held-out 20% — music the model never trained on.
  • If that validation accuracy beats the previous best, we save the model to model.pt.
  • If train accuracy ≫ validation accuracy, it's memorising (overfitting) — which is exactly why we use Dropout.

So model.pt always holds the best version seen, not just the last.

one command does all of this → uv run genre train

Making a guess

From recording to answer

  • Record 15–30s and split it into 3-second segments.
  • Turn each segment into a spectrogram and run it through the CNN.
  • The CNN outputs 10 scores; softmax turns them into probabilities that add up to 100%.
  • Average the probabilities across all segments — one quiet intro won't fool the whole prediction.

The genre with the highest average wins, and we show the confidence.

Example output

# $ uv run genre record
Genre: HIPHOP  (84% confident)
Analysed 10 segment(s):
  hiphop   84.1% █████████████████
  pop       7.3% ██
  reggae    4.0% █
  rock      2.1% ▌

Confidence matters: 84% is a strong call; 30% means “honestly, not sure”.

Your turn

Run it yourself

Five commands

# 1. install everything
uv sync

# 2. (download GTZAN into data/ first)
# 3. train the CNN
uv run genre train

# 4. classify a file...
uv run genre predict song.wav

# 5. ...or record live!
uv run genre record --seconds 30

See the magic

# save the "picture of sound" as a PNG
uv run genre demo song.wav

Then experiment

  • Edit config.py: more epochs, longer segments, fewer mel bands.
  • Add a confusion matrix — which genres get mixed up?
  • Try data augmentation (pitch-shift, time-stretch).
  • Build a little web UI that records in the browser.
The code

model.py — the CNN, explained

This is the whole model in ~30 lines of PyTorch. It has two halves: a feature extractor (conv blocks that find patterns) and a classifier (dense layers that vote on the genre).

class GenreCNN(nn.Module):          # our model = a CNN
    def __init__(self):
        # 1) FEATURES — 4 conv blocks find patterns
        self.features = nn.Sequential(
            conv_block(1,  16),     # 1 channel  -> 16 maps
            conv_block(16, 32),     # deeper & richer...
            conv_block(32, 64),
            conv_block(64, 64),
        )
        # 2) squash ANY size to a fixed 4x4 grid
        self.pool = nn.AdaptiveAvgPool2d((4, 4))
        # 3) CLASSIFIER — fully-connected layers vote
        self.classifier = nn.Sequential(
            nn.Flatten(),              # 64x4x4 -> 1024
            nn.Linear(1024, 128),
            nn.ReLU(),
            nn.Dropout(0.3),           # anti-overfit
            nn.Linear(128, 10),       # 10 genre scores
        )

    def forward(self, x):           # x: (N, 1, 128, 130)
        x = self.features(x)          # -> (N, 64, 8, 8)
        x = self.pool(x)              # -> (N, 64, 4, 4)
        return self.classifier(x)     # -> (N, 10)
# one repeatable conv block
def conv_block(in_ch, out_ch):
    return nn.Sequential(
        nn.Conv2d(in_ch, out_ch, 3, padding=1),
        nn.BatchNorm2d(out_ch),
        nn.ReLU(),
        nn.MaxPool2d(2),
    )

Inside each conv block

  • Conv2d(…,3,padding=1) — slides 3×3 filters to find patterns; padding keeps the size.
  • BatchNorm2d — steadies the numbers so training is faster & stable.
  • ReLU — zeroes negatives → the non-linearity that lets layers stack.
  • MaxPool2d(2) — keeps the strongest value per 2×2 → halves width & height.

The classifier head — why each piece?

  • AdaptiveAvgPool2d((4,4)) — averages to a fixed 4×4 whatever the input size, so the next layer always gets exactly 64×4×4 = 1024 numbers.
  • Flatten — lines the 1024 numbers into a single list (dense layers need a flat input).
  • Linear — fully-connected “voting” layers that mix features into 10 genre scores.
  • Dropout(0.3) — randomly ignores 30% of neurons while training → fights overfitting.

Why it's built this way

  • Channels grow (16→64) while the grid shrinks — the net trades “where” for “what”.
  • nn.Sequential just runs the listed layers in order — clean, readable stacking.
  • forward() defines the data flow; PyTorch calls it on every prediction and auto-computes gradients for training.
  • The 10 scores become probabilities later via softmax (in predict.py).
The code, visualised

A conv block as a (simplified) network

Here's the model's very first block — self._conv_block(in_ch=1, out_ch=16). The single grayscale spectrogram channel fans out into 16 as 16 different 3×3 filters each make their own feature map:

self._conv_block(in_ch=1, out_ch=16) 16 filters → 16 maps input 1 channel Conv2d 3×3 1 → 16 maps BatchNorm2d 16 maps ReLU 16 maps MaxPool2d(2) 16 maps · ½ size

Only Conv2d grows the channels (1 → 16, the fan-out). BatchNorm and ReLU keep 16, and MaxPool keeps 16 but halves the width & height.

# the model's first block — turns 1 channel into 16 feature maps
self._conv_block(in_ch=1, out_ch=16)

def _conv_block(in_ch, out_ch):
    return nn.Sequential(
        nn.Conv2d(in_ch, out_ch, kernel_size=3, padding=1),  # 1 → 16
        nn.BatchNorm2d(out_ch),   # keep numbers stable
        nn.ReLU(),                # add non-linearity
        nn.MaxPool2d(2),          # shrink the image
    )
The code

train.py — the training loop

The whole loop is just a handful of lines. The four inner steps are exactly the Guess → Score → Blame → Adjust cycle from before.

# set up once
model = GenreCNN().to(device)
criterion = nn.CrossEntropyLoss()      # the loss
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

for epoch in range(EPOCHS):           # repeat 30x
    for inputs, targets in train_loader:   # one batch
        optimizer.zero_grad()          # reset
        outputs = model(inputs)        # 1 guess
        loss = criterion(outputs, targets)  # 2 score
        loss.backward()                # 3 blame
        optimizer.step()               # 4 adjust

What the key lines do

  • CrossEntropyLoss — measures how wrong the genre guesses are.
  • Adam — the optimizer that decides how to nudge each weight.
  • model(inputs) — the forward pass (the guess).
  • loss.backward() — backprop: works out every weight's blame.
  • optimizer.step() — applies the update; the model improves.

After each epoch we check accuracy on the held-out set and save the best model. That's the entire idea of training.

🎧✨

You've got the whole picture

Sound becomes an image, a CNN learns the patterns, and the model names the genre. Now go train it, break it, and improve it.

README.md → setup & commands src/ → read the code

Happy learning! Press ← to review any slide.

Cheat sheet

Words you now know

Neuron

A unit that weighs inputs, sums them, and fires through an activation.

Weights & bias

The tunable knobs. Training adjusts them to reduce errors.

Activation (ReLU)

The squiggle that lets networks learn curves, not just lines.

Loss

A number measuring how wrong the model is. Lower is better.

Gradient descent

Rolling downhill on the loss to find good weights.

Epoch

One full pass over all the training data.

Convolution

Sliding a small learned filter to detect local patterns.

Pooling

Shrinking a feature map while keeping the strongest signals.

Overfitting

Memorising the training data instead of truly learning.

Spectrogram

A picture of sound: time × pitch × energy.

Softmax

Turns raw scores into probabilities that sum to 100%.

Validation set

Held-back data for an honest accuracy check.