ML Foundations · Part 1

How Models Actually Learn: Gradient Descent and Backpropagation from First Principles

Every impressive thing a neural network does — writing code, recognizing faces, translating languages — comes down to one shockingly simple loop repeated billions of times: make a guess, measure how wrong it was, nudge every parameter a tiny bit in the dire…

Good Omens Studio7 min readMachine Learning

Every impressive thing a neural network does — writing code, recognizing faces, translating languages — comes down to one shockingly simple loop repeated billions of times: make a guess, measure how wrong it was, nudge every parameter a tiny bit in the direction that would have made it less wrong. That’s it. Everything else is detail. But the details are where the magic lives, and understanding them changed how I think about every model I use.

A model is just a function with knobs

Strip away the mystique and a neural network is a function: numbers in, numbers out. An image becomes a grid of pixel values; a sentence becomes a sequence of token IDs; the network transforms them through layers of matrix multiplications and simple nonlinear functions, and produces numbers at the other end — class scores, next-token probabilities, whatever the task demands.

The function has parameters — the weights — and modern networks have billions of them. Untrained, they’re random noise, and the function computes garbage. Training is the process of finding values for those knobs such that the function computes something useful. Framed that way, learning is a search problem: somewhere in a billion-dimensional space of possible weight settings, there are configurations that behave intelligently. How do you find one?

The loss function: turning “wrong” into a number

You can’t optimize what you can’t measure. The loss function takes the model’s prediction and the correct answer and returns a single number: how bad was this? For classification and language modeling, the standard choice is cross-entropy: it heavily punishes the model for being confidently wrong and rewards it for putting probability mass on the right answer.

This is a quietly profound move. The entire, fuzzy notion of “being good at the task” gets compressed into one scalar. Training doesn’t try to make the model “smart” — it tries to make one number go down. Every capability we observe is a side effect of that pressure. This is also why loss function design matters so much: the model becomes whatever the loss rewards, not what we intended.

Gradient descent: skiing downhill in a billion dimensions

Picture the loss as a landscape: every point is one possible setting of all the weights, and the altitude at that point is the loss. Training means finding a low valley. The landscape has billions of dimensions and we can’t see it — but at any point where we’re standing, calculus can tell us the slope.

The gradient is the vector of partial derivatives of the loss with respect to every weight — for each of the billions of knobs, it says “turning this knob up increases the loss by this much.” Gradient descent is then the obvious move: step in the exact opposite direction. Multiply the gradient by a small number — the learning rate — and subtract it from the weights. Repeat.

The learning rate is the most consequential hyperparameter in all of deep learning. Too large, and you leap over valleys, bouncing chaotically or diverging. Too small, and training crawls, possibly getting stuck in mediocre regions. In practice we don’t even keep it constant: schedules warm it up at the start (when gradients are wild) and decay it toward the end (to settle gently into a minimum).

Two more refinements make the raw idea practical:

Stochastic mini-batches. Computing the true loss over the entire dataset for every single step would be absurdly slow. Instead we estimate it on a small random batch of examples. The estimate is noisy — and the noise turns out to be a feature, not a bug: it jitters the trajectory enough to shake out of sharp, brittle minima and settle into wide, flat ones, which generalize better.

Momentum and Adam. Plain SGD is easily rattled by noisy or ill-scaled gradients. Momentum accumulates a running average of past gradients — like a heavy ball rolling downhill, it smooths out jitter and powers through small bumps. Adam goes further, adapting the effective step size per parameter based on the history of that parameter’s gradients. Adam (and its weight-decay-corrected variant AdamW) is the default optimizer for nearly every large model trained today.

Backpropagation: computing a billion derivatives efficiently

Gradient descent needs the gradient — the derivative of the loss with respect to every weight. For a billion weights, computing each derivative independently would take a billion forward passes. Backpropagation computes all of them in roughly the cost of two.

The insight is the chain rule from calculus, applied systematically. A network is a composition of simple operations, and the derivative of a composition is the product of the derivatives of its pieces. So: run the forward pass and remember the intermediate values at every layer. Then walk backward from the loss, layer by layer, multiplying local derivatives as you go. Each layer receives “how much does the loss change if my output changes?” from the layer above it, uses its stored intermediates to convert that into “how much does the loss change if my weights change?” (that’s the gradient it needed), and passes “how much does the loss change if my input changes?” down to the layer below.

One backward sweep, and every parameter in the network knows exactly how it contributed to the error. This is the algorithm — popularized for neural networks by Rumelhart, Hinton, and Williams in 1986 — that makes deep learning computationally possible at all. Modern frameworks like PyTorch automate it completely: you write the forward pass, and automatic differentiation constructs the backward pass for you. But the memory cost of storing those intermediate activations is real, and it’s why tricks like gradient checkpointing (recompute instead of store) exist for training very large models.

Why deep learning shouldn’t work — and does

Here’s what puzzled me most when I learned this. The loss landscape of a deep network is wildly non-convex — riddled with hills, valleys, and saddle points. Gradient descent is a purely local, greedy algorithm. It should get stuck in terrible local minima constantly. Classical optimization theory basically predicts deep learning shouldn’t work.

Yet it does, and the emerging understanding is fascinating: overparameterization is the cure, not the disease. In very high dimensions, true local minima that are much worse than the global one become vanishingly rare — most critical points are saddle points, which have escape directions, and the noise of stochastic gradients finds them. With far more parameters than strictly necessary, there isn’t one needle in a haystack; there are astronomically many good solutions, and huge connected regions of low loss between them. The bigger the model, the easier the optimization gets. This is one of the deep reasons scale works.

The second puzzle is generalization: a network with a billion knobs can memorize its training set outright, and classical statistics says such a model must fail on new data. But gradient descent has an implicit bias — among all the settings that fit the training data, it tends to find simple ones, smooth functions that interpolate rather than contort. Regularization techniques (weight decay, dropout, data augmentation, early stopping) reinforce this preference, but a surprising amount of it comes free with the optimizer itself.

Watching it happen

The most concrete way to internalize all this is to watch a loss curve. Early in training, loss drops fast — the model is learning the easy statistical regularities (in language: common words, basic syntax). Then it slows into a long grind where the real capabilities form. If the training loss keeps falling but validation loss turns upward, the model has stopped learning and started memorizing — overfitting — and it’s time to stop, regularize, or get more data.

Every advanced topic I’ll cover in later posts — fine-tuning, LoRA, RLHF, distillation — is a variation on this same loop with a different loss or a different set of trainable knobs. Once you truly see training as “a number goes down, and every parameter is nudged by its share of the blame,” nothing in modern ML is mysterious. Complicated, yes. Mysterious, no.