Before any of this existed, the obvious way to build a thinking machine seemed to be writing down rules by hand. If a photo has whiskers and pointed ears, call it a cat. If an email contains certain words, call it spam. For decades, that’s exactly what engineers tried, and for decades it kept failing the moment reality got messy. Language has too many exceptions to list. Vision has too many angles, too much lighting, too many breeds of dog that only sort of look like the last one. The rule-writing approach hit a ceiling that no amount of extra effort could push past. What eventually broke through was giving up on written rules and replacing them with something that could tune itself. That something is the weight, and the full set of them is what people mean by a model’s parameters.
What a weight actually is
A weight is a single adjustable number that controls how much one piece of information should influence the result.
Picture an old stereo mixing board covered in small dials, one for bass, one for treble, one for each instrument's volume. Every dial on that board is doing the same job a weight does: deciding how loudly one particular input gets to speak in the final mix. Turn a dial up and that input matters more. Turn it down and it fades toward silence. A model is built from an enormous version of that same board, except instead of six or eight dials tuned by a sound engineer, there are billions of them, and no person has ever touched a single one directly.
So who sets them, if not a person? Nobody sets them. They settle into place through repetition. Before training begins, every dial sits near a random position, deliberately meaningless. Training is the long process of nudging each dial a tiny amount, checking whether the whole board's output got a little better or a little worse, and adjusting again. Repeat that cycle across billions of examples and the dials gradually settle into positions that, together, produce useful answers.
But nobody designs that final configuration. How does it end up being anything coherent?
It's discovered, one small correction at a time, the way a musician's fingers find the right position on a fretboard after enough repetition. The difference is that here the practice happens automatically, and at a scale no human could match.
The Nook of Wonder Theorems & Beautiful Patterns — what "random, then nudged" actually means
Weights typically start as small random values drawn from a distribution centered on zero, often written:
w ~ Normal(0, σ²)
Read as: "each weight w is drawn from a bell-curve distribution centered at 0, with some small spread σ." Starting small keeps early outputs from exploding into huge, meaningless numbers before training has taught the dials anything useful.
"Nudging" a weight during training means updating it in the direction that reduces error, scaled by a small step size called the learning rate:
w_new = w_old − η · (∂Loss / ∂w)
η (eta) is the learning rate, how big a step to take. The fraction is the gradient: how much the model's error would shift if this one weight moved slightly. Part 10 of this series unpacks gradients properly, worked example and all.
How a single dial actually gets used
Every input a model receives gets multiplied by its weight, and those results get added together to produce one number.
Imagine deciding whether to bring an umbrella, weighing a few pieces of evidence: how cloudy it looks, what the forecast says, whether your knee is aching, an old and slightly embarrassing but surprisingly reliable signal. You don't treat those three clues equally. The forecast probably counts for more than your knee. Somewhere in your head, each clue carries an informal weight, and you're combining them into a single decision. A model does the same arithmetic, with numbers standing in for the clues.
Why multiply, rather than picking whichever input looks biggest? Multiplication lets one weight express three different jobs. A high weight says pay close attention to this input. A weight near zero says mostly ignore it. A negative weight says this input should push the answer the other direction entirely, the way a sudden cold front might make you doubt a forecast that otherwise sounded sunny. One operation, three behaviors.
Stack that pattern across billions of dials, and what does it actually add up to?
Something capable of weighing an enormous number of subtle, interacting factors at once, far more than any hand-written rule could track.
The Nook of Wonder Theorems & Beautiful Patterns — the actual neuron equation, worked out
This multiply-and-add pattern has a name. It's a weighted sum, and it's the core operation of a single artificial neuron:
z = (x₁·w₁ + x₂·w₂ + x₃·w₃) + b
x values are inputs, w values are weights, and b is a bias, a constant nudge added after the sum that lets the neuron shift its whole output up or down regardless of the inputs.
Worked example, three inputs, three weights, one bias:
x = [2, 0.5, -1] w = [0.4, 1.2, 0.9] b = 0.1
z = (2 × 0.4) + (0.5 × 1.2) + (−1 × 0.9) + 0.1
z = 0.8 + 0.6 − 0.9 + 0.1 = 0.6
That single number, 0.6, usually passes through one more step called an activation function (commonly written σ or ReLU) before moving to the next neuron. That step is what lets networks capture curved, bendy patterns instead of only straight-line ones. Layers of these neurons, stacked and wired together, get built up into full networks in part 7.
the other half of the word "parameters"
Weights and parameters, are they the same thing?
Weights are the majority of a model’s parameters, but the bias terms sitting alongside them count too.
Weights
The dials that decide how much one input matters to another. Almost all of a model's parameters are weights.
"this input × 0.83"
Parameters
The umbrella term: every learned number in the model, weights and biases together.
weights + biases = parameters
Why bother naming biases separately if weights already do most of the work? Think back to the umbrella example. Suppose you’re a cautious person, someone who’d lean toward bringing an umbrella even on a middling day, before any of the actual evidence gets weighed. Weights alone can’t express that built-in leaning, because a weight only acts on an input that’s present. A bias is the model’s version of that baseline caution: a constant nudge applying no matter what the inputs say, shifting the whole decision up or down before the weighted evidence gets added in. In casual conversation people say “weights” and “parameters” interchangeably and it rarely causes confusion. Parameter count is the full tally, weights and biases combined, and weights make up the overwhelming majority.
The Nook of Wonder Theorems & Beautiful Patterns — actually counting the parameters in one layer
Real models organize their weights into grids called matrices (formalized as tensors in part 5). If a layer takes in n inputs and produces m outputs, its weight matrix needs one weight for every input-output pair:
parameters in one layer = (n × m) weights + m biases
Worked example: a layer taking 1,000 inputs and producing 1,000 outputs needs a 1,000×1,000 grid of weights, which is 1,000,000 weights, plus 1,000 biases, for 1,001,000 parameters in that one layer alone. Stack dozens of layers like this, many of them far larger, and "70 billion" stops sounding mysterious and starts sounding like simple accumulation.
so what does "70 billion parameters" mean
Putting the size in perspective
Parameter count is a rough proxy for how much capacity a model has to capture subtle patterns, though it’s far from the whole story.
Does a bigger number always mean a better model? No, and this turns out to be one of the more counterintuitive lessons in the field. More dials generally mean more room to represent nuance, in the same way a mixing board with sixty channels can capture a more detailed sound than one with six. A sixty-channel board is useless without good source material to feed it, and a model is no different. Past a certain scale, the quality and range of the training data starts to matter as much as, sometimes more than, the raw count of dials. Part 11 digs into exactly when a smaller model beats a larger one in practice.
Small enough to run on a phone or laptop.
The usual local-workload size, and the point where quantization matters.
Server-scale. Bigger is not automatically better for every task.
The Nook of Wonder Theorems & Beautiful Patterns — why parameter count drives memory size
Each parameter has to be stored as a number with a certain precision, meaning how many bits it's saved with. A common default is 16-bit floating point, 2 bytes per parameter:
model file size ≈ parameter count × bytes per parameter
Worked example: a 7-billion-parameter model at 2 bytes each needs roughly 14 billion bytes, about 14 GB, just to store the weights, before accounting for anything else the software needs while running. That memory bill is the whole reason quantization matters for running models on smaller hardware.
The lever for that bill is quantization, storing each weight with fewer bits than it was trained with.quantizationholding each weight in fewer bits than it was trained with, to shrink the model That topic gets its own proper treatment later in the series.
where this shows up
- Model names: The “7B” or “70B” in a model’s name is its parameter count
- File size: More parameters generally means a bigger download
- Hardware needs: Why some models run on a phone and others need a server rack