Full fine-tuning of a large language model is brutally expensive — not because of the forward pass, but because training multiplies memory. Every weight needs a gradient; the Adam optimizer keeps two more running statistics per weight; suddenly a 7B model that fits on a gaming GPU for inference wants 70+ GB just to learn. For years the assumption was: that’s the price of specialization. Then LoRA (Hu et al., 2021) showed the price was mostly imaginary — you can freeze the entire model, train a set of tiny add-on matrices amounting to a fraction of a percent of the parameters, and match full fine-tuning quality on most tasks. This article is about how it works and, more interestingly, why something so small is enough.
The memory problem, precisely
Where does training memory actually go for a 7B model in 16-bit? Weights: 14 GB. Gradients (one per weight): another 14 GB. Adam’s two moment estimates, typically kept in 32-bit: ~56 GB more. Plus activations saved for backprop. Rough total: north of 84 GB to fine-tune a model that needs 14 GB to run. The multiplier — roughly 6× the weight memory — is why fine-tuning felt like a datacenter activity. And every fine-tuned variant is a full 14 GB copy to store and load, so ten specialized versions means ten complete models.
Notice what the arithmetic implies: the cost isn’t the model — it’s the trainable-ness of the model. Shrink the number of trainable parameters and gradients, optimizer states, and per-task storage all shrink with it. The question becomes: how few trainable parameters can you get away with?
The insight: updates are low-rank
Here’s the empirical observation LoRA is built on, and it’s genuinely deep. When you fully fine-tune a pretrained model, you get an update ΔW for each weight matrix — the difference between final and initial weights. Analyze those updates and they turn out to have very low intrinsic rank: the change is not an arbitrary reshuffling of millions of numbers, but something highly structured, describable in far fewer dimensions. Earlier work (Aghajanyan et al.) had shown the same thing from another angle — fine-tuning succeeds even when constrained to a random subspace of just a few thousand dimensions. The picture that emerges: pretraining did the hard work of building general-purpose features; adaptation is a small rotation and re-weighting of what already exists. You’re not teaching the model new machinery; you’re steering machinery it has.
If the needed change is low-rank, why parameterize it with a full-rank matrix? A rank-r update to a d×d matrix can be written as the product of two thin matrices: ΔW = B·A, where A is r×d and B is d×r. For d = 4096 and r = 8, that’s ~65k parameters instead of ~16.7 million — a 256× reduction — for that one matrix, while still being able to express any rank-8 change. LoRA simply does this everywhere it matters: freeze W, learn B and A, and compute each layer’s output as W·x + B·A·x (scaled by a constant α/r, a knob I’ll dig into in article 3).
Two small choices in the design carry real weight. A is initialized with random noise and B with zeros — so B·A starts as exactly zero, meaning at step one the model is precisely the pretrained model, and training departs smoothly from a known-good point rather than from a random perturbation. And the adapters are additive side-cars: the frozen weights are never touched, which is what makes everything in article 4 (merging, swapping, stacking) possible.
What this buys you, concretely
Run the memory math again for that 7B model with LoRA at rank 16 on the attention matrices: trainable parameters drop from 7B to a few tens of millions — well under 1%. Gradients: tiny. Optimizer states: tiny. The dominant memory cost left is just the frozen weights held for the forward pass (14 GB — and shrinking that is exactly what QLoRA does, next article). Practical consequences, in rough order of how much they changed the ecosystem:
Fine-tuning on one GPU. What needed a node now fits on a single consumer or prosumer card. This single fact created the open fine-tuning community — the explosion of specialized variants of every open model traces directly to LoRA’s memory arithmetic.
A fine-tune becomes a file, not a model. The artifact you produce is the adapter: tens of megabytes. Sharing, versioning, and hoarding task-specific specializations became as cheap as sharing images. (The generative-art world ran with this hardest — “a LoRA” is simply the unit of customization for diffusion models: a style, a character, a concept, each a small downloadable file.)
Zero inference penalty — by construction. Because B·A has the same shape as W, you can merge the adapter into the base weights (W ← W + B·A) after training and serve a completely ordinary model. LoRA’s cost exists only during training. This is the property that separates it from earlier adapter methods (which inserted extra layers and paid latency forever) and from prompt/prefix-tuning (which spends context and steers more weakly).
Regularization for free, mostly. With <1% of parameters trainable and the base frozen, catastrophic forgetting is structurally damped — the model cannot drift arbitrarily far. Studies comparing LoRA to full fine-tuning find LoRA forgets less and stays closer to base-model behavior, at the price of somewhat less capacity to absorb genuinely large distribution shifts. “LoRA learns less and forgets less” is the honest one-line summary of the literature, and for most practical adaptations — style, format, domain, task procedure — that trade is exactly the right one, for exactly the reason the LIMA result suggested: most adaptation is low-information steering, not new knowledge.
The limits, stated honestly
Low-rank is a hypothesis about your task, and it can be false. Where the needed change is genuinely high-rank — teaching a model a new language it barely saw, continued pretraining on a large novel corpus, deep new capabilities rather than redirection of existing ones — full fine-tuning (or high-rank LoRA, which converges toward the same cost) still wins, and papers measuring the gap find it concentrated precisely in those regimes, with code and math sitting in the middle. The working heuristic I’ve adopted: the more your task resembles “select and stylize what the model already knows,” the better LoRA does; the more it resembles “install knowledge that isn’t there,” the more rank — or real fine-tuning, or RAG — you need.
There’s also a subtle scientific point buried in LoRA’s success that I keep thinking about: it’s evidence about what fine-tuning fundamentally is. If a rank-8 nudge to a few matrices can turn a base model into a chatbot, a SQL specialist, or a medical-tone assistant, then those behaviors were already latent in the pretrained network — reachable, sitting a small rotation away. Fine-tuning is less like education and more like tuning a radio: the stations were always broadcasting; you’re adjusting the dial. That framing (which the interpretability world would state as: capabilities are directions in representation space, and LoRA learns pointers to them) is, to me, the real intellectual payoff of the method — and the cleanest explanation of why it works at all.
Next article: the second half of the memory story. LoRA removed the cost of training the parameters; QLoRA removes most of the cost of holding the frozen model underneath — and it’s the reason a 65B-parameter fine-tune fits on hardware you can buy at a normal electronics store.