The first three articles in this series were about making an adapter. This one is about what makes adapters a genuinely different kind of artifact from a fine-tuned model: what you can do with them afterward. Because a LoRA is a small, separable, additive object — ΔW = B·A, sitting beside frozen weights — it can be merged, unmerged, combined, weighted, and hot-swapped in ways a monolithic fine-tune never could. This composability turned out to be LoRA’s most underrated property, and it’s the foundation of both the creative-AI ecosystem and some elegant production infrastructure.
Merging: folding the adapter into the base
The simplest post-training operation: compute W′ = W + (α/r)·B·A for every adapted matrix, save the result, delete the adapter. The output is an ordinary model — same architecture, same size, same speed as the base — with the fine-tune permanently baked in. This is the standard move for single-task deployment: zero added latency, no adapter-aware serving code, compatible with every downstream tool (including quantization for deployment, which you do after merging).
Two pieces of merge hygiene the community learned the hard way. First, the QLoRA trap from article 2: adapters trained against a 4-bit base learned to fit that slightly-lossy model. Merge them into the pristine 16-bit weights and the combination differs subtly from what you evaluated during training. Usually it’s fine; occasionally it measurably drifts. The careful options: evaluate the merged 16-bit model explicitly (not just the training-time checkpoints), or merge into a dequantized copy of the exact base you trained against. Second, merging is a one-way door in practice — W′ − W recovers the update only if you kept the original base pristine and versioned. Treat bases as immutable artifacts and adapters as diffs, and you keep the whole system reversible; sloppy base management is how teams end up with mystery models nobody can reproduce.
Stacking: adapter arithmetic
Because adapters are additive, multiple adapters can be applied at once: W + ΔW₁ + ΔW₂, each with its own weighting coefficient. The generative-image world turned this into a folk art — a checkpoint plus a style LoRA at 0.8 plus a character LoRA at 0.6, dialed by feel — and it works far more often than it has any right to, for a reason that connects back to my embeddings article: adapters are directions in a very high-dimensional space, and in high dimensions, independently-trained directions are nearly orthogonal by default. Style-ness and character-ness barely overlap, so they superimpose with only mild interference.
“Nearly” is load-bearing. Stack adapters trained on similar data, or too many at once, or at aggressive weights, and the interference shows up — quality degradation, concept bleed, models pulled in contradictory directions. The LLM world hits this when combining, say, a domain adapter with a style adapter: sometimes free lunch, sometimes measurable regression on both tasks, and the only reliable arbiter is your eval set. A more principled cousin — task arithmetic — treats fine-tune deltas as vectors you can add and even subtract (subtracting a “toxic completion” delta as a detox operation is the famous party trick), and merge methods like TIES and DARE (trim small-magnitude changes, resolve sign conflicts, then combine) demonstrably reduce interference when merging several task adapters into one. It’s a genuinely useful toolbox, with one boundary worth stating plainly: merging composes behaviors, cheaply and lossily. It does not compose knowledge reliably — two adapters each knowing half a domain don’t merge into one that knows the whole domain.
Multi-adapter serving: the systems payoff
Here’s where the separation of base and adapter becomes real infrastructure. Consider a platform fine-tuning per customer — support bots, per-tenant styles, per-product specialists. With full fine-tunes, a hundred customers means a hundred multi-gigabyte models; serving them concurrently means a hundred GPU allocations, mostly idle. Economically impossible below enterprise scale.
With adapters: one frozen base model in GPU memory, plus a hundred adapters at tens of megabytes each. Adapters swap in milliseconds (they’re small enough to keep resident or stream on demand). Better still, requests for different adapters can be batched together: systems in the S-LoRA / Punica lineage — and the multi-LoRA support now built into mainstream servers like vLLM — run the shared base computation as one big efficient batch, then apply each request’s tiny B·A path individually with custom kernels. The base model, which is 99%+ of the FLOPs, is computed once per batch regardless of how many distinct fine-tunes are represented in it.
The economics this unlocks deserve a pause: the marginal cost of an additional fine-tuned model drops to roughly zero — megabytes of storage and a rounding error of compute. This is precisely why fine-tuning-as-a-service products can offer cheap per-customer models, and it resolves what would otherwise be an ugly tension in the whole customization story: specialization used to mean fragmentation (every fine-tune a fork, every fork a serving bill), and adapters turn it into configuration — one shared foundation, per-tenant behavior as an attachable file. The mental model I keep returning to: the base model is an operating system; LoRAs are apps.
The router pattern, and where this is all heading
Once serving many adapters is cheap, a design space opens above it: which adapter should handle this request? A lightweight classifier routing queries to a math adapter, a code adapter, a writing adapter — each small and sharp — starts to resemble a mixture-of-experts assembled after the fact, from independently trained parts. Research along these lines (adapter libraries with learned routing, retrieval over banks of thousands of task adapters) keeps reporting the same encouraging shape: a well-chosen small specialist beats a mediocre generalist, and choosing can be automated. I don’t think it’s settled how far this modular vision goes — frontier labs still bet primarily on monolithic scale — but as the economical path to mass customization, it has already won: it’s how the ecosystem works today, from image-generation communities trading style files to SaaS platforms shipping a fine-tune per customer.
Closing the series where it started: LoRA began as a memory optimization — a way to afford fine-tuning. What it actually did was change the ontology of fine-tuning. A specialization stopped being a model and became a file: diffable, shareable, weightable, stackable, swappable at runtime, priced near zero at the margin. The deepest ideas in this field keep having that character — a compression trick that turns out to be a statement about structure (updates are low-rank because adaptation is steering), and an efficiency hack that turns out to redesign the ecosystem built on top of it. Whatever I fine-tune next, I’ll be thinking of it not as changing a model, but as writing one of these files.
Next and final series before I circle back to the big optimization guide: quantization — the other half of making models fit, and the deepest rabbit hole of the bunch.