Training a capable model is only half the battle. The model that comes out of pretraining or fine-tuning is almost never the model you actually want to ship. It’s too big, too slow, too expensive to serve, and often carries capacity it doesn’t need for your use case. Model optimization is the discipline of closing that gap — taking a model that works and turning it into a model that works in production, on your hardware, within your latency and cost budget.
This article walks through the full optimization stack: knowledge distillation, quantization, pruning and sparsity, parameter-efficient adaptation, and the inference-time optimizations that tie it all together. The goal is not just to name the techniques but to explain why each one works, what it costs you, and how to decide which combination fits your situation.
Why optimization is possible at all
Every optimization technique exploits the same underlying fact: neural networks are massively overparameterized. Modern training needs that overparameterization — large, redundant parameter spaces make optimization landscapes smoother and let stochastic gradient descent find good solutions. But once training is done, most of that capacity is dead weight. The lottery ticket hypothesis showed that small subnetworks inside large trained models can match the full model’s performance. Intrinsic dimensionality research showed that fine-tuning effectively happens in a tiny subspace of the full parameter space. Quantization research showed that weights stored in 16 bits of precision often carry far less than 16 bits of useful information.
In other words: the information a trained model actually uses is much smaller than the container it lives in. Optimization is the art of shrinking the container without spilling the contents.
There are three broad families of approach, and they compose:
- Change the model itself — distillation, pruning, architecture changes.
- Change the numeric representation — quantization of weights, activations, and KV cache.
- Change how inference runs — batching, caching, speculative decoding, compilation.
A serious deployment usually uses techniques from all three.
Knowledge distillation: transferring capability into a smaller body
Distillation trains a small “student” model to imitate a large “teacher” model. The classic formulation, from Hinton, Vinyals, and Dean’s 2015 paper, is deceptively simple: instead of training the student on hard labels (the single correct answer), train it on the teacher’s full output distribution — the soft targets.
Why does this work better than just training the small model from scratch on the same data? Because the teacher’s probability distribution contains what Hinton called “dark knowledge.” When a well-trained image classifier sees a picture of a BMW, it assigns high probability to “BMW,” a small-but-nonzero probability to “garbage truck,” and an astronomically smaller one to “carrot.” Those relative probabilities encode the similarity structure of the world — which mistakes are reasonable and which are absurd. Hard labels throw that structure away; soft targets preserve it, giving the student a far richer training signal per example.
In practice, the student’s loss combines two terms: a standard cross-entropy against the true labels and a divergence term (usually KL divergence) between the student’s and teacher’s softened distributions, with a temperature parameter that flattens both distributions to expose the small probabilities that carry the dark knowledge.
Distillation in the LLM era
For language models, distillation has evolved into several distinct practices:
Logit-based distillation is the classic method applied to next-token prediction: the student matches the teacher’s token distribution at every position. This requires access to the teacher’s logits, so it works when you own both models. DistilBERT is the canonical early example — roughly 40% smaller and 60% faster than BERT while keeping about 97% of its language understanding performance, using a combination of soft-target loss, standard masked language modeling loss, and a hidden-state alignment loss.
Sequence-level / data distillation is what most people mean today when they say “distillation” for LLMs: use the teacher to generate training data — instructions, reasoning traces, preference labels — and fine-tune the student on those outputs with an ordinary supervised objective. No logit access needed, which is why this became the dominant mode when strong teachers were only reachable through APIs. The entire family of models fine-tuned on frontier-model outputs works this way. Reasoning distillation — training a small model on a large model’s chain-of-thought traces — has proven especially potent: a compact model that imitates a strong reasoner’s process, not just its answers, recovers a surprising fraction of the reasoning ability.
Intermediate-representation distillation matches hidden states or attention maps between teacher and student (TinyBERT, MobileBERT). This transfers more structure but requires architectural compatibility and more engineering.
What distillation costs you
Distillation is a training procedure, so it costs training compute, data, and iteration time. The student has a genuine capacity ceiling: a 1B model cannot absorb everything a 70B teacher knows, so you must decide which capabilities matter. Distillation is at its best when your deployment needs a specialist — you distill exactly the behaviors your product uses, and the student can match or even exceed the teacher on that narrow slice while being an order of magnitude cheaper.
Quantization: doing the same math with smaller numbers
Quantization reduces the numeric precision of a model’s parameters (and optionally its activations) from 16- or 32-bit floats to 8-bit, 4-bit, or even lower representations. It attacks the two costs that dominate LLM inference: memory footprint and memory bandwidth. A 70B-parameter model at FP16 needs ~140 GB just for weights; at 4 bits it needs ~35 GB and fits on a single high-end GPU. And because autoregressive decoding is memory-bandwidth-bound — every generated token requires streaming the entire weight matrix through the compute units — halving the bytes per weight can nearly double token throughput even with no change in arithmetic.
The core operation maps a range of float values onto a small integer grid: pick a scale (and possibly a zero-point) per tensor, per channel, or per small group of weights, round each value to the nearest grid point, and store the integers plus the scales. The finer the granularity (per-group scales of 32–128 weights are standard at 4-bit), the better the fidelity, at a small storage overhead.
The outlier problem
Naive quantization of large language models fails in a characteristic way, and understanding why is the key to understanding the whole modern quantization literature. Above roughly 6–7B parameters, transformer activations develop systematic outlier features — a small number of hidden dimensions with magnitudes up to 100× larger than the rest. A single scale must stretch to cover the outliers, crushing all the normal values into a few grid points and destroying accuracy.
Every major method is a different answer to this problem:
- LLM.int8() keeps the outlier dimensions in 16-bit and quantizes everything else to 8-bit, doing mixed-precision matrix multiplication. Simple, accurate, no calibration needed.
- SmoothQuant observes that outliers live in activations, while weights are well-behaved, and migrates the difficulty: scale activations down and weights up by a per-channel factor, making both easy to quantize to INT8 and unlocking fast integer kernels.
- GPTQ treats quantization as a layer-wise optimization problem: quantize weights one column at a time, using second-order (Hessian) information from a small calibration set to update the remaining weights and compensate for each rounding error. This is what made 3–4 bit weight quantization of very large models practical.
- AWQ notes that a tiny fraction (~1%) of weight channels are disproportionately salient — determined by activation magnitudes, not weight magnitudes — and protects them by per-channel scaling before quantization, achieving strong 4-bit accuracy with no backpropagation and good hardware efficiency.
- The GGUF k-quant and i-quant families (the formats behind llama.cpp) use block-wise quantization with hierarchical scales, mixed precision across layers, and importance-matrix weighting, bringing usable 2–5 bit models to consumer CPUs and GPUs.
PTQ vs. QAT, and where the floor is
Everything above is post-training quantization (PTQ): take a finished model, quantize it in minutes-to-hours with a small calibration set. Quantization-aware training (QAT) instead simulates quantization during training with a straight-through estimator, letting the network learn weights that are robust to rounding. QAT costs real training compute but buys accuracy at aggressive bit-widths, and it’s how the most extreme results are achieved — BitNet b1.58 trains models with ternary weights {-1, 0, +1} (~1.58 bits) from scratch and matches full-precision quality at moderate scales, hinting at a future where matrix multiplication becomes mostly addition.
As practical guidance: 8-bit weight quantization is essentially free. 4-bit weight quantization with a good method (AWQ, GPTQ, or modern GGUF k-quants) costs little on most tasks and is the standard deployment point today. Below 4 bits, degradation becomes noticeable and task-dependent — acceptable for casual use, risky for reasoning-heavy or precision-sensitive workloads. Larger models degrade more gracefully than smaller ones at the same bit-width, which produces the well-known rule of thumb that a bigger model at 4-bit usually beats a smaller model at 16-bit for the same memory budget.
Don’t forget the KV cache: for long contexts and large batches, the attention cache can rival the weights in memory. Quantizing the KV cache to 8-bit is generally safe and often the difference between fitting a workload and not.
Pruning and sparsity: removing what the model doesn’t use
Pruning deletes parameters outright. Unstructured pruning zeroes individual weights (classically, the smallest-magnitude ones) and can remove a large fraction with minimal accuracy loss — but irregular sparsity is hard to accelerate on GPUs, so the wins are mostly in storage unless you have hardware support (NVIDIA’s 2:4 semi-structured sparsity being the notable exception, giving a real ~2× math throughput). Structured pruning removes whole units — attention heads, neurons, layers — producing a genuinely smaller dense model that runs faster everywhere, at the cost of more accuracy loss per parameter removed.
For LLMs, one-shot pruning methods like SparseGPT and Wanda showed that 50% of weights can be removed from large models without retraining using clever layer-wise reconstruction or activation-aware importance scores. The most effective modern recipe, though, combines pruning with distillation: prune a large model down (depth and width), then distill the original model into the pruned one to heal the damage. This prune-and-distill pipeline is how several production “small” models are actually made — carved from a larger sibling rather than trained from scratch, at a fraction of the cost.
Parameter-efficient adaptation: LoRA and friends
Optimization isn’t only about inference — it’s also about the cost of specializing a model. Full fine-tuning of a large model duplicates the entire parameter set per task and requires optimizer states several times the model’s size in memory. LoRA (Low-Rank Adaptation) sidesteps this by freezing the base weights and learning small low-rank update matrices: for a weight matrix W, learn ΔW = BA where B and A are thin matrices of rank r (often 8–64). Trainable parameters drop by ~10,000× versus full fine-tuning, and because BA can be merged into W after training, there is zero inference latency penalty.
QLoRA pushed the memory frontier further: quantize the frozen base model to 4-bit (using the NF4 data type, double quantization, and paged optimizers) and train LoRA adapters in 16-bit on top. This made fine-tuning 65B-class models on a single 48 GB GPU possible and effectively democratized large-model fine-tuning.
For deployment, adapters have a second superpower: multi-tenancy. Because the base model is shared and each adapter is tiny (megabytes), a single server can hold one base model and hot-swap hundreds of task- or customer-specific adapters, batching requests across all of them. One base model, many specialized behaviors, marginal cost per specialization near zero.
Refinements like DoRA (decomposing updates into magnitude and direction) and rank-adaptive variants close most of the remaining quality gap with full fine-tuning. The practical picture in 2026: LoRA-family methods are the default for adaptation, and full fine-tuning is reserved for cases with large data and large behavioral shifts.
Inference-time optimization: the serving stack
The final family of techniques changes nothing about the model’s parameters — only how inference executes. These often deliver the largest real-world wins.
Continuous batching. Autoregressive generation means requests finish at different times. Naive batching waits for the slowest request; continuous (in-flight) batching admits new requests into the batch the moment any sequence finishes, keeping the GPU saturated. This alone can improve throughput by an order of magnitude in serving systems.
PagedAttention / KV-cache management. vLLM’s key insight was that KV-cache memory was being wasted by fragmentation and over-reservation. Managing the cache in small pages, virtual-memory style, with copy-on-write sharing for common prefixes, dramatically raises the number of concurrent sequences a GPU can hold. Prefix caching extends this across requests: shared system prompts and few-shot preambles are computed once and reused.
FlashAttention. Attention is memory-bound; FlashAttention restructures the computation into tiles that stay in fast on-chip SRAM, avoiding materialization of the full attention matrix. It computes exact attention faster and in linear memory, and it’s now simply the default kernel everywhere.
Speculative decoding. Decoding is sequential and bandwidth-bound, but verification is parallel. So let a small draft model propose several tokens cheaply, then have the large model verify them all in one forward pass, accepting the longest correct prefix via a rejection-sampling scheme that provably preserves the large model’s output distribution. Typical speedups are 2–3× with zero quality loss. Variants like Medusa (extra decoding heads) and self-speculative approaches remove the need for a separate draft model.
Compilation and kernel fusion. Graph compilers (TensorRT-LLM, torch.compile, ONNX Runtime) fuse operations, eliminate framework overhead, and pick optimal kernels per shape and hardware. Routine 1.5–3× gains, more on small models where overhead dominates.
Architecture-level choices made before training also count as optimization: grouped-query attention (GQA) shrinks the KV cache several-fold, mixture-of-experts (MoE) decouples parameter count from per-token compute, and sliding-window or hybrid attention variants tame long-context cost.
Putting it together: a decision framework
Optimization techniques compose, and the right stack depends on your constraint. A practical way to think through it:
1. Define the target. Hardware (one A100? a phone? a CPU server?), latency budget (interactive chat needs fast time-to-first-token and ~20+ tokens/sec; batch jobs only need throughput), and the specific capabilities your application uses. Every downstream decision follows from these.
2. Right-size the model first. The cheapest optimization is not running capacity you don’t need. If your task is narrow, a distilled or pruned-and-distilled specialist at 3–8B often matches a general 70B model on that task. Distillation gives you the biggest single reduction — 10× or more — but costs a training project.
3. Quantize by default. Weights to 4-bit with AWQ/GPTQ/k-quants (or 8-bit if you have headroom and want zero risk), KV cache to 8-bit for long-context workloads. This is hours of work for 2–4× memory and bandwidth savings. Evaluate on your tasks, not just perplexity — quantization damage is task-dependent and perplexity hides reasoning regressions.
4. Specialize with adapters, not copies. LoRA/QLoRA for every task variant; merge adapters for single-task deployment, hot-swap them for multi-tenant serving.
5. Fix the serving layer. Run a modern engine (vLLM, TensorRT-LLM, SGLang, or llama.cpp at the edge) to get continuous batching, paged KV cache, FlashAttention, and prefix caching for free. Add speculative decoding if latency matters and you can pair a draft model.
6. Measure the right things. Track time-to-first-token, inter-token latency, throughput at your real concurrency, memory headroom, and task accuracy — together. Every technique trades among these, and the trade-offs interact: quantization speeds up decoding but can slightly lower speculative acceptance rates; aggressive batching raises throughput but hurts tail latency; pruning plus 4-bit quantization can compound damage that neither causes alone. Always evaluate the final composed stack, not each step in isolation.
The mental model to keep
Every technique in this article is a different way of spending one budget to save another: distillation spends training compute to save inference compute forever; quantization spends a little accuracy to save memory and bandwidth; LoRA spends a little expressiveness to save adaptation cost; speculative decoding spends parallel FLOPs to save sequential latency. Optimization is not a single trick but a portfolio, and the teams that ship the fastest, cheapest models are the ones that understand which budget is actually binding for their workload — and spend everything else against it.
The direction of travel is clear: precision keeps falling (4-bit is standard, ternary is coming), small specialized models keep closing the gap on large generalists through better distillation, and more of the optimization stack keeps moving into training itself (QAT, natively quantized and MoE architectures). The container keeps shrinking. The contents, remarkably, keep fitting.