Every LoRA config file has the same handful of lines — r, alpha, dropout, target_modules, learning rate — and almost everyone (me included, at first) fills them in by copying a config that worked for someone else. This article is my attempt to actually understand each knob: what it does mechanically, what the research and accumulated community experience say, and how to reason about it instead of cargo-culting it. This is the most practical piece in the series, so I’ll organize it knob by knob.
Rank (r): how much capacity you’re renting
Rank is the width of the adapter bottleneck — the r in ΔW = B·A. It caps the expressiveness of the update: rank 8 can only make rank-8 changes to each targeted matrix. So it sounds like the quality dial, and the natural instinct is “more rank = better fine-tune.”
The consistent empirical finding — from the original paper onward — is more interesting: quality saturates with rank, usually early. For style, format, tone, chat behavior, and modest domain adaptation, ranks 4–16 typically capture nearly everything; doubling to 64 or 128 adds parameters, memory, and overfitting surface while moving evals barely or not at all. This is the low-intrinsic-rank story from article 1 cashing out in practice: if the change you’re asking for is genuinely low-dimensional, extra rank is empty capacity — and empty capacity on a small dataset doesn’t stay empty, it memorizes.
When does rank matter? When the adaptation is genuinely high-information: a new language, heavy domain shift (legal, medical corpora), complex structured skills, or continued-pretraining-flavored jobs. Community rules of thumb that match the published ablations: 8–16 for behavior and style; 32–64 for domains and skills; 128+ only when you have lots of data and evidence the lower rank is the bottleneck — at which point it’s worth asking whether you’re in full-fine-tuning territory anyway. The honest procedure is boring: start at 16, check evals, double only if underfitting is demonstrated. Rank is the knob where “bigger” most reliably wastes money silently, because it rarely hurts benchmarks — it just stops helping.
Alpha (α): the volume dial with a confusing label
The adapter’s output is scaled by α/r before being added to the frozen path. That coupling to r is the perennial confusion: change rank without touching alpha and you’ve also changed the effective strength of your adapter. Two conventions circulate: keep the α/r ratio fixed (commonly α = 2r, so the scale is a constant 2), or fix α at some value and accept that the scale shifts with r. The first is the sane default — it makes rank experiments actually be about rank.
There’s a subtler, well-established wrinkle: with the standard α/r scaling, high-rank LoRAs are effectively under-scaled — their per-direction learning slows as r grows, which is part of why naive high-rank runs disappoint. The rsLoRA fix scales by α/√r instead, and if you’re experimenting above rank ~64 it’s the theoretically better-grounded choice, usually one flag in modern libraries. The mental model that survives all the conventions: alpha (relative to rank) is the volume of the adapter’s voice against the frozen model. Too quiet and training must fight the scale; too loud and small parameter noise becomes large behavior noise. Fix the ratio, tune the learning rate instead, and treat exotic alpha values in copied configs with suspicion.
Target modules: which matrices get adapters at all
The original paper adapted only attention’s query and value projections, and that minimal recipe became folklore. The accumulated evidence since — QLoRA’s ablations prominently among it — points the other way for language models: targeting all linear layers, attention and MLP, matters more than raising rank. The intuition fits everything from my Transformer article: attention routes information, but the MLPs are where most parameters — and by strong evidence, most stored knowledge and computation — live. Adapting only attention lets you re-route what the model does; adapting the MLPs lets you adjust what it computes. For behavior-only tweaks the attention-only recipe still works and is cheaper; for anything domain- or skill-shaped, all-linear is the modern default (and is exactly what the mainstream fine-tuning frameworks now ship as their preset). Rank spent everywhere at 16 generally beats rank 128 spent on q and v alone.
Learning rate, dropout, and the supporting cast
Learning rate is the knob that actually breaks runs. LoRA wants rates roughly 10× higher than full fine-tuning — 1e-4 to 3e-4 is the well-worn band for SFT-scale jobs — because you’re training a tiny, zero-initialized side-car, not nudging billions of pretrained weights. Symptoms are legible: loss plateauing high with bland outputs means too low; loss spiking or the model turning weird and repetitive means too high. A cosine or linear decay with brief warmup is standard and rarely worth fighting.
An underrated companion finding (from the “LoRA learns less and forgets less” line of work and others): the two matrices aren’t symmetric — B (zero-initialized) and A behave differently under training, and giving B a higher learning rate (the LoRA+ recipe) measurably speeds convergence, especially at low rank. Again usually one flag.
Dropout on the adapter path (0.05–0.1) is cheap overfitting insurance for small datasets and mostly unnecessary past ~50k examples. Bias training is generally left off; embedding/lm-head adaptation only matters if you’ve added tokens or need vocabulary-level shifts. And remember the non-LoRA hyperparameters still exist: epochs (1–3 for SFT; small datasets overfit fast, and eval-loss creep is your stop signal), sequence length, and batch size do their usual jobs regardless of the adapter math.
The variants worth knowing (and the ones to ignore)
The LoRA-variant literature is enormous; the shortlist that has actually earned adoption:
DoRA decomposes each weight update into magnitude and direction, training them separately — analysis of full fine-tuning shows those components move differently, and vanilla LoRA entangles them. Result: consistently closer-to-full-FT quality, biggest gains at low rank, at a modest training-speed cost and zero inference cost (it merges like anything else). If quality at small rank is the goal, DoRA is the first upgrade to try. rsLoRA and LoRA+, covered above, are near-free fixes with solid theory. AdaLoRA allocates rank per-layer adaptively — conceptually right, practically niche. Beyond these, my filter for any new variant is unchanged: does it beat the boring recipe at matched total parameters and tuned LR? Most published wins evaporate under that control.
The config I’d actually start from
Pulling it together — for a typical instruction/domain fine-tune with QLoRA: r=16, α=32, all linear layers targeted, dropout 0.05, LR 2e-4 with cosine decay and warmup, 1–3 epochs, eval set held out per the data-series rules, and one deliberate ablation (r=32 or DoRA) only if evals show a gap. It’s almost anticlimactic — but that’s the actual lesson of this deep-dive. The knobs interact less mysteriously than the folklore suggests: rank is capacity (match it to the information content of the task), alpha-to-rank is volume (fix it and forget it), targets are coverage (all-linear unless you’re sure), and learning rate is the one that needs your respect. Everything I copied for months turns out to have a reason — and knowing the reasons is what lets you deviate when your task isn’t the average task.
Last article in the series: what happens after training — merging adapters into bases (and the QLoRA merge trap), stacking multiple LoRAs, and the genuinely elegant systems trick of serving hundreds of fine-tunes from one GPU.