Most fine-tuning posts end with the fine-tuned model winning. This one doesn’t, and I think that makes it more useful.
For Dragon Commentary Studio — my AMD AI Hackathon video-captioning agent with four dragon personas — I wanted each dragon to have its own fine-tuned voice: per-dragon LoRA adapters on Qwen3-1.7B, routed at inference so each caption style used its own specialist model. I built the whole thing. Then I turned it off for the submission and shipped the base model. Here’s the full story.
The dataset pipeline
The plan: 6 categories × 4 dragons × 500–600 examples each, for 11,200 total training examples in OpenAI chat JSONL format.
| Category | Per dragon | Type |
|---|---|---|
| Introductions | 500 | Single-turn |
| Waiting lines | 300 | Single-turn |
| Reactions (single) | 600 | Single-turn |
| Reactions (multi) | 400 | 3–5 turn multi-dragon |
| Waiting conversations | 500 | 3–5 turn multi-dragon |
| Reaction conversations | 500 | 3–5 turn multi-dragon |
The generation pipeline:
- Seed data came from 46 real pipeline runs — actual canonical observations and captions from the working system, not invented scenarios.
- Prompt variation: a generator script produced 6,200 prompt variants across the categories.
- Response generation via kimi-k2p6 with strict in-character enforcement per dragon.
- Validation: rule checks plus an LLM judge scoring each example, rejecting anything below 7/10.
- Assembly into per-dragon training splits (90/10 train/val), with a Modal fallback path (Qwen2.5-7B on A10G) for GPU-accelerated generation.
Training ran on Fireworks’ supervised fine-tuning API: Qwen3-1.7B base, LoRA rank 16, 3 epochs, learning rate 1e-4.
Nothing in this pipeline was wasted, by the way — even the rejects’ complement had value: 7,014 clean intro and waiting lines were extracted into categorized quick-chat pools that power the web UI’s live dragon banter. The dataset paid rent even before training started.
The result: subpar
With per-dragon LoRA routing wired up and auto-fallback to the base model in place, the honest assessment was that the fine-tuned adapters’ analysis of the messages was subpar. Not broken — subpar. The base text model, prompted with the rich persona profiles and few-shot style examples, was simply more reliable.
Why? My best analysis:
- The dataset was too small for the job. 500–600 examples per category sounds like a lot until you split it across a persona’s full behavioral range. The multi-turn conversational categories especially were asking a 1.7B model to learn complex multi-dragon dynamics from a few hundred samples each.
- Distribution mismatch. The training data taught persona voice, but inference-time inputs were canonical observation JSON requiring analysis — reading structured facts and reasoning about them before styling. The fine-tune sharpened the voice while the analytical capability is exactly what a small model has least to spare.
- Compute constraints compounded it. This mirrors what happened on my previous project, Caro5, where I optimistically queued 7 personality models against Qwen3-4B overnight and learned that $20 of budget disagrees with ambition. Small budgets push you toward smaller models and fewer epochs, which is precisely when fine-tuning gains get fragile.
The decision: base model ships, adapters postponed
The hackathon scored a hidden 12-clip evaluation set. That framing made the decision for me: for hidden-set reliability, boring beats bespoke. The base model with strong persona prompting had a known, stable quality floor. The LoRA adapters had a higher ceiling on their best outputs and a lower floor on their worst — and a hidden test set is a machine for finding your worst outputs.
So the submission path defaulted to the base model, with LoRA routing left in the codebase as an optional flag, and per-dragon fine-tuning postponed until the dataset is bigger and the training budget allows more than a light pass.
What I actually learned
- A validation pipeline doesn’t guarantee a sufficient dataset. My judge filtered for quality per example; nothing filtered for coverage of the inference distribution. Those are different failure modes.
- Fine-tune for the task you’ll run, not the vibe you want. I trained on styled outputs and needed styled analysis of structured input. The gap between those two swallowed the gains.
- Evaluate against the deployment condition. The arena for this project wasn’t “which model writes the best dragon line” — it was “which system fails least on twelve clips I’ll never see.”
- Ship the reversible decision. The adapters aren’t deleted; they’re a flag. When the dataset grows, flipping it back on is trivial.
Fine-tuning wasn’t the wrong idea. It was the wrong week for it, with the wrong dataset size, judged by a metric that punishes variance. Knowing the difference — and being willing to turn off the thing you spent days building — is, I suspect, most of what “engineering judgment” means.
The dataset pipeline (prepare → generate → validate → assemble → finetune, all-in-one) lives in the Dragon Commentary Studio repo. My earlier fine-tuning work — where I did ship LoRA-trained commentary models, quantized under 1GB — is documented in my Caro5 posts.