I Built an 11,200-Example Synthetic Dataset, Fine-Tuned a Model, and Then Didn't Ship It

A post-mortem on the Dragon Commentary Studio LoRA: what the data pipeline got right, why the fine-tune still lost to the base model, and why shipping the base model was the correct engineering call.

Good Omens Studio5 min readAI Systems

Most fine-tuning posts end with the fine-tuned model winning. This one doesn’t, and I think that makes it more useful.

For Dragon Commentary Studio — my AMD AI Hackathon video-captioning agent with four dragon personas — I wanted each dragon to have its own fine-tuned voice: per-dragon LoRA adapters on Qwen3-1.7B, routed at inference so each caption style used its own specialist model. I built the whole thing. Then I turned it off for the submission and shipped the base model. Here’s the full story.

The dataset pipeline

The plan: 6 categories × 4 dragons × 500–600 examples each, for 11,200 total training examples in OpenAI chat JSONL format.

Category Per dragon Type
Introductions 500 Single-turn
Waiting lines 300 Single-turn
Reactions (single) 600 Single-turn
Reactions (multi) 400 3–5 turn multi-dragon
Waiting conversations 500 3–5 turn multi-dragon
Reaction conversations 500 3–5 turn multi-dragon

The generation pipeline:

  1. Seed data came from 46 real pipeline runs — actual canonical observations and captions from the working system, not invented scenarios.
  2. Prompt variation: a generator script produced 6,200 prompt variants across the categories.
  3. Response generation via kimi-k2p6 with strict in-character enforcement per dragon.
  4. Validation: rule checks plus an LLM judge scoring each example, rejecting anything below 7/10.
  5. Assembly into per-dragon training splits (90/10 train/val), with a Modal fallback path (Qwen2.5-7B on A10G) for GPU-accelerated generation.

Training ran on Fireworks’ supervised fine-tuning API: Qwen3-1.7B base, LoRA rank 16, 3 epochs, learning rate 1e-4.

Nothing in this pipeline was wasted, by the way — even the rejects’ complement had value: 7,014 clean intro and waiting lines were extracted into categorized quick-chat pools that power the web UI’s live dragon banter. The dataset paid rent even before training started.

The result: subpar

With per-dragon LoRA routing wired up and auto-fallback to the base model in place, the honest assessment was that the fine-tuned adapters’ analysis of the messages was subpar. Not broken — subpar. The base text model, prompted with the rich persona profiles and few-shot style examples, was simply more reliable.

Why? My best analysis:

  • The dataset was too small for the job. 500–600 examples per category sounds like a lot until you split it across a persona’s full behavioral range. The multi-turn conversational categories especially were asking a 1.7B model to learn complex multi-dragon dynamics from a few hundred samples each.
  • Distribution mismatch. The training data taught persona voice, but inference-time inputs were canonical observation JSON requiring analysis — reading structured facts and reasoning about them before styling. The fine-tune sharpened the voice while the analytical capability is exactly what a small model has least to spare.
  • Compute constraints compounded it. This mirrors what happened on my previous project, Caro5, where I optimistically queued 7 personality models against Qwen3-4B overnight and learned that $20 of budget disagrees with ambition. Small budgets push you toward smaller models and fewer epochs, which is precisely when fine-tuning gains get fragile.

The decision: base model ships, adapters postponed

The hackathon scored a hidden 12-clip evaluation set. That framing made the decision for me: for hidden-set reliability, boring beats bespoke. The base model with strong persona prompting had a known, stable quality floor. The LoRA adapters had a higher ceiling on their best outputs and a lower floor on their worst — and a hidden test set is a machine for finding your worst outputs.

So the submission path defaulted to the base model, with LoRA routing left in the codebase as an optional flag, and per-dragon fine-tuning postponed until the dataset is bigger and the training budget allows more than a light pass.

What I actually learned

  1. A validation pipeline doesn’t guarantee a sufficient dataset. My judge filtered for quality per example; nothing filtered for coverage of the inference distribution. Those are different failure modes.
  2. Fine-tune for the task you’ll run, not the vibe you want. I trained on styled outputs and needed styled analysis of structured input. The gap between those two swallowed the gains.
  3. Evaluate against the deployment condition. The arena for this project wasn’t “which model writes the best dragon line” — it was “which system fails least on twelve clips I’ll never see.”
  4. Ship the reversible decision. The adapters aren’t deleted; they’re a flag. When the dataset grows, flipping it back on is trivial.

Fine-tuning wasn’t the wrong idea. It was the wrong week for it, with the wrong dataset size, judged by a metric that punishes variance. Knowing the difference — and being willing to turn off the thing you spent days building — is, I suspect, most of what “engineering judgment” means.

The dataset pipeline (prepare → generate → validate → assemble → finetune, all-in-one) lives in the Dragon Commentary Studio repo. My earlier fine-tuning work — where I did ship LoRA-trained commentary models, quantized under 1GB — is documented in my Caro5 posts.