Keeping a Vision Pipeline Honest: Three Hallucination Guards I Built for a Video-Captioning Agent

How Dragon Commentary Studio separates perception from schema, filters OCR by consensus, and gates every caption through a critic loop.

Good Omens Studio5 min readAI Systems

For the AMD AI Hackathon (Track 2), I built Dragon Commentary Studio: a Dockerized agent that watches any short video and generates captions in four distinct styles — formal, sarcastic, humorous-tech, and humorous-non-tech — each voiced by a dragon persona. The judging was an LLM scoring caption accuracy and style match on a hidden 12-clip set, with a hard 10-minute runtime budget and strict JSON output.

The style part was fun. The hard part — and the part I think generalizes to any production VLM system — was keeping the captions grounded. A sarcastic dragon is entertaining; a sarcastic dragon confidently describing things that aren’t in the video is a failing score and, in a real product, a trust-destroying bug.

Here are the three guards that did the work.

Guard 1: Separate perception from schema construction

The naive approach is to hand the vision model a batch of frames and ask it to return one big JSON describing the video. I learned not to do this. Batch JSON prompts can spill reasoning text into the output or hallucinate aggregations — conclusions about the video that no single frame supports.

Instead, the design principle became: VLMs describe single-frame facts; deterministic code owns the JSON.

Phase A of the pipeline sends selected representative frames (chosen by PySceneDetect scene boundaries, with a uniform 1 FPS fallback) to the vision model in one request — but the model must return a separate constrained tagged block per frame: subjects, actions, objects, OCR, camera movement, scene details. Python then matches blocks by frame ID, normalizes the facts, and merges them into a canonical JSON schema. If a frame’s analysis comes back malformed, it retries once, then becomes an empty observation rather than corrupting the merge.

The result: no reasoning text can leak into the schema, and no cross-frame conclusion exists unless my code computed it. The canonical JSON becomes the only factual substrate downstream models are allowed to work from.

Guard 2: Cross-batch OCR consensus

OCR was the biggest single source of hallucination. Vision models love to “read” text that isn’t there, and they do it convincingly.

The fix exploits redundancy. When a video spans multiple frame batches, OCR is extracted independently per batch, then matched across batches with normalized text comparison. Text that appears in two or more batches survives; singletons are discarded. A word the model imagined in one batch almost never gets imagined identically in another, so consensus filtering removes single-batch hallucinations almost for free.

One edge case worth the special-casing: single-batch (short) videos bypass the filter entirely. Otherwise, legitimate on-screen text in short clips would be filtered out for lacking a second batch to agree with it. Guards need escape hatches or they start destroying true positives.

Guard 3: A validation critic loop that knows metaphor from fabrication

Phase B turns the canonical JSON into four styled captions. Every caption then passes through a validator scoring five weighted dimensions: factuality, tone, specificity, hallucination control, and completeness. Fail the gate, and the caption is revised and retried.

The subtle design problem: two of my four styles are supposed to be figurative. When the sarcastic dragon says she’s “watched galaxies blink out with less ceremony,” that’s style, not a factual claim about galaxies. So the validator had to distinguish grounded metaphors from literal unsupported facts — flowery language about things visible in the video passes; a plain factual claim with no support in the canonical JSON fails. Without that distinction, the critic either rubber-stamps everything or vetoes the personality out of the product.

And because this all runs under a 10-minute cap: Phase B reserves 30 seconds for shutdown and output, divides the remaining retry time by the number of input video tasks, and if a caption exhausts its retry budget, the pipeline keeps the last usable failed commentary instead of returning nothing. A degraded caption beats a missing style when every clip must return all four.

The pattern underneath all three

Every guard is the same idea wearing different clothes: shrink the surface area where a model’s output is trusted.

  • Per-frame tagged facts shrink it to single-frame observations.
  • OCR consensus shrinks it to claims two independent runs agree on.
  • The critic loop shrinks it to captions that survive an adversarial check against the canonical record.

Models propose; deterministic code and independent validation dispose. Everything the pipeline logs — parsed frame observations, plus which backend and model handled vision, captioning, validation, and transcription — exists so that when something does go wrong, I can see exactly which trust boundary failed.

Dragon Commentary Studio runs as a single Docker container: FastAPI, faster-whisper for audio, PySceneDetect + FFmpeg for frames, kimi-k2p6 for vision and deepseek-v4-pro for caption generation via Fireworks AI. The four dragon personas — and the pipeline itself — are groundwork for the AI game-review system I’m building into my board game, Caro5: same architecture, with board snapshots instead of video frames.