The Bit Explainers

The Model Already Knew the Answer — It Just Couldn't Stop Talking Long Enough to Say It

A look at a provocative new finding: two random vectors, no training, and a small open-weight model's accuracy nearly doubles. Here's what the claim is, how it holds up, and why "if true" is doing a lot of work in that sentence.

Good Omens Studio15 min readMachine Learning

The claim in one paragraph

A small team working under the name Irys ran Qwen3-4B, a compact open-weight model, on a set of multi-step arithmetic problems. Greedy decoding — the standard way these models generate text — got the right answer only 32% of the time. But when the researchers looked at whether the correct answer showed up anywhere in the model’s internal computation, that number jumped to roughly 80%. The model, in other words, appeared to “know” the answer far more often than it “said” the answer. Their fix: prepend two vectors of random noise to the model’s embedding space before it starts generating — no fine-tuning, no retraining, one line of code. Accuracy on a single pass rose to 51.6%. Run ten noisy passes and take the most common answer (plurality vote), and accuracy hit 72%. At least one of those ten passes contained the correct answer 100% of the time.

If that holds up, it’s a genuinely interesting result. The rest of this piece is about what “holds up” should mean here.

What Qwen3-4B computes vs. what it actually says 80% gets the right answer somewhere inside its computation 32% actually states the answer under ordinary greedy decoding the missing 48 points

The gap between these two bars is the entire paper. Whatever’s wrong with this model, it isn’t a lack of arithmetic — it’s a failure to land the plane.

The mechanism they’re proposing

The explanation offered is a version of a known LLM failure mode: autoregressive lock-in. Language models generate one token at a time, and each token constrains what can plausibly come next. The claim is that the first ~20 tokens disproportionately determine the shape of the entire response — including whether the model settles into rigid, heavily-formatted “answer templates” (numbered steps, LaTeX, headers) that it then dutifully fills in, whether or not the underlying computation was ever finished. The paper’s framing: the model commits to performing the aesthetic of solving the problem rather than actually solving it, then runs out of token budget formatting an answer it never computed.

Random noise injected at the very start, the argument goes, knocks the trajectory off that templated rail and into what they call “exploratory computation mode” — messier, more like scratch arithmetic (“ok so 45+23 thats 68, then times 17…”) than a structured proof. That messier path is shorter and, empirically, more often correct.

Supporting this, they report that in 500 sampled responses, every single response that reached a natural stopping point (rather than getting cut off at the token limit) was correct — a suspiciously clean result worth treating with real skepticism given the small sample, but a striking one if it replicates.

Where the answer gets lost — and how noise reroutes it Prompt begins First ~20 tokens (decisive) greedy decoding locks into a rigid formatting template ✕ runs out of tokens formatting, never states the answer + 2 random vectors nudges it into exploratory computation ✓ finishes naturally, states the answer

Why “direction doesn’t matter” is the strangest part of the story

The most counterintuitive detail: it doesn’t seem to matter what the random noise actually is. Carefully optimized perturbation vectors and pure random noise produced statistically indistinguishable results (Mann-Whitney p = 1.0, reportedly repeated across two different geometric setups). The authors’ interpretation is that the noise isn’t adding information — it’s adding energy. Their working theory is stochastic resonance, a real phenomenon from physics and signal processing where adding noise to a system that has a signal sitting just below some activation threshold can push that signal over the threshold and make it detectable. Mapped onto a transformer, the idea is that the correct reasoning path already exists in the model’s weights but sits just under the threshold that greedy, deterministic decoding would ever select — and noise gives it the nudge it needs.

They point to three other independent papers (on random soft prompts, meaningless filler tokens, and repeated “…” tokens) that arrive at similar conclusions through different routes, which is a reasonable point in favor of there being something real here, even if the underlying mechanism is still guesswork.

The dose matters, even if the direction doesn’t

One detail that’s easy to miss: more noise is not better noise. The team swept the prefix length from 0 to 8 tokens, and accuracy peaks sharply at 2 and then falls as more noise is added.

Accuracy vs. number of random prefix tokens 50% 32.0% 42.7% 51.6% 44.0% 44.4% 0 tokens 1 token 2 tokens 3 tokens 8 tokens peak — the "sweet spot"

The proposed reason for the drop-off is almost funny: too much noise gives the model too many directions to explore, so at 8 tokens it starts branching into several partial strategies and blows its token budget before finishing any of them. The fix for a broken formatting habit turns out to have its own version of the same failure mode if you overdo it.

Where this actually seemed to help

The arithmetic benchmark was the controlled experiment, but the more interesting anecdotes are elsewhere:

  • On a Redis debugging task, the baseline model produced 14 incoherent words and stopped. Every noise-perturbed run instead produced a complete, 600+ word diagnostic writeup.
  • On 12 legal-reasoning tasks (contract review, GDPR classification, misclassification analysis, and similar), the best of five noisy runs beat the plain baseline on 11 of 12, with some fairly large jumps in judged quality.
  • Smaller models with perturbation reportedly matched or beat larger models in the same family more often than not — the headline economic claim, since a 4B model is dramatically cheaper to run than a 14B or 32B one.

Not all voting methods are equal — and one of them backfires

Running ten noisy passes only helps if you combine them sensibly. The paper’s comparison of voting strategies is one of its more useful practical findings:

Ten noisy passes, four ways to combine them 32% baseline (no noise) 40% majority vote (>50% agree — backfires) 72% plurality vote (most common answer) 100% oracle (best of 10 — theoretical ceiling)

Majority voting — requiring more than half the seeds to agree — actually performs worse than doing nothing, because when most individual seeds are still more likely to be wrong than right, demanding a true majority just filters out cases where the answer never reached consensus. Plurality voting (simply taking whichever answer showed up most) sidesteps this, because correct answers tend to converge on the same value while wrong answers scatter randomly across different (truncated, garbled) outputs. The oracle bar is the important gut-check here: it’s not something you get automatically — it says “the right answer was in there on every task,” not “we have a reliable way to always pick it out.” That gap between the gold (oracle) and green (plurality) bars is, honestly, the whole unsolved part of this research.

The economics argument

The piece leans hard into cost. Using rough cloud-inference math, ten perturbed passes on a small model cost around $0.008–0.009 per query, versus an estimated $0.45 for a frontier “thinking” model burning tens of thousands of reasoning tokens on the same query — roughly a 56x difference.

Cost per query — bars capped for readability (~50x gap) $0.009 10 perturbed passes on the 4B model $0.45 one frontier call (thinking-token billing)

At real production volumes (tens of thousands of queries a day), that’s the difference between a few thousand and several hundred thousand dollars a month. The broader argument is that the industry’s obsession with scaling up model size to fix quality problems may be missing a cheaper lever: fixing how a model is asked to compute, not making it bigger.

Reasons for caution — and the authors are fairly upfront about most of these

This is where “if true” earns its place in the framing:

  • Tiny sample sizes. 25 arithmetic problems, 5 planning tasks, 12 legal tasks. The statistics reported (McNemar’s test, p < 0.001) tell you the effect is probably not pure noise, but they don’t tell you the effect size will hold at scale — the authors say as much themselves, noting an earlier 3-seed scout run overestimated the effect before a 10-seed rerun brought it back down.
  • It doesn’t generalize to every model or setting. A near-ceiling model (already scoring 76%) got worse with perturbation. Quantization matters a lot — 4-bit weights reportedly killed the effect almost entirely, while 8-bit preserved it, and that comparison rests on one model.
  • The scoring/voting method is unresolved. Oracle accuracy (best of 10 seeds) hit 100%, but mean accuracy was much lower, and the “scorer” meant to automatically pick the best seed was reportedly broken on most of the legal tasks. The gap between what’s possible and what’s reliably retrievable is still wide.
  • A prior claim in the same research line was a bug. The authors disclose that an earlier verbosity finding traced back to a code defect — a genuinely good transparency signal, but also a reminder that this is fast-moving, unreviewed, self-published research.
  • It’s not peer-reviewed. This is a Substack post from an independent newsletter, not a paper that’s been through external review. The GitHub repo is open, which is the right move if the claims are meant to be checked, but “open-source and statistically significant on a small sample” is a different bar than “replicated.”

The bigger-picture pitch, and the part worth separating from the data

The article’s closing argument is really two claims stacked together: a narrow empirical one (random embedding noise measurably helps a stuck small model finish reasoning it already started) and a much broader economic one (this is an overlooked “second revolution” in AI, on par with cheaper compute unlocking whole new industries the way smartphones did). The first claim is testable and, on the evidence presented, plausible enough to be worth independent replication. The second is more of a thesis about where the industry’s incentives point, and it’s worth reading as opinion and framing rather than as something the arithmetic benchmark actually established.

Reading between the numbers: what this might actually mean

Setting aside whether the exact percentages replicate, a few implications are worth sitting with regardless of how the follow-up experiments turn out.

1. Small-model benchmarks may be measuring the wrong thing. If a model “knows” the answer 80% of the time but “says” it 32% of the time, then a leaderboard score isn’t really measuring the model’s intelligence — it’s measuring the model’s default decoding path. That’s a meaningfully different claim. A lot of “small open-weight models just aren’t smart enough yet” conclusions may actually be “small open-weight models weren’t asked in a way that let them finish.” Worth remembering the next time a benchmark table is used to justify buying a bigger, more expensive model.

2. This is a new data point in an old argument about what alignment costs you. Instruction-tuning and RLHF teach models to reach for reassuring scaffolding — numbered steps, headers, clean notation — because that’s what raters reward. This paper is indirect evidence that the scaffolding itself can crowd out the computation, especially in smaller models with less capacity to do both at once. An open question the authors don’t answer: does a raw, non-instruction-tuned base model show the same lock-in, or is this specifically a side effect of post-training? That would tell you whether the fix belongs at inference time (like this one) or earlier, in how models are trained to respond.

3. It reframes “test-time compute” around where you spend it, not just how much. Frontier labs’ answer to a hard query is more thinking tokens — compute spent linearly, at the output stage. This result suggests a second, nearly-free lever: intervene at the very first tokens, before the trajectory commits to anything. If that generalizes, “spend more” and “start differently” turn out to be two separate knobs, and the second one is currently almost unexplored outside a handful of papers.

4. The real bottleneck isn’t generation — it’s picking the right answer out of what you generated. Oracle accuracy hit 100% on the arithmetic benchmark; the actual usable accuracy without a working scorer was much lower. That’s the same shape as a pattern showing up across a lot of current AI research: generating many candidate solutions is becoming cheap, but verifying which candidate is correct remains the hard, unsolved part. This paper doesn’t solve that problem — its own scorer was reportedly broken on most legal tasks — but it does sharpen exactly where the next real gain has to come from.

5. A caution worth naming, since the paper doesn’t address it. All of this was tested on arithmetic, debugging, and legal-reasoning tasks — domains where “finish the formatting and state the number” is exactly the failure mode you want to break. Nobody has yet tested whether the same prefix-perturbation trick interacts with safety-relevant behavior (refusals, guardrail phrasing) the same way it interacts with a math answer template. A technique that reliably knocks a model off its default template is worth understanding fully before assuming it only ever helps.

Bottom line

If the effect replicates outside this one team’s setup — across more models, more tasks, and larger sample sizes — it would be a genuinely useful, cheap technique: a way to unstick small models on tasks where they have the knowledge but choke on execution, without any retraining. That’s a real, practical, and fairly modest claim. The much bigger claim the piece makes — that this represents an overlooked economic shift in how AI capability gets delivered — is a lot to hang on 25 arithmetic problems and a dozen legal tasks. Worth watching, not yet worth banking on.