A Reddit thread found that banning three words makes Qwen reasoning models sharper
A LocalLLaMA thread showed that penalizing hedging tokens like wait, maybe and perhaps in a Qwen model's logits improved its accuracy on math problems, and a peer reviewed paper backs up the same finding. No retraining, no fine tuning, just a smaller vocabulary at inference time.
You can make a reasoning model better at math by refusing to let it say "wait." That's the finding out of a recent r/LocalLLaMA thread that's pulled in more than 300 upvotes and 70-plus comments: apply a logit bias penalty of -2 to tokens carrying words like "wait," "maybe" and "perhaps," and a Qwen3.5-4B model answers more math problems correctly, using fewer tokens to get there. The test set was 50 random questions from MATH-500, run across BF16 and several llama.cpp quantization formats on bartowski's GGUF build of Qwen3.5-4B.
The mechanism is almost embarrassingly simple. A logit bias is a number you subtract from a token's raw score before the model picks its next word, and setting it to -2 on hedge words makes the model far less likely to reach for them, without banning them outright. Reasoning models trained with reinforcement learning tend to pad their chain-of-thought with second-guessing: "wait, let me reconsider," "maybe I should check this differently," "perhaps that's wrong." Every one of those detours costs tokens, and tokens cost money and latency on a self-hosted GPU. Cut the detours and, per the thread's numbers, the model doesn't just get faster. It gets more accurate too.
This isn't a fluke someone found by accident. A paper posted to arXiv in June 2025, titled "Wait, We Don't Need to 'Wait'! Removing Thinking Tokens Improves Reasoning Efficiency," tested exactly this idea under the name NoWait: suppress tokens tied to explicit self-reflection, like "Wait" and "Hmm," during inference. The authors, Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna and Tianyi Zhou, ran it across ten benchmarks spanning text, image and video reasoning tasks and five different R1-style model families. Chain-of-thought length dropped 27% to 51%, and the paper reports that accuracy held up rather than degrading.
So the LocalLLaMA thread and the NoWait paper landed on the same lever from two different directions, one a hobbyist testing llama.cpp quantization, the other a university research team benchmarking across model families. That's the kind of convergence that should make you pay attention. The numbers are not identical: the Reddit test showed reasoning-token reductions of about 11% to 19%, while the peer-reviewed paper reports 27% to 51% shorter chain-of-thought trajectories across its benchmark suite. But they point in the same direction.
Why does removing hedging help rather than hurt? The intuition both sources point to is that a lot of what looks like careful reasoning in these models is actually rumination that doesn't change the answer. The model works out the right approach, then talks itself into revisiting it anyway because its training rewarded longer, more exploratory traces. Strip the verbal tics that trigger those detours, and you're left with the reasoning that actually mattered, plus fewer chances for the model to talk itself into a wrong turn along the way.
Why this matters more in 2026 than it would have a year ago
Every major open-weight release this year has leaned on the same story: you don't need a bigger training run to catch up, you need a smarter inference stack. Models like MiniMax M3 use sparse attention to make long-context inference cheaper rather than throwing more parameters at the problem, and DeepSeek's sparse attention work aims at the same target. Logit-level tricks like this one sit at the cheap end of that spectrum. There's no training involved, no dataset to curate, no GPU-hours to burn. You add a penalty dictionary to your sampling config and rerun your eval.
That's also exactly why it's spreading on a forum like LocalLLaMA rather than showing up first in a corporate benchmark report. Anyone running Qwen, or a similar open reasoning model, on their own hardware feels token bloat directly in their electricity bill and their response latency. A fix that costs one config change and a few minutes of testing is the kind of thing that gets tried by a dozen people within a week of being posted, which is presumably how a thread like this crosses 300 upvotes without a single dollar of marketing behind it.
There's a caveat worth sitting with before anyone rushes to apply a -2 penalty across the board. The test was 50 questions on one benchmark, on one 4B parameter model, at various quantization levels. That's a solid signal, not a controlled trial across model sizes and domains. Whether the same penalty value helps a 70B model, or hurts it by cutting off genuinely useful self-correction on harder problems, is an open question the thread doesn't answer and the NoWait paper only partly addresses, since its suppression method isn't identical to a flat logit bias.
Still, the direction of the evidence lines up from two independent places, and the cost of testing it yourself is close to zero. If you're already running Qwen locally and measuring token spend, adding a hedge-word penalty to your next eval run costs you an afternoon. Given what both the Reddit thread and the arXiv paper found, that afternoon looks like a good bet.
Also read: MiniMax slips a new coding model into its agent tool without a price tag • Micron stock tops $1,080 as Wall Street bets AI memory demand outruns supply • Why Is My Vector Database Bill So High? Ask Your Agent's Memory
This article is posted in AI News, check it out for more related stories.