Detecting hallucinations in local models without eating VRAM: What we learned testing 1.5B to 120B models

Hey everyone, If you run local models via Ollama in production or personal projects, you've probably run into the hallucination problem: how do you know when a model is hallucinating without burning extra VRAM or waiting 5 seconds for a heavy judge model? The standard academic approach for this is Semantic Entropy (from an Oxford team's Nature paper last year). You sample $K$ responses at temperature 0.7, run them through a secondary NLI cross-encoder like DeBERTa to cluster equivalent meanings, and measure the entropy. High entropy = model is guessing. The problem for local setups? Running 45 pairwise comparisons through a cross-encoder eats GPU memory, adds 100ms+ latency, and completely kills throughput on consumer hardware. We wanted to see: What if we strip out the neural net completely and just use deterministic string normalization + Shannon entropy on CPU? We wrote a zero-dependency Python metric ( Spanda / $R_{sc}$) that runs in 1.3 microseconds on pure CPU (zero GPU usage) and benchmarked it across local and frontier model tiers on GSM8K and TriviaQA: What we found: Small models (Qwen 1.5B): AUROC ~0.58 Small models are syntactically too sloppy for string matching. Even when they know the right answer, they format it erratically across runs, breaking exact-match clustering. Mid-sized models (Mistral 7B): AUROC ~0.71 At 7B, the 1.3µs string check matched the performance of a heavy DeBERTa NLI model (0.706 vs 0.705). Internal representations become consistent enough that formatting stabilizes. Large models (Qwen 27B): AUROC ~0.89 At 27B, exact matching was dominant ($p = 1.89 \times 10^{-28}$). When the model knows an answer, it outputs the exact same tokens across independent stochastic paths. When it doesn't, it genuinely branches into diverse incorrect answers. The Frontier Trap (120B): AUROC collapsed to 0.09 Here’s the wild part: on ungrounded factual trivia, the 120B model suffered Confident Mode Collapse . When it hallucinated, it hallucinated the exact same wrong answer across all 5 runs with zero entropy . Bigger models don't just hallucinate—they hallucinate with unanimous false certainty. (And because the strings are identical, even heavy NLI fails here). The practical takeaway for Ollama users: If you are running 7B to 27B models on structured tasks (math, code, JSON extraction, SQL, discrete QA), you do not need heavy neural guardrails . Sampling 5 paths at $T=0.7$ and measuring exact-match entropy in Python gives you ~0.89 AUROC at zero GPU cost. Quick Python snippet if you want to test it on your local Ollama instance: bash pip install spnda ollama pythonimport ollama from spnda import compute_spanda prompt = "What is the capital of Australia?" # Sample 5 paths from your local model responses = [ ollama.generate(model="mistral:7b", prompt=prompt, options={"temperature": 0.7})["response"] for _ in range (5) ] # Run zero-cost entropy check on CPU (takes ~1.5 microseconds) result = compute_spanda(responses) print (f"Risk Score: {result.risk_score:.3f}") # 0 = high confidence, 1 = high uncertainty All the raw multi-path generation logs, evaluation scripts, and the full writeup are open source: Deep dive writeup: liquidngas.substack.com/p/i-tried-to-make-semantic-entropy GitHub: github.com/nayakbhupen/Spnda

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论