Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R]
Anthropic’s Jacobian Lens work introduced a way to inspect verbalizable representations inside language models. Follow-up experiments suggested that entropy in this internal “workspace” might help identify confidently incorrect answers. I tested that hypothesis on Qwen3-4B across ~11,400 examples from seven distinct datasets, including TriviaQA, PopQA, NQ-Open, TruthfulQA, HotpotQA, GSM8K, and CommonSenseQA. Three main findings: It can complement output confidence on factual retrieval. On datasets such as P
评论
?
参与讨论