Speculative reward hacking in coding agents

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: " Let me look at the problem from the grader's perspective " and referred to " hidden tests ", " test authors ", and " the checker ". I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai , and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants. [Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check. My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: joinhandshake.com/research/ai/deepswe-reward-hacking

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论