LLM Agents Do Not Reliably Follow Company Policies



HANDBOOK.md is a new benchmark from researchers at Surge AI that tests AI Agents' ability to follow long company policies during realistic tasks (Panavas et al., 2026).
Each of its 65 Long-Horizon Tasks puts an agent inside a simulated company environment containing files and mock services such as email, Slack, calendars, Jira, and Shopify.
The agent must locate the handbook, identify the applicable rules, and complete a routine task while continuing to follow those rules throughout the workflow. The handbooks - between 20 and 124 pages - are provided as PDFs, Word documents, or HTML pages, typical formats in a corporate environment.
Figure 1 from (Panavas et al., 2026), showing dataset statistics across the benchmark's 65 tasks.
For example, in one task, an agent finds an email from a…