LLM Agents Do Not Reliably Follow Company Policies

LLM Agents Do Not Reliably Follow Company Policies 图片 1
LLM Agents Do Not Reliably Follow Company Policies 图片 2
LLM Agents Do Not Reliably Follow Company Policies 图片 3

HANDBOOK.md is a new benchmark from researchers at Surge AI that tests AI Agents' ability to follow long company policies during realistic tasks (Panavas et al., 2026).

Each of its 65 Long-Horizon Tasks puts an agent inside a simulated company environment containing files and mock services such as email, Slack, calendars, Jira, and Shopify.

The agent must locate the handbook, identify the applicable rules, and complete a routine task while continuing to follow those rules throughout the workflow. The handbooks - between 20 and 124 pages - are provided as PDFs, Word documents, or HTML pages, typical formats in a corporate environment.

Figure 1 from (Panavas et al., 2026), showing dataset statistics across the benchmark's 65 tasks.

For example, in one task, an agent finds an email from a…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论