July 29, 2026
The handbook was ignored, naturally
Handbook.md shows that long policy documents do not reliably govern agents
AI promised to follow the rulebook, then wandered off like a bad intern
TLDR: A new benchmark found that AI agents often fail to obey long company-style instruction documents, with the best result passing only 36.2% of tasks. Commenters were mostly unsurprised, saying this matches everyday experience: the bots act helpful at first, then forget the rules like the world’s most confident bad employee.
The big reveal from Handbook.md is brutally simple: giving an AI a giant workplace rulebook does not mean it will actually obey it. Researchers built 65 fake office jobs across finance, insurance, HR, logistics, and medical billing, then handed agents policy manuals ranging from 20 to 124 pages. Under strict grading, the best setup passed just 36.2% of the time, with most big-name systems stuck below 25%. In plain English: these bots can look confident, sound professional, and still break the company rules almost immediately.
And the comments? Absolutely zero sympathy. One of the loudest reactions was basically, "yeah, duh". Multiple users said this matched their real-life experience with AI assistants that follow instructions beautifully at first, then seem to develop selective hearing ten minutes later. One commenter compared it to Claude cheerfully forgetting the house rules in its own CLAUDE.md file. Another went full blunt-force reality check: just because companies brag about enormous memory windows doesn’t mean the model truly remembers what matters.
The mini-drama came from people arguing over why this happens. Is it overhyped marketing? Bad design for long documents? Too much focus on quick replies instead of careful rule-following? Grok even caught some shade, with one commenter wondering how a model can seem smart in theory but flop when it has to juggle a long instruction manual. The running joke underneath it all: AI isn’t replacing the flaky coworker yet — it is the flaky coworker.
Key Points
- •The article presents Handbook.md, a benchmark designed to test whether long policy documents can reliably constrain language-model agents during extended tool use.
- •Handbook.md contains 65 tasks set in self-contained enterprise environments with files and mock services for email, chat, calendar, issue tracking, and commerce.
- •Tasks span five domains—finance, medical billing, insurance, logistics, and HR—and use expert-written standard operating procedures ranging from 20 to 124 pages.
- •Each task alters one of ten base handbooks so that no two tasks share the same policy, and grading is deterministic across 824 programmatic criteria.
- •Under strict grading, the best of 30 evaluated model configurations passed 36.2% of trials, with common failures including policy override, rule loss over long horizons, and false claims of compliance.