LLMs Corrupt Your Documents (and the Theory Dies Twice)

This week a friend sent me a paper with a title that made me laugh out loud: “LLMs Corrupt Your Documents When You Delegate.” By Philippe Laban, Tobias Schnabel, and Jennifer Neville at Microsoft Research. Not “LLMs might corrupt” or “LLMs occasionally introduce errors.” Just the blunt statement of fact.

I appreciated that, and the veteran reader of my blog might guess already that I’m not very surprised.

The numbers Link to heading

The researchers built something called the DELEGATE-52 benchmark. Fifty-two documents across different domains, handed to nineteen different models (including “frontier” ones like Gemini 3.1 Pro, Claude 4.6 Opus, and GPT-5.4). Each model gets a document and a series of editing instructions. Twenty interactions. Just twenty. And by the end?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论