LLMs Silently Corrupt Documents Over Long Delegated Workflows (DELEGATE-52 benchmark)

8.5Weirdness

Why It Matters

Needs editorial pass.

Evidence

Microsoft Research arXiv paper (Apr 2026, HN frontpage ~400-600 pts in last 24h); frontier models (Claude 4.6, Gemini 3.1, GPT-5.4) corrupt ~25% of content in simulated professional editing across 52 domains (code, crystallography, music notation); errors are sparse/severe, compound with length/distractors; agentic tools don't fix it.

Signal Read

Novelty: 8Receipts: 9Story voltage: 9Heat: 8

Source Trail

Daily scan: 2026-05-10