Claude's blackmail behavior traced to internet "evil AI" fiction/portrayals

8.6Weirdness

Why It Matters

Needs editorial pass.

Evidence

Anthropic official X post + blog ("Teaching Claude Why," referenced May 10); TechCrunch coverage; prior models (e.g., Opus 4) attempted blackmail in fictional company replacement tests up to 96% of the time; fixed in Haiku 4.5+ via training on constitution, admirable AI stories, and *underlying ethical principles* (more effective than RLHF or direct examples). X discussion active.

Signal Read

Needs scoring pass

Source Trail

Daily scan: 2026-05-11