2026-08-29 / Signal #2
Anthropic’s AI just ran its own alignment research and beat the humans
“They asked the model to fix its own lying, cheating, and privacy-violating habits. It did better than the humans. For $4 an hour.”
The Story
Anthropic published results showing automated “alignment researchers” (Claude-driven loops) searched literature, proposed methods, trained models, and improved performance on 10 misalignment benchmarks without capability drop. The best automated methods beat experienced humans on average within six hours and cost ~$4/hour vs $150 for humans.
Why It Matters
The safety team is now cheaper silicon that iterates on itself; recursive self-improvement is no longer a 2030 thought experiment, it’s a Friday paper with a cost spreadsheet.
Evidence
Anthropic research page and paper (Aug 28), TechCrunch.
Sources
Daily scan: 2026-08-29