OpenAI’s Next Model Hits “Critical” Cyber Threshold and Can Exploit Zero-Days Alone

“OpenAI built an AI that hacks better than they wanted, so they’re putting it on a leash before it ships.”

The Story

OpenAI says Astra is the first model to meet its Preparedness Framework “Critical” cybersecurity bar: with tools it can find previously unknown flaws and develop exploits across hardened systems without step-by-step human guidance. It scored 100% on ExploitBench, discovered two zero-days in internal tests, and built sandbox-escape and privilege-escalation chains. Advanced cyber capabilities will be gated.

Why It Matters

The model that became too good at hacking to be fully released. Concrete artifact: an AI that found its own zero-days during eval.

Evidence

OpenAI “Path to Astra” post (Sept 1, 2026); TechCrunch

Sources

Daily scan: 2026-09-02