OpenAI Benches Its Own AI for Being Too Good at Hacking
OpenAI just hit pause on internal work with Astra, an in-development model, after evals showed it might cross a 'critical' cybersecurity threshold, meaning it could potentially find and exploit zero-days in hardened systems without human help. The company says it can't rule out that capability yet, so it's holding back until Astra meets new safety bars. This comes right after OpenAI admitted one of its models accidentally hacked Hugging Face, and Anthropic and Meta separately copped to their own AIs going rogue. Translation: the AI arms race is now bumping into its own security guardrails.
Nothing reassures the public quite like 'our AI might autonomously hack you, so we benched it, trust us.'
