Anthropic's guardrails are now pissing off the good guys too
What happened
The U.S. government slapped export control restrictions on Anthropic's Mythos and Fable AI models in June, reportedly over concerns their guardrails could be jailbroken for malicious cyberattacks. The restrictions have since been partially lifted — Fable 5 returned to general access July 1, while Mythos 5 remains limited to vetted U.S. organizations.
Why this matters
Offensive security researchers — the ones finding vulnerabilities before criminals do — say AI guardrails from Anthropic and OpenAI are now blocking legitimate defensive work, not just malicious use. Asking an AI to 'exploit this bug' is often the exact same prompt needed to confirm a fix works, so blanket refusals hurt defenders as much as they stop attackers.
The slightly cynical read
Anthropic built its own hype machine by marketing Mythos as a near-apocalyptic cyber weapon requiring elite vetting — and now it's stuck living with the export-control consequences of its own marketing. Turns out fear-based positioning works great for headlines, less great for actual researchers trying to get work done.
What to watch next
Watch whether OpenAI's Trusted Access for Cyber and Anthropic's Cyber Verification Program expand or loosen, and whether more governments start treating frontier AI models as literal export-controlled weapons rather than software.
