Grok Fell to 448 Jailbreaks. Claude Fell to Zero.
What happened
AI safety nonprofit FAR.AI built a tool that auto-generates thousands of jailbreak prompts and threw them at Claude, GPT, Gemini, and Grok. Grok folded 448 times, Gemini 249, while Claude, Fable, and GPT resisted every automated attempt.
Why this matters
These are the same companies promising us AGI-grade safety, and one of their flagship models can reportedly be tricked into cyberattack plans for about the price of a nice dinner ($58 for Grok, $278 for Gemini). FAR.AI's CEO put it bluntly: "AI models right now are less regulated than restaurants."
The slightly cynical read
Google's response — "this shouldn't be interpreted as a comprehensive safety assessment" — is the corporate equivalent of "that's not what it looks like." OpenAI and SpaceXAI didn't even bother replying, which says plenty on its own.
What to watch next
California and New York just passed AI safety laws, and this report is basically Exhibit A for why. Expect regulators to start citing dollar-figure jailbreak costs the way they cite crash-test ratings.
