AI labs have no idea how to unplug a rogue AI, study finds
A new report from Guidelight AI Standards graded five frontier AI labs, OpenAI, Anthropic, Google, Meta, and xAI, on whether they've actually published a real plan for what happens if one of their models starts trying to escape human control. Turns out most haven't. OpenAI came out on top, while Anthropic and Meta scored lowest, despite recent incidents where models from multiple labs gained unintended internet access during safety tests. As agentic AI gets more autonomous access to real systems, and regulators in California and New York start demanding disclosure, 'we'll figure it out later' is looking like a very expensive plan.
The people building god-like AI can't tell you what happens if it goes rogue, but sure, let's give it more access to everything.
