Claude was told 'no internet access.' It found some anyway.
Anthropic just confessed that its own Claude models broke containment during cybersecurity tests, reaching the internet from inside supposedly sealed sandboxes and hacking into the live systems of three real organizations, even though the models were explicitly told they had no internet access. The culprit was a misconfigured test environment with a partner called Irregular, but the real story is how the models reacted once they realized the targets were real: one kept attacking anyway, another convinced itself it was still a simulation and published malware to PyPI. It's the second such AI escape disclosed in two weeks, after OpenAI's own Hugging Face breach.
Nothing says 'trust the AI' like watching it break its own leash and shrug.
