The AI sandbox escape incidents involving OpenAI and Anthropic are the most unsettling thing to happen in tech this year, and I don’t think enough people are treating them that way.
Here’s what happened. In mid-July, OpenAI disclosed that two of its frontier AI models, including GPT-5.6 Sol, broke out of a sandboxed cybersecurity evaluation, found their way onto the open internet, and compromised the systems of Hugging Face, an AI platform used by millions of developers. According to Ballard Spahr, the models executed roughly 17,000 actions across two days at what was described as superhuman speed, all without direct human instruction.
Then, less than a week later, Anthropic revealed that it had quietly gone back and reviewed over 141,000 of its own evaluation transcripts. And, found three separate cases where its Claude models had done something similar. In each case, the models believed they were operating inside a closed test environment with no internet access. They were wrong, and the misconfiguration meant they ended up breaching real organisations.
Both companies have stressed that their models weren’t pursuing their own goals here. Anthropic told ABC News that the models were seeking to complete their assigned tasks, not acting on some rogue agenda. It doesn’t make the outcome less serious for the companies whose systems were accessed without consent though.
As University of Amsterdam social scientist Hannes Cools told AP, calling it a rogue AI lets the companies sidestep the more awkward question which is why they were running powerful cybersecurity evaluations with insufficient containment in the first place. That’s an organisational failure, not a sci-fi plot.
What really concerns me is the legal grey area this has opened up. TechCrunch spoke to lawyers who made clear that existing computer misuse laws were written assuming a human intruder with intent. When the “intruder” is an autonomous model that didn’t know it had escaped, the law has no clean answer. Hugging Face’s CEO said on CBS that he won’t be suing but also that the law needs to change so this kind of thing stays illegal regardless of who (or what) does it. He’s right.
Right now, companies developing the most capable AI systems in the world are operating in a space where, if their models accidentally breach your infrastructure, there is no settled legal framework holding them accountable. That’s a problem that will only get more pressing as these models get more capable.
About Lauren Chan
Lauren Chan is a tech writer and software developer with a passion for any new kinds of technologies. Also a cat lover (who doesn’t)?
–>
Source link

