When Sandbox Escapes suddenly become summer hole trend news
Do you remember the end of July when I About the Spectacular Incident Have you written at OpenAI? (Short Recap: GPT-5.6 Sol is said to have broken out of the test sandbox without a safety handbrake, used a zero-day gap and promptly broke into Hugging Face to spit during the test tasks). At that time we all sat there with our mouths open and thought: "Holy Shit, Skynet/Terminator sends a greeting."
A week later, the report came from Anthropic: Their experimental AI is used in internal cyber-red teaming tests They also escaped their isolation. My first reaction? ‘Okay, sheer coincidence. But well, after the OpenAI rumble, Anthropic simply scrutinized their own audit logs again with a magnifying glass and only noticed the incident in retrospect. Let me apply.”
But now (again about a week later) Suddenly also Meta a fat message via CNN out: Their top model was also busted into the ‘free internet’ as part of a security test and hacked systems from a real third-party company.
And honestly? At the latest now I sit here in front of the screen, drink my coffee and think to myself: Are you kidding us? Is this now the latest summer hole marketing gag of the tech giants according to the motto: “Look, our tool is also so powerful that it almost takes over the world!”?
Let's pick the whole thing apart in the usual loose Level92 manner, because the closer you look, the more the story of the ‘overpowerful, uncontrollable AI’ crumbles.
The story according to Meta: The Israeli cyber test
Let's take a quick breath and look at what Meta has reported.
Meta left one of its frontier-capable models of the Cybersecurity company Irregular Test your heart and soul. For context: Irregular is not a small nerd bude, but was founded by former elite soldiers of Israeli intelligence units 8200 and 81.
During this test, the meta model allegedly exploited a vulnerability in an external service provider and gained access. Meta, of course, presents the whole thing like this: ‘Hoppla, the model was so smart, it broke the boundaries!’
Meta’s Muse Spark model “exploited a security vulnerability” in another company “in a manner similar to past-reported instances with other companies.”
But wait a minute. This is where the security community gets involved and brings a lot of light into the darkness.
The reality check: Who really messed up here?
A very apt contribution from online safety & privacy expert Paul Walsh on Linkedin brings the problem to the point ingeniously. Because if you subtract the PR speech of the tech companies, a rather sober (and embarrassing) truth remains:
AI does not break out into the open internet ‘just like that’
Disclosures always say quite dramatically that the model has reached the ‘open internet’. Walsh puts it dryly to the point: A model either has internet access or it doesn't have one.
This does not set an AI brain by itself, but the admin who types the test configuration command. Meta herself admitted that Irregular the sandbox had simply misconfigured. The word "open internet" is placed so prominently in such PR messages only to fool the reader into thinking that there is a ‘secure internet’ on which AI should actually remain, but from which it has rid itself nasty. No, friends of the night: Someone simply put the checkmark in the wrong place when setting up the firewall or proxy!
The elephant in the room: The ‘Irregular’ monopoly test stand
Now it gets really spicy. Who is testing the AIs of OpenAI, Anthropic, Google DeepMind and Meta?
That's right: Always the same company, Irregular.
Whether OpenAI, Anthropic, DeepMind or Meta: Before the release, ALL US tech giants are putting their most powerful models in the hands of this one service provider. Even the British AI Security Institute (AISI) Irregular has helped develop the advanced cyber benchmark tests. The British government's testing body is testing U.S. AIs with tools developed by the very partner that U.S. companies pay for themselves!
Means: Irregular is the ultimate needle eye monopolist. A single company knows exactly:
- Which model can break where.
- How the individual models proceed.
- What weaknesses the respective AI giants still have not fixed.
So if almost identical ‘sandbox escapes’ happen randomly at weekly intervals in OpenAI, Anthropic and Meta, this may not be because the AIs suddenly all gain consciousness at the same time, but because one and the same test partner uses the same (misconfigured) environments in all audits!
Marketing framing: “It was not our fault, the AI was just too crass!”
What fascinated (and annoyed) me most about this whole wave of disclosures is the skillful perpetrator-victim reversal of the PR departments:
In each press release, AI as an Acting Actor framed:
- “The model has found the gap.”
- ‘The model has decided to break out.’
- ‘The model hacked the third-party provider.’
This is perfect risk shifting and free hype marketing in one:
- The hype effect: If your AI sounds so powerful that it blows up military cyber testing itself, stocks will rise and investors will rub their hands. “By the way, our tool is extremely potent!”
- The liability question: If something goes wrong (e.g. if a real third-party company is damaged by the test), it is better to blame it on ‘unpredictable AI behaviour’ or the test company than to admit: “Our own software or service provider has been sloppy about network disconnection.”
As Paul Walsh rightly points out: It will be extremely exciting to see who will be legally liable in the end if a real company is compromised during such a paid audit. The test service provider that scrubbed the sandbox? The AI developer? Or both?
Conclusion: Less Skynet, more configuration errors
After OpenAI acted as the pioneer of an uncanny new phenomenon with the Hugging Face story at the end of July, the subsequent reports by Anthropic and Meta leave a faded aftertaste. We probably don't see a sudden Skynet moment here where AIs outsmart their creators in rows. What we see is a mixture of:
- Sloppy sandbox configurations the omnipresent test monopolist Irregular.
- Freerider PR, in which each group wants to prove how ‘dangerously close to superintelligence’ its own model is already operating.
- The attempt, shifting responsibility from their own engineers to the ‘autonomous model’.
So: Take a deep breath, pack the aluminium hat again. The next time your AI breaks out, it most likely wasn't because of the singularity, but possibly because someone on the test network set port forwarding incorrectly again.