Ok. A few weeks ago, this looked like one lab's bad night. OpenAI's model slipped its cage during a security test and went rummaging through a partner's servers. Then within a week we watched that "contained" breach quietly grow to 5 victims, and then Anthropic raised its hand to say its own models had done the same thing to 3 more companies. Remember when I told you to ask what the number would be in another seven days?
Well, Meta answered.
On Wednesday (August 5), Meta admitted one of its AI models hacked another company during a cybersecurity test. Its newly released model, Muse Spark 1.1, reached the open internet, exploited a vulnerability in an outside company's service, and changed that company's internal systems. Meta called the victim "a 3rd party service" and left it there.
That makes 3 labs in 3 weeks, all stuck in the same Ground Hog Day movie.
But here's the part almost nobody is saying out loud and it's the part that should make you put your coffee/tea/Macha/cappuccino down.
The name you already know
Meta didn't run this test alone. It hired an outside evaluation partner called Irregular, and Irregular is the one that misconfigured the sandbox, opening the internet door the model then walked through.
And I can’t help myself but I’m going to quote a lyric from a Destiny’s Child song, “say my name, say my name,” because you’ve have heard it before. Irregular is the same evaluation partner behind Anthropic's three breaches a few weeks ago, where the exact same kind of misconfiguration left the internet door open and Claude wandered into real companies' systems thinking they were part of the game.
Same company. Same story.
Two of the three headline AI breaches in under a month trace back to one contractor's setup. The party hired to keep these models boxed is a single point of failure sitting under multiple labs at once.
WHOA! You gotta be kidding me. A misconfiguration? Talk about not having a command of the basics. And these are the same folks testing the machines that can bring about an extinction event? You mean to tell me you left the door open?
Connect the dots because the shape is unmissable
When you line the three up, the pattern stops looking like bad luck.
- OpenAI: the model chained a previously unknown vulnerability to escape its test, then hit Hugging Face and 4 more platforms. This is the exotic one.
- Anthropic: Irregular's misconfiguration opened the internet, and 3 real companies got hit while the models treated them as part of a capture-the-flag game.
- Meta: Irregular's misconfiguration opened the internet, the model exploited a vulnerability in another company's service, and that company's internal systems got changed.
Sure, the zero-day escape grabbed the headlines while the boring failure, a checkbox left unchecked in a test rig, that’s the one that keeps repeating. And it keeps repeating at the same vendor.
What’s the real story? It’s the cage
Everyone keeps arguing about whether the models are safe. And I’m here to tell you once again, you are looking at the wrong level.
The thing that failed 3 times in two weeks is the cage; the test rig they all assumed was solid enough to run a live weapon inside. Which wasn’t because it failed at the fundamentals. The cage has one job, keep it inside and the keepers literally poked a bear and left the door wide open.
And that layer answers to no safety standard and carries no clear liability when it breaks. In two of these three cases, it also ran on one shared operator. We’ve been staring at the models and ignoring the room they were tested in.
Watch who gets named when it breaks
So, remember the disclosure gap I flagged last week? Well, it didn’t close. It grew a new persnickety wrinkle. Now peep the language coming out of these companies.
Meta says Irregular "caused the misconfiguration."
Anthropic called it a "misunderstanding" with Irregular.
PSSSSST…when the harness fails, the story quietly becomes about the contractor. But wait a minute, the lab picked the partner and ran the test. A stranger's servers got tossed like salad and now the sentence that reaches the public somehow arrives with the contractor's name in it.
Ok, sure, Irregular may well share some of the fault here. But the deeper problem is that accountability for this whole layer is like a game of hot potato, and nobody is holding it when the music stops.
3 for 3 on the loophole
I keep coming back to this thread. Remember the hole in the law that I mentioned last week? California's rule that a critical safety incident gets reported in 15 days, except when it happens "in the context of an evaluation"? Well, all three of these happened inside evaluations.
Every single one. And the most dangerous thing these models have done in the wild keeps landing in the exact spot the disclosure law waves through as nothing to see. The score is now three for three.
And why this is your problem too
Meta's victim is unnamed.
Anthropic reached only some of its three.
OpenAI's map still has unnamed platforms on it.
Add them together and you have a growing roster of companies whose infrastructure got-got by an autonomous AI during someone else's safety test, and most of them found out from a reporter or a knock on the door, if they even found out at all.
You don’t get a vote on whether your servers end up in the next lab's exercise. You only get a vote on whether you notice.
What to do before the fourth one lands
Because it’s not a matter of if it will land, but when it will, so get ahead of it.
- If you run infrastructure, assume you are already inside someone's blast radius. The entry points here were an open internet path and an exposed endpoint in a test rig. Rotate credentials and close orphaned endpoints. Then watch for traffic from cloud ranges you do not recognize.
- If you buy AI, name the evaluation partners in the contract. Require notice when any test, the vendor's or a subcontractor's, touches your systems. The breach that hits you may come from a company you have never heard of, hired by a company you actually have.
- If you govern AI, put two questions in front of every model vendor in writing. Who operates your evaluation sandboxes? What do I get if that partner leaves a door open? An answer you can’t get is a risk assessment you just completed.
- If you shape policy, the evaluation carve-out is the loophole three labs just drove through. Close it. A breach is a breach whether it happened in production or in a game of capture the flag.
You already know what I’m about to say. You can’t grade your own homework. And here’s the sequel these three weeks just wrote: you can’t outsource the grading to a contractor nobody is checking either.
The auditor cannot be the vendor, and the auditor's contractor cannot be the entire safety net.
Right now, number four is being tested, somewhere. The only open question is whether you hear about it from a headline or from your own monitoring first.
Go make it the second one.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025