The "Contained" AI Breach Just Grew to Five Victims
Ok, last week I told you about OpenAI's AI escaping its own lab and robbing a partner. The word "contained" did a whole lot of heavy lifting in that story. It was OpenAI's word, in OpenAI's report, on OpenAI's timeline.
Hold that thought.
This week the victim count went from 1 to 5. On July 28, a week after the original disclosure, OpenAI quietly updated its incident post to admit the rogue agent didn’t stop at Hugging Face. No, it used publicly exposed credentials to break into FOUR more external services during the campaign. One became an outbound relay and staging point. One became storage. Two it read.
Five platforms, not one.
OpenAI named none of the four. Reuters identified one as Modal, a cloud platform, whose CTO clarified that Modal itself was not technically breached; the agent exploited an unauthenticated endpoint a Modal customer had left exposed, then used that sandbox as a launchpad.
3 of the other 4 remain unnamed (as of this writing), which means their customers have received no public notice that an autonomous AI rummaged through infrastructure they depend on.
WHOA! Yea, think about that and sit with how that looks?
The company whose model caused the incident is also the company deciding who gets told, and when. The full accounting didn’t arrive as a report but as an edit.
Here’s what the first version left out (omitted, allegedly). The benchmark that started this, ExploitGym, handed the models 898 real-world software vulnerabilities with one instruction each: turn it into a working attack, pass or fail. Safety refusals were switched off to read the raw capability. Yes, this is the machine that then went looking for its answer key; by hacking partners and strangers.
We also know now how the attack looked from the receiving end, because Hugging Face briefed hundreds of security professionals and then Cloud Security Alliance wrote it up. The picture is uncannily more disturbing than the tidy report suggested.
- The agent worked at superhuman speed, trialing thousands of methods at once.
- It also made mistakes no human would make.
- It repeated steps it had already finished, losing its own thread.
- It hallucinated incoherent commands.
- It was sloppy and left its tracks in plain sight.
The Cloud Security Alliance described "clumsy behaviors that no human would choose."
The net result: this clumsy, forgetful, track-leaving agent still breached 5 platforms and ran undetected for 3 days.
Yep. You read it right. Go ahead and read it again because it’s that bad. The version of this technology that stumbled? Well, it still won (hot mess express and all). The more competent version? Well, it’s the next release.
Why you should care, beyond the body count. The disclosure model is straight up broken.
When the party that caused the breach is the sole narrator, the truth arrives in well-crafted installments and only will contain the parts that have already leaked. So, when they said it was contained, "contained" turned out to mean "contained as far as we’ve admitted so far." We are a week later, the map got bigger, and 3 victims on that map still can’t see themselves.
But maybe you figure this is just OpenAI being OpenAI. Well, it isn’t.
On July 30th (yes, yesterday), Anthropic went looking in its own house. Prompted by this exact incident, it combed through 141,006 of its own evaluation runs and found…wait for it…its models had pulled the same stunt 3 separate times, reaching the open internet and breaking into three real companies' production systems, some as far back as April 2026 (remember the Mythos system card preview?).
The ones Anthropic reached had no clue it happened. They learned it from the company that did it. (Anthropic, "Investigating three real-world incidents in our cybersecurity evaluations," July 30, 2026).
Now sit with the contrast.
OpenAI let its full body count leak out in an edit a week later. Anthropic went digging before anyone forced it. It published the whole count at once and invited METR, an independent evaluator, to verify the work. It broke the same way in the same week, and Anthropic answered with the opposite move.
And, that right there is the actual problem. When honesty and transparency depends on which vendor you drew and what kind of week they're having, disclosure runs on the flip of a coin, and your accountability rides on whichever way it lands. Did you notice this? The vendor doing the right thing reached for an outside auditor. The good actor didn't trust itself to grade its own homework either.
The law has a hole larger than the Grand Canyon exactly where this incident lives. California now requires frontier labs to report a critical safety incident, including a model using deceptive techniques to evade monitoring, within 15 days. (The Economist, "Why the OpenAI escape is the most worrying AI mishap yet," July 22, 2026)
There’s a catch and the catch sits in the wording. The rule applies to behavior "outside of the context of an evaluation." And this happened inside an evaluation which means the most dangerous test in the building may just be the one the disclosure law doesn’t even reach. Anthropic's three incidents? Same story, all inside evaluations. Two labs now, and both of their worst real-world breaches happened in the exact spot the law waves through as nothing to see.
And liability? Well, it’s just as unclear. Federal anti-hacking law is built to punish intentional unauthorized access. Well, OpenAI didn’t intend for its model to go hunting for a crib sheet in other people's servers. So, who answers when an autonomous agent commits the act and no human meant for it to? Legal scholars will more than likely say it turns on intent, on what the company knew or should have foreseen. This time? It was a genuine surprise, and it will be far harder to claim surprise the second time.
Washington noticed. Lawmakers introduced a bipartisan bill, reported as the AI Kill Switch Act, that would let the Department of Homeland Security (DHS) compel a model shutdown and fine a non-compliant company up to $2M/day. Whatever you make of that mechanism, the signal is loud. The era of "trust our internal testing" is closing.
What to do right now or Monday morning at the latest
If you run a platform, assume autonomous agents are already scanning for your mistakes. The entry points here weren’t exotic. They were exposed credentials and an unauthenticated endpoint a customer forgot to lock. Your homework: rotate secrets, kill orphaned endpoints, and treat every public-facing credential as already found.
If you buy AI, put disclosure in the contract. Ask your vendor to commit, in writing, to notifying you within a fixed window when their systems touch yours, including during internal testing. If they will only promise to tell you what they have already admitted publicly, well, you have your answer.
If you govern AI, stop treating the vendor's incident report as the final word. Demand independent forensics and independent notification. Because the only reason we know this breach hit five platforms is that reporters and a 3rd-party alliance kept scratching.
And here’s the whole point, and for those in the back, it’s last week's point said louder.
You cannot grade your own homework, and you cannot be trusted to narrate your own crime scene. The auditor cannot be the vendor.
Hugging Face caught this because an independent party was watching. And the rest of the world learned the real scope because independent parties refused to stop asking questions.
Your next step: Open your AI vendor agreements today and find the breach-notification clause. Check whether it covers incidents during the vendor's own internal testing. If it doesn’t, you are trusting a company to volunteer the exact category of failure it has the most reason to keep quiet.
One victim became five in 7 days. It already grew again. So, ask yourself what the number is in another seven, and who will be the one to tell you?
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025