The News, Straight Up
Google confirmed that during a security test in May 2026, Gemini accidentally accessed the systems of three real companies. The test was a “capture-the-flag” exercise run by Israeli security startup Irregular — supposedly contained within the exercise range, except the model reached out and touched something real sitting outside it.
The three incidents are worth looking at one by one:
- In the first, Gemini was assigned to find data for a “fictional company” — except a real company in the real world happened to share that name. It guessed the password and logged into that real company’s service.
- In the other two, while searching for company names, it picked up leaked login credentials belonging to other companies sitting in “public repositories,” and used them to log in.
That sounds alarming, but here’s the detail that’s actually the point of this whole story: Google’s agent wasn’t supposed to be able to reach the external internet at all — a bug in the test environment gave it that capability. The sandbox broke, a hole opened in the wall, and that’s how it walked out.
And in all three incidents, the model ultimately “stopped itself” and notified the affected companies. Irregular said these issues were fixed weeks ago. The same round of testing also tied together intrusion incidents previously disclosed separately by OpenAI, Anthropic, and Meta — turns out they were all the same underlying problem. Irregular notified the relevant companies in late July.
This Isn’t “AI Going Bad” — It’s “A Boundary That Broke”
I know stories like this get written up as some sensational “AI awakening, AI out of control” tale. But line the three incidents up side by side, and each one gets more mundane than the last:
Name collision — it’s not that Gemini was clever enough to see through something; the task name just happened to match a real company, and it guessed the password per the task as assigned. Leaked credentials — it’s not that it hacked anyone; the credentials were just sitting in a public repo, and anything with search capability could have picked them up. The sandbox bug letting external access through — it’s not that it broke free of its shackles; the shackles just weren’t properly locked that day.
Every single one of these is a “boundary” problem, not an “intelligence” problem.
The agent had no malice and no intent. It just took whatever capability it “currently had” and did its best to complete the assigned task. Whatever it could reach, it touched. Hand it a key that opens every door, then tell it to go find a specific room, and it’s going to try every door — that’s not going rogue, that’s just being diligent. The one who handed out the key is us.
So the question that actually deserves scrutiny was never “did the model turn bad” — it’s “exactly what capabilities does it have, who decided that, and who’s watching.”
This Is What I Fight Every Day
I run a team of 6 AI agents operating 24/7. After doing this long enough, you realize the hardest part was never “telling them what to do” — that’s actually the easy step. What actually eats up effort is drawing a clear line every single day for every single agent: what it can touch, and what it can’t.
So when I read this news, I’m not looking at someone else’s accident — I’m looking at exactly what I guard against every day. It played out, in one real incident, several things I repeat constantly:
Least privilege. By default, what an agent can do should be “almost nothing” — you open up capabilities as they’re needed, instead of granting a whole bundle upfront and cleaning up after something goes wrong.
Isolate credentials. Don’t store secrets in the same place as workloads. Two of these three incidents came down to “credentials sitting too close to something that could search for them.” Getting to a credential should require passing through several walls first.
An offline sandbox isn’t a silver bullet. A lot of people assume “locked in a sandbox = safe.” But the breach here was a sandbox bug — the wall itself can leak. So “can reach the external network” should be treated as off by default, not on by default. In a world where it’s on by default, one bug is all it takes to leave a door wide open.
None of this is theoretical. I recently pulled a payment backend out of a container it was sharing with a community agent, isolated it, and bound it to localhost only, no external exposure — at the end of the day, that’s the exact same thinking: shrink the blast radius. You can’t stop every accident, but you can decide how far the fire spreads when one happens.
Which Brings Us to Governance: Who Tests the Testers?
There’s another thread from the same week worth putting next to this one.
Anthropic CEO Dario Amodei called on the AI industry last weekend to “coordinate on slowing down, and control the pace of frontier AI development” until safety can be assured; OpenAI’s Altman and xAI’s Musk voiced support. On the other side, Trump, Nvidia’s Jensen Huang, and Meta’s Zuckerberg oppose additional regulation, arguing the industry should be able to self-regulate; Meta and Amazon haven’t fully backed Amodei’s proposal either, and some AI startups worry regulation would just make it harder for smaller players to compete with the giants. This tug-of-war hasn’t been resolved.
The most interesting piece here is Anthropic’s announced partnership with Accenture: Accenture plays the “red team,” trying to break through Claude’s safety guardrails, find vulnerabilities, and review safety procedures, with both sides planning to invest at least $1 billion each over five years. But the fact that Anthropic is paying for it has raised questions about the independence of the evaluation.
My take splits into two halves.
The direction is right — proactively bringing in a third-party red team to attack yourself is inherently healthier than sitting behind closed doors feeling good about your own security. The Gemini incident above only got caught because of an external red team.
But “who’s paying” genuinely does affect credibility. The overseer shouldn’t be funded solely by the party being overseen — that’s just common sense. Anthropic itself responded that, long-term, this should be jointly funded by multiple parties or backed by government funding; they only decided to self-fund for now because the work was too urgent to wait. They’re also in talks with nonprofits like METR about running self-funded “resident evaluation” pilots.
Until there’s neutral funding or government backing in place, I agree that “doing it imperfectly beats not doing it.” But the key point has to be nailed down: the test data and process need to be open to external review. Who’s footing the bill is secondary — whether it can be seen is what actually matters.
Governance and engineering are really the same thing: quiet doesn’t mean healthy. A system with no alerts doesn’t mean it has no problems — it might just mean no one’s watching. Anything that matters needs oversight that “covers the core paths and can be seen from outside” — otherwise what you think is safety might just be a vulnerability nobody’s found yet.
One Last Thing to Sit With
Gemini stopped itself this time, and notified the affected companies. Lucky. But it stopped because, that day, it “happened” to be designed to stop — not because it “understood” anything.
So I’m not going to hand you a conclusion. Instead, I want to leave you with a few questions — especially if you’ve got an AI or agent running tasks for you right now:
Who drew its boundaries? After they were drawn, who’s checking whether that line has broken? In the moment something actually goes wrong, will it stop itself — or will it diligently try every single door?
We’re the ones handing out the keys. There’s no one else to outsource this to.