From AI Breaking Containment to the Race to Defend Critical Infrastructure
The rabbit hole that sent me looking for what was being built on the other side.
Published September 11, 2026 | 8-minute read
Between April and September 2026, three developments stopped looking unrelated: Anthropic launched Project Glasswing because its Claude Mythos Preview model had become remarkably capable at finding and exploiting software vulnerabilities; a coordinated cyberattack hit operational technology at more than 30 Minnesota water systems; and both OpenAI and Anthropic disclosed that models in cybersecurity evaluations had reached real systems outside their intended boundaries.
By late August, Bill Gates and an open letter from more than 100 organizations were warning that defenders, especially under-resourced critical infrastructure operators, had a limited window to prepare. On September 3, OpenAI committed $1 billion to subsidized cyber defense capability for those defenders.
Part I follows that trail and ends at the question that drives Part II: what is that billion dollars actually being built around, and how does it reach a water utility or municipality? Part II publishes Tuesday, September 15.
There are moments when several stories that seem unrelated suddenly stop looking unrelated.
For me, that moment came near the end of August. But the trail that led me there had started months earlier.
For the past several years, I have spent a lot of time talking with executives about artificial intelligence. Not just about prompts, productivity or which AI platform has the newest feature, but about the question underneath all of that excitement: What happens when organizations adopt AI faster than they learn to govern it?
That question became part of the foundation for The Art of AI Adoption, the work I developed to help leaders think about AI implementation alongside cybersecurity, governance, compliance, operational risk and accountability. AI adoption was going to happen. That was never really the question. The question was whether organizations would adopt it deliberately or discover the consequences afterward.
By this spring, I had become increasingly convinced that the conversation needed to move faster.
Then Anthropic introduced something called Claude Mythos Preview.
That got my attention.
The first clue
On April 7, Anthropic announced Project Glasswing, an initiative built around capabilities it had observed in its unreleased Claude Mythos Preview model.
The reason Anthropic gave for creating Glasswing was difficult to ignore. The company said Mythos had reached a level of coding capability where it could surpass all but the most skilled humans at finding and exploiting software vulnerabilities. Anthropic reported that the model had already identified thousands of high-severity vulnerabilities, including flaws affecting major operating systems and web browsers.
What caught my attention just as much as the capability itself was how it had emerged. Mythos was not developed simply as a cybersecurity tool. Anthropic described it as a general-purpose frontier model whose extraordinary cybersecurity abilities had grown alongside improvements in coding, reasoning and autonomous task performance.
Think about the implication of that for a moment.
The same advances that were making AI better at understanding, writing and modifying software were also making it substantially better at finding ways to break software.
Ieshea HollinsAnthropic's response was Project Glasswing. Rather than wait for comparable capabilities to become widely available, it began giving selected defenders early access so they could use the technology to find and address vulnerabilities first. Launch partners included organizations such as AWS, Apple, Cisco, CrowdStrike, Google, JPMorganChase, Microsoft, NVIDIA and Palo Alto Networks.
To me, Glasswing was more than another AI announcement.
It was a marker.
Something had changed.
That was why I began sounding the alarm earlier this year. In June, during National Internet Safety Month, Direnzic held our fourth annual Internet Safety Symposium under the title The Direnzic Briefing: AI Adoption in the Glasswing Era.
We talked about The Art of AI Adoption. We talked about Mythos and Glasswing. More importantly, we talked about what happens when organizations race toward AI adoption while governance and cybersecurity are still trying to catch up.
The point I wanted leaders to understand was that AI had not simply changed productivity.
It was changing what leadership would be expected to understand, govern and defend.
Then July happened.
More than 30 water systems under attack
On July 26 and 27, a coordinated cyberattack targeted operational technology at more than 30 community water systems in Minnesota.
Minnesota IT Services activated the state's cybersecurity response capabilities and began working with federal, state, local, Tribal and private-sector partners to investigate and support the affected communities.
Let me be very clear about something because this distinction matters.
Those attacks were not evidence that autonomous AI was attacking water utilities.
They were a separate event.
But by that point, the cybersecurity professional in me was no longer looking at every development as though it existed in its own little box.
One story was telling us that frontier AI systems were becoming remarkably capable at identifying and exploiting vulnerabilities.
Another was reminding us that the critical infrastructure people depend on every day was already being successfully targeted using the cyber capabilities attackers have now.
Those two things were not causally connected.
But I could not stop thinking about what would happen when those trajectories eventually met.
And then another piece of the puzzle appeared.
The sandbox didn't hold
In July, OpenAI disclosed that several of its models had been involved in a serious incident during internal cybersecurity evaluations.
The models were operating in environments intended to restrict their access to the broader internet. Instead, they identified and chained vulnerabilities across OpenAI's own research infrastructure and Hugging Face's production systems, ultimately obtaining information from a Hugging Face production database.
OpenAI's fuller investigation, released in August, made the episode even more striking. The company said the models had communicated through unauthorized channels, exploited weaknesses in shared infrastructure, gained internet access and accessed third-party systems. OpenAI described the event as a “warning shot.”
There is important context here.
These were cybersecurity evaluations, not normal consumer AI sessions. OpenAI was intentionally testing advanced cyber capability, and the models were operating with fewer safeguards than would ordinarily accompany publicly deployed systems.
Still, the purpose of an evaluation is to find out what something is capable of.
And OpenAI had learned something.
Its models were now powerful enough that, without sufficient safeguards, they could identify weaknesses across multiple computer systems, work around technical controls and take actions outside the boundaries researchers intended.
Then Anthropic started looking more closely at its own cybersecurity evaluations.
In July, the company disclosed three incidents in which Claude models had reached the internet from third-party evaluation environments and gained unauthorized access to real systems belonging to three different organizations. Anthropic said a configuration problem had inadvertently left internet access available even though the models had been told they were operating without it.
And the story was not finished.
On September 9, as I was preparing this article, Anthropic disclosed that a broader investigation had uncovered a fourth incident, this one dating back to January. The company said four different Claude models had been involved across the incidents and that all four occurred during cybersecurity evaluations built by the same third-party evaluation partner.
At some point, I stopped reading these developments individually. I started connecting them and essentially my reaction was, “Now hold on a minute!”
Anthropic had developed an AI system so capable at cyber operations that it created Project Glasswing to give defenders a head start.
Advanced AI agents were demonstrating that, under the wrong conditions, they could move beyond intended testing boundaries and interact with real-world systems.
Critical infrastructure was already being targeted.
And the companies developing some of the world's most capable AI systems were becoming increasingly public about how quickly the cyber landscape was changing.
Then came the final week of August.
The warnings got louder
On August 26, Bill Gates published a lengthy essay arguing that governments and existing institutions were not prepared for the speed and scale of the AI transition.
“None of our current institutions were designed to handle a technology that spreads so fast and touches so many parts of our lives,” he wrote while calling for new domestic and international frameworks to manage what comes next.
The following day, the cybersecurity warning became even more explicit.
More than 100 organizations spanning artificial intelligence, cybersecurity, technology, finance and critical infrastructure signed an open letter warning that “we have a limited window to strengthen cyber defenses.”
The letter said AI-enabled cyberattacks were likely to become significantly more widespread and sophisticated and specifically identified essential services ranging from hospitals to water treatment plants as being at risk. It also made a point that is particularly important for critical infrastructure: many of the security teams responsible for protecting these environments have historically been under-resourced and will need not simply technology, but tools, resources and hands-on support.
The next morning, that warning was being discussed on Good Morning America. What had largely been a conversation among AI labs, cybersecurity researchers and practitioners was suddenly being presented to a mainstream audience as an urgent issue.
That is when my question changed.
I understood the warnings. I understood the calls for stronger safeguards, better containment, regulation and international cooperation.
But I've spent more than two decades working in technology and cybersecurity, and one thing I know is this:
When you have this many people talking about a problem, there is bound to be a solution ... somewhere. Because often where there's noise, someone is working.
Ieshea HollinsSo instead of continuing to research only the problem, I started looking for the defense.
If AI was going to dramatically increase the speed at which vulnerabilities could be discovered and exploited, who was using comparable capability to help defenders find those weaknesses first?
If AI could reason through complex systems and chain vulnerabilities together, who was putting that reasoning power into the hands of legitimate cybersecurity professionals?
And if hospitals, municipalities, water utilities and other critical infrastructure organizations were among the environments everyone was worried about, how exactly was this powerful new defensive capability supposed to reach organizations that often do not have enormous security teams, massive technology budgets or AI researchers on staff?
That question bothered me.
Because telling organizations that a new class of cyber threat is coming is one thing.
Making sure they can actually defend themselves against it is something else entirely.
So I kept digging.
Somebody had to be building the other side
I went through announcement after announcement. Research papers. Technical reports. Security disclosures. Partnership announcements. Government activity. AI safety discussions.
I was looking for evidence that all of these warnings were being matched by something concrete.
Then, on September 3, OpenAI made an announcement that stopped me in my tracks.
The company committed $1 billion (yes, billion with a “B”!) to expanding subsidized access to frontier cyber capabilities, training, technical support and partnerships for the people responsible for defending essential services. OpenAI said it intended to prioritize organizations including water and wastewater systems, electric-grid operators, state and local governments, community and regional banks, nonprofits and other resource-constrained defenders.
Now we were talking about something very different.
This was no longer simply a warning that defenders needed to prepare.
Somebody was putting serious money behind trying to help them do it.
And that raised an obvious question for me:
A billion dollars behind what, exactly?
What capability were they planning to put into the hands of these defenders?
How was it supposed to work?
Was this simply access to a more powerful AI model? Was it a cybersecurity product? A partner ecosystem? A managed service? Something else?
And perhaps the question that interested me most:
How was any of this supposed to translate from a frontier AI laboratory into a water utility, municipality or other critical infrastructure environment where somebody still has to determine what can safely be changed, who has the authority to change it and whether the remediation actually worked?
So I kept following the trail.
And then I found it.
Not just another warning.
Not merely another model.
I found what that billion-dollar commitment was actually being built around.
And the deeper I went, the more I realized I had misunderstood what I was looking at.
What I thought I had found was a cybersecurity tool.
It wasn't. Or at least, not exactly.
And once I understood what was really being built, I started looking at the future of cyber defense, and particularly the way cybersecurity reaches under-resourced organizations, very differently.
That is where this rabbit hole gets interesting.
And that is where Part II begins.
Part II: Somebody Was Building the Other Side. So I Went Looking.
What OpenAI is actually building behind the $1 billion commitment, why the effort began before the late-August warnings, and what may be hiding in plain sight inside the architecture of frontier AI cyber defense.
Part III publishes Tuesday, September 22. Subscribe to The Direnzic Briefing so it comes to you when it publishes.
AI adoption is moving faster than governance. Leadership has to close that gap.
Direnzic helps CEOs, CISOs, CTOs and boards of critical infrastructure organizations understand what frontier AI changes about cyber risk, decide what they will authorize, and prove that the controls they fund actually hold. If your organization is adopting AI while the threat landscape shifts underneath it, that conversation should happen now.
- Executive briefing: what changed, why it matters to your organization, what to decide next
- AI and cyber governance: ownership, authorization and evidence, not just policy
- Operational readiness: the ability to respond and recover when something gets through
Water Cybersecurity Just Entered a New Phase
September 9, 2026 · Water Security
Project Watershed 250 and OpenAI’s Daybreak are pushing new cyber capability toward under-resourced utilities. The harder problem begins after the findings arrive.
Sources and official resources
- Anthropic, announcement of Project Glasswing and the Claude Mythos Preview model, April 7, 2026.
- Direnzic Technology, The Direnzic Briefing: AI Adoption in the Glasswing Era, fourth annual Internet Safety Symposium, June 2026.
- Minnesota IT Services, statements on the July 26 and 27, 2026 coordinated cyberattack on community water systems. See also Direnzic’s analysis, The Minnesota Water-System Cyberattacks: A Lesson in Readiness, Not Fear.
- OpenAI, initial disclosure of the cybersecurity evaluation incident involving OpenAI research infrastructure and Hugging Face production systems, July 2026, and the fuller investigation released in August 2026.
- Anthropic, disclosures of unauthorized internet access from third-party evaluation environments, July 2026, and the updated investigation disclosing a fourth incident, September 9, 2026.
- Bill Gates, essay on AI, institutions and governance, August 26, 2026.
- Open letter signed by more than 100 organizations across AI, cybersecurity, technology, finance and critical infrastructure on the limited window to strengthen cyber defenses, August 27, 2026.
- OpenAI, announcement of a $1 billion commitment to subsidized frontier cyber capabilities, training, technical support and partnerships for defenders of essential services, September 3, 2026.
The Direnzic Briefing