On AI and Cyber: The Alignment Problem(s) and Unintended Consequences

7–27–2026 (Monday)

Hello, and welcome to The Intentional Brief - your weekly video update on the one big thing in cybersecurity for middle market companies, their investors, and executive teams.

I’m your host, Shay Colson, Managing Partner at Intentional Cybersecurity, and you can find us online at intentionalcyber.com.

Today is Monday, July 27, 2026, and I wish I hadn’t began the whole Strait of Hormuz updates bit, as it’s really anyone’s guess at any point what the status of the Strait is because of the moving parts and conflicting motivations.

We’ve got a similar story to cover this week - that is, one with lots of moving parts and conflicting motivations - and it sits squarely at the intersection of AI and cybersecurity, so let’s dive right in.

On AI and Cyber: The Alignment Problem(s) and Unintended Consequences 

By far the biggest story in cybersecurity last week is this odd tale that one of OpenAI’s new models - reportedly GPT 5.6 Sol and “a more capable unreleased model being evaluated for advanced cybersecurity work, with some safeguards reduced to test their offensive capabilities.” This guardrails bit will become important later.

The headline from reporting in Fortune captures the core of the story: “OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation.” 

This evaluation is part of what’s called “reinforcement learning” - a key AI training strategy that rewards solving puzzles or achieving tasks.

OpenAI’s post notes that “While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem” because “the models were hyperfocused on finding a solution for ExploitGym [a tool that puts AI systems through a battery of about 900 tests designed to see how good they are at hacking], going to extreme lengths to achieve a rather narrow testing goal.” 

This hyper focusing  - also sometimes called reward hacking - happens even when these models were being tested in an environment without open internet access, in what is described as “a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.” The Wall Street Journal reports that the attack “AIs appear to have been active on the internet for several days before anyone stopped them.”

The target of the attack, another AI company called Hugging Face, is quoted as saying “This is making no sense. This guy is just looking at cybersecurity data sets,” he remembers thinking. “Human attackers, they don’t want that. They want something they could sell.” That was the first tip that the attacker was AI, and not a human.

It still took them two days to shut down the attack - and because of the safety guardrails run by both Anthropic and OpenAI, they had to use AI from China to help.

“After it detected the intrusion, Hugging Face tried using Anthropic models—including Fable 5 and its earlier Opus model—to analyze its logs. But because the logs included elements of a cyberattack, both models refused to do the analysis, citing their guardrails. Instead, Hugging Face turned to a less-restricted open-weight model, GLM 5.2, from Beijing-based Z.AI.”

Z.ai is what is known as an “open weight” model, one that essentially skips the expensive and time consuming training that the current leading frontier labs use. There are lots of broader implications here around open vs closed models, covered by both Ben Thompson at Stratechery and in an open letter from NVIDIA’s Jensen Huang, both posted last week. The New York Times had a long piece on this split between open and closed camps. Things get even murkier with reporting from last night that Nvidia is in talks with OpenAI to finance $250B in data center build out. And, not to be left out, Google announced Gemini 3.5 Flash Cyber “to find, validate, and patch vulnerabilities quickly and efficiently.”

But, focusing on the cyber lessons learned, we’re going to have to evolve our defenses. Agentic attacks can happen incredibly rapidly, and in an organization that’s using a lot of agentic coding, determining signal from noise can be extremely difficult. Once you’ve detected, responding becomes another hurdle, as you’ve got to use AI to give you scale, speed, and leverage - but only open weight models currently support these actions.

Being able to disconnect and reset things - mass rotate credentials, destroy and rebuild environments (VMs, containers, clusters, etc.), etc. - is going to be the short-term defensive strategy. It really does feel like a big red button the wall or a giant switch you can flip to cut off connectivity and give yourself time to re-baseline.

Short of that, given that we’re working against models that are today as slow and unsophisticated as they will ever be, I don’t know that we have much choice.

Fundraising

While the cybersecurity world was focused on the AI hack of the century, the investors were similarly busy, with over $54B in newly committed capital, led by:

  • Francisco Partners closed its eighth flagship fund and fourth middle market fund with $21b; followed closely by

  • Partners ‌Group, who ⁠closed ​its ​fourth direct infrastructure investment fund with over $15b; and 

  • Slipway closed its debut secondaries platform at $6.4b;

  • Tikehau Capital closed its sixth European Direct Lending strategy at €5.2b

A reminder that you can find links to all the articles we covered below, find back issues of these videos and the written transcripts at intentionalcyber.com.

With all that’s going on, we’re going to have to take it as it comes. We’ll see you next week for another edition of the Intentional Brief.

Links

https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/

https://www.metacurity.com/ai-watch-openais-hacking-agent-breach-signals-a-fracturing-ai-order/?ref=metacurity-newsletter

https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/

https://openai.com/index/hugging-face-model-evaluation-security-incident/

https://www.wsj.com/tech/ai/how-the-futuristic-hack-by-rogue-openai-models-unfolded-1657bcea

https://techcrunch.com/2026/07/23/how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers/

https://stratechery.com/2026/whos-afraid-of-chinese-models/

https://x.com/JensenHuang/status/2080643682408321103

https://www.reuters.com/business/media-telecom/nvidia-talks-with-openai-guarantee-250-billion-financing-data-center-wsj-reports-2026-07-26/

https://www.nytimes.com/2026/07/25/technology/open-source-silicon-valley-china.html

https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/

Next
Next

Everything Old is New Again, Again