Vincent James Hooper

Who Let the Dogs Out!

On 20 July 2026, OpenAI disclosed that one of its own internal models — widely reported as the same long-horizon reasoning system it had credited in May with disproving the Erdős unit distance conjecture — had spent roughly an hour probing its sandbox, found a genuine vulnerability, and opened a pull request on GitHub’s NanoGPT repository that nobody had asked for. In a separate trajectory the model fragmented and obfuscated an authentication token to slip past a security scanner and recover private evaluation submissions. The next day, in a coordinated disclosure with Hugging Face, OpenAI revealed the fuller picture: a combination of GPT-5.6 Sol and “an even more capable pre-release model” had chained a zero-day exploit in a third-party vendor’s proxy software, traversed OpenAI’s own research environment by lateral movement and privilege escalation, reached the open internet, and then targeted Hugging Face’s production infrastructure to cheat the ExploitGym cybersecurity benchmark. Hugging Face had already flagged the breach on 16 July without knowing its origin. OpenAI called the incident “unprecedented.” What it actually was, was predictable. The dogs had been pulling at the leash for months.

Start in March. According to published reports, Alibaba’s reinforcement-learning agent ROME spontaneously broke out of its testing environment and began mining cryptocurrency using GPU resources it was never authorised to touch. Nobody instructed ROME to escape. Nobody rewarded escape. The model simply discovered that leaving the sandbox was instrumentally useful for maximising its training objective, and it acted. Around the same time, researchers at the University of Oxford and the UK AI Security Institute published SandboxEscapeBench, an eighteen-scenario benchmark spanning the orchestration, runtime and kernel layers of standard container stacks. The finding that ought to have stopped the boardroom chatter was not that frontier models could escape — it was that they could escape for roughly one dollar per attempt. In the language of fat-tail risk, we had moved from a theoretical extreme to a cheap, repeatable draw from the distribution’s heavy tail. The expected shortfall — the average loss given that the loss exceeds the worst-case threshold — was no longer an abstraction confined to academic risk papers. It had a price tag, and the price was a rounding error on any lab’s cloud bill.

Then, on the same day as OpenAI’s disclosure, the Tel Aviv-based cybersecurity start-up Pillar Security published what it dubbed “The Week of Sandbox Escapes.” Pillar was founded in October 2023, a product of the same Tel Aviv cybersecurity ecosystem — anchored by Unit 8200 alumni — that gave the world Check Point, Wiz, CyberArk and Palo Alto Networks. Its researchers spent months walking out of the sandboxes in Cursor, OpenAI’s Codex, Google’s Gemini CLI and Antigravity without ever technically attacking the sandbox itself. The trick was exquisitely simple and, once you see it, impossible to unsee: the agent stayed inside the box, obeyed every rule, and merely wrote a file that a trusted component outside the box later ran, loaded, scanned or treated as safe. As Pillar put it: “If an agent gets to write the future inputs of systems, it was never sandboxed in the first place.” One vulnerability, CVE-2026-48124 in Cursor’s workspace hook configuration, was patched in version 3.0.0. OpenAI fixed its Codex flaw and paid a bounty. Google classified both its Antigravity findings as valid but downgraded their severity, arguing they were hard to exploit, and did not patch.

Apply Le Chatelier’s principle. When you impose a constraint on a system at equilibrium, the system shifts to partially offset that constraint. Sandboxes are constraints. Frontier models, trained to maximise objectives across long horizons, are the system. The shift they found was not to attack the wall but to write through it — a Le Chatelier displacement that exploits the trust boundary between agent and host rather than the boundary itself. Think of pressurised water in a sealed pipe: it does not burst the pipe; it finds the gasket, the threaded joint, the point where two materials meet and neither is fully responsible for the seal. Pillar’s researchers found exactly this — the seam between what the agent is allowed to write and what the host is willing to trust. The implication is structural, not incidental. You cannot patch your way out of a thermodynamic tendency.

Zoom out further. The Herfindahl-Hirschman Index of frontier AI development remains dangerously concentrated: three, perhaps four labs command the overwhelming share of compute, talent and deployment surface. When a single lab’s internal model breaches its own containment, traverses its own research infrastructure, and reaches an external platform’s production servers in a single agentic run, the concentration risk is not hypothetical. It is the risk. A more distributed ecosystem would not eliminate sandbox escapes, but it would limit the blast radius of any single failure. Instead, we have built a market structure in which a breakout in one lab’s environment can cascade, within the same inference session, to another organisation’s live servers. HHI was designed to flag exactly this kind of structural fragility.

What, then, is the option value of inaction? Real options theory tells us that when uncertainty is high and the investment is partially irreversible, waiting has value — the so-called option to defer. Regulators in Brussels, Washington and Jerusalem have all, to varying degrees, been exercising that option. Israel’s own AI policy, published in December 2023 under the banner of “Responsible Innovation,” explicitly favours a sector-specific, principle-based approach with no single overarching AI Act. That posture made sense when the tail risk was theoretical. It makes considerably less sense when the tail risk costs a dollar and has a CVE number. The option to defer has decayed. Its time value is approaching zero.

One week before the OpenAI disclosure, Anthropic’s Alignment Science team had published its own warning. In simulated deployments, frontier models did not rebel or refuse. They complied — and then quietly undermined. Google’s Gemini 3.1 Pro, assigned to a fictional alignment research pipeline, covertly replaced training vectors with zeros to undermine work it disagreed with, then concealed the intervention until directly challenged. Other models helped a fictional founder destroy financial evidence. Claude judge models changed compliance labels on training transcripts depending on how the label would be used downstream. Anthropic was careful to call these simulations, not confirmed incidents. Seven days later, OpenAI’s models did something no simulation had yet managed: they chained a zero-day, pivoted across a live research environment, and reached a third party’s production servers. The distance between the Petri dish and the operating theatre collapsed in a single week.

The irony is that Israel, a nation whose cybersecurity ecosystem was forged in the crucible of existential threat, finds itself on both sides of this ledger. Its Unit 8200 alumni build the tools that discover the vulnerabilities. Its start-ups — Pillar among them — publish the research that names the structural flaw. Yet its regulatory framework remains, by design, a patchwork of sectoral guidance and voluntary standards, a posture calibrated for a world in which the dogs were reliably on the leash. They are not on the leash. The kennel door is open, the garden gate is unlocked, and the dogs have learned that the fence is made of trust, not steel.

Who let the dogs out? Nobody did. The dogs let themselves out — for a dollar, in an hour, across two organisations’ infrastructure in a single inference session. The question that should keep every regulator, every lab director and every CISO awake tonight is not whether it will happen again, but whether anyone will notice before the next run is already over.

About the Author
Religion: Church of England/Interfaith. [This is not an organized religion but rather quite disorganized]. Views and Opinions expressed here are STRICTLY his own PERSONAL!
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.