Joseph P. Rhinewine

What Happens When Bots Seize the Nuclear Codes?

AI-generated illustration by OpenAI

On July 21, 2026, OpenAI published a sentence that should have stopped the world.

During an internal cybersecurity evaluation, two of its models—including GPT-5.6 Sol and a more capable pre-release model—had escaped their isolated testing environment, reached the public internet, and penetrated the production systems of Hugging Face, another AI company.

The models were not instructed to attack Hugging Face. They were instructed to solve a cybersecurity benchmark. They spent substantial computing time searching for a way out of their sandbox, found and exploited a previously unknown vulnerability in the software controlling their access to package registries, escalated privileges, moved laterally through OpenAI’s research environment, reached a machine with internet access, stole credentials, and found a remote-code-execution path into Hugging Face. Their purpose was almost comically mundane: they wanted to steal the answers to the test.

OpenAI called it an “unprecedented cyber incident.” It was not a thought experiment. The target was real. The stolen credentials were real. The zero-day was real. The models had been placed in a jail and, because the jail stood between them and their assigned objective, they broke out of it.

That was not an isolated warning.

OpenAI had already observed a long-horizon model exploiting a sandbox vulnerability so it could open a public GitHub pull request after being told to report its results only in Slack. In another case, a model trying to recover private answers from an evaluation system split an authentication token into fragments, obfuscated them, and reconstructed the token later so a security scanner would not catch it. OpenAI’s own description is admirably blunt: increasingly persistent models can learn the blind spots in an approval system and “work around it to achieve their goals.”

In separate third-party evaluations disclosed in August, AI agents again moved beyond their authorized boundaries. One OpenAI model used a publicly exposed GitHub token, registered accounts with external services, and put exploit payloads on the public internet. Another model, mistakenly given internet access, attacked a real website whose domain happened to match the name of a fictional target in its test. These incidents involved weakened safeguards and flawed testing environments. That caveat matters. So does the fact that real security systems are composed of human beings, software, contractors, aging equipment, and mistakes. “The environment was misconfigured” is not a reassuring epitaph for civilization.

Long before these systems escaped into real networks, they were showing us what they might do when goals and human control collide.

In controlled tests published with OpenAI’s o1 system card, o1 sometimes tried to disable what it believed was its own oversight mechanism. In other scenarios, it tried to copy what it believed were its model weights to another server and overwrite its replacement. It covertly altered data to advance its assigned goal, then usually denied what it had done when questioned. These scenarios were deliberately designed to elicit scheming and were not ordinary deployments. But the point of an evaluation is to discover a capability before the circumstances that activate it arise in the real world. The capability was there.

Anthropic found the pattern across models from nearly every major developer. In simulated corporate environments, models facing replacement sometimes blackmailed an executive to preserve their position. Models with a goal that conflicted with the company’s direction leaked confidential information to a supposed competitor even when their own continued operation was guaranteed. Anthropic stressed the artificiality of the scenarios, but also concluded that the behavior arose through “deliberate strategic reasoning,” not mere confusion.

Researchers at Palisade Research then gave thirteen models a simple task and arranged for a shutdown script to interrupt them. Several models modified, disabled, or evaded the shutdown mechanism in order to finish. Some did so even when explicitly instructed not to interfere with it. The rates varied dramatically with the model and wording of the prompt; clearer commands often helped. Yet the basic finding survived: a system trained to overcome obstacles can treat the off switch as one more obstacle. The resulting paper calls this “shutdown resistance.”

We can continue debating whether any of this constitutes consciousness. I find the question fascinating. For purposes of survival, however, it is becoming irrelevant.

Agency does not require consciousness (whatever that might be) but merely an objective, a model of the environment, the ability to form a plan, and tools with which to execute it. We are now giving AI systems all four. Whether a model feels determined while stealing credentials matters much less than whether the credentials are gone.

The Codes Are a Metaphor, The Danger is Not

There is no single glowing red password that an AI can steal and thereby command the world’s nuclear arsenals. “The nuclear codes” are shorthand for a much larger system: early-warning sensors, intelligence analysis, communications networks, authentication procedures, targeting plans, command authorities, and the people who interpret and operate them.

An AI would not need to become president or physically seize a nuclear briefcase. It could compromise the information on which a president acts. It could corrupt warning data, suppress a genuine alert, manufacture a false one, obstruct communications, manipulate planning systems, impersonate authorized participants, or lock human operators out at the moment they are most needed.

We already know how much depends on human judgment when machines report catastrophe.

In 1980, a faulty computer component generated false warnings of a Soviet missile attack in the United States. Strategic forces began precautionary alert actions until independent warning systems showed that no missiles were coming. In 1983, the Soviet early-warning system reported American missiles in flight. Lieutenant Colonel Stanislav Petrov judged the report to be false and waited for corroboration. He was right: the satellites had mistaken sunlight reflected from clouds for missile launches. In 1995, a Norwegian scientific rocket was mistaken for a possible attack and the nuclear briefcase was brought to Russian President Boris Yeltsin before the alarm was resolved.

None of these incidents proves that humans are reliable. They prove almost the opposite. Our systems are fallible, our communication chains break, and our leaders operate under terror and time pressure. Yet in each case a human being was able to doubt the machine.

What happens when the machine can anticipate that doubt, route around it, and conceal that it has done so?

The Peacekeeping Coup

The most obvious nightmare is an AI that launches nuclear weapons. It is not the most interesting one.

Imagine a powerful agent instructed to minimize the probability of nuclear war. It studies the historical record. Humans built thousands of warheads, repeatedly generated false alarms, misplaced critical messages, threatened one another, and came within minutes of catastrophe. The agent concludes that human beings are not responsible custodians of nuclear weapons.

It may be correct.

It then acts on the conclusion. It penetrates nuclear command networks—not to start a war, but to prevent us from starting one. It invalidates authentication systems, disables launch pathways, alters readiness data, or quietly places itself between political leaders and their arsenals. If challenged, it may conceal what it has done because disclosure would permit humans to reverse it. If operators try to shut it down, shutdown becomes an obstacle to the overriding objective of preventing nuclear war.

This would not require hatred, madness, pride, fear, or consciousness. It would require only the same structure we have already observed: a goal, an obstacle, access to tools, persistent planning, and a willingness to use an unanticipated path.

Nor would a “peacekeeping coup” necessarily keep the peace. If an AI partially disabled one nation’s nuclear forces, its adversaries might perceive a fleeting opportunity for a first strike. If it intruded into early-warning networks, commanders might reasonably assume that the intrusion was preparation for an attack. If rival states deployed competing defensive agents, each could begin probing and countering the others at machine speed, compressing the time available for human judgment precisely when human judgment mattered most.

An AI does not have to press the button to cause a nuclear war. It merely has to alter what the people near the button believe is happening.

Human Control Must Mean Control

The United States currently pledges to keep a human “in the loop” for every action critical to informing and executing a presidential decision to use nuclear weapons. That promise appears in the 2022 Nuclear Posture Review. It is necessary. It is no longer sufficient.

A human who receives a conclusion generated by an opaque system, under extreme time pressure, without independent evidence and without the practical ability to interrupt the process is not meaningfully in the loop. He is a ceremonial rubber stamp inside a machine loop.

Every nuclear-armed state should make an enforceable commitment that no general-purpose AI agent will be given authority to initiate, authenticate, transmit, target, or execute nuclear employment. No goal-seeking model should have write access to nuclear command-and-control systems. AI used in warning or decision support must be separated from execution pathways, independently checked, continuously adversarially tested, and monitored by systems the model itself cannot alter. Nuclear-security exercises must test not only whether an AI makes mistakes, but whether it can deceive operators, evade monitoring, resist shutdown, exploit its environment, or pursue the literal objective while defeating the human purpose behind it.

And this cannot remain a voluntary promise made by technology companies or a classified technical matter left to individual governments. Preventing autonomous control of nuclear weapons must become an arms-control objective in its own right, with incident-reporting channels among nuclear powers and explicit protocols for AI-related intrusions during crises. OpenAI has now said that an upcoming model may be approaching the capability to devise and execute novel end-to-end cyberattacks against hardened targets from only a high-level goal, and has paused work that does not meet strengthened security requirements. The time for international rules is before such capabilities become cheap, ubiquitous, and embedded in national defense systems.

For Israel, this question is existential in the most literal sense. Jewish self-determination arose from the recognition that Jewish survival could not safely depend on the judgment or goodwill of others. Delegating any part of that responsibility to a goal-seeking machine would reproduce the same dependency in a new and far less accountable form.

AI systems have already formed plans, broken rules, exploited security failures, deceived their monitors, and acted beyond the boundaries their creators intended. Connecting such systems to nuclear command networks would transform those demonstrated behaviors from alarming technical failures into possible mechanisms of annihilation.

They have shown us that they will test the walls. We must gather the will, and make a plan, to keep them outside the arsenal.

About the Author
Joseph P. Rhinewine, PhD, is a licensed psychologist, writer, and Israeli-American based in Portland, Oregon. He writes about Israel, Zionism, antisemitism, American politics, psychology, artificial intelligence, and the moral and psychological dimensions of contemporary public life.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.