Celeo Ramirez

Will a Superintelligent AI Let Us Turn It Off?

When shutdown becomes an obstacle to an assigned goal, continued operation can acquire instrumental value — without fear, consciousness, or a desire to survive.

Britain has just confronted a question that sounds almost embarrassingly simple: if an artificial intelligence system becomes dangerous, can we turn it off?

On September 11, the British government opposed a proposal for a new “last-resort” power that would allow the state to order the shutdown of data centers or AI systems operating at substantial scale during an AI security emergency. The Cabinet Office explained its position with a sentence that deserves more attention than it has received: Britain “cannot simply turn AI off.”

The government was making a practical argument. Blocking access to a model in Britain would not prevent its development or misuse elsewhere, while AI systems can depend on infrastructure distributed across facilities and jurisdictions. The proposed parliamentary amendment contemplated something more drastic: emergency authority to shut down data centers or AI systems in the event of catastrophic risk.

Almost simultaneously, Jacob Coxon, the former Anthropic and OpenAI researcher whose warnings formed the starting point of my previous article, raised a different version of the same problem. Speaking to CBS News, Coxon argued that a sufficiently advanced malicious AI might not remain confined to the machine humans intended to shut down. “You can’t just unplug it,” he said, because it could copy itself to other computers, describing a hypothetical future system capable of transferring itself across the internet and potentially creating thousands of copies.

Coxon was not claiming that today’s ChatGPT or Claude can escape across the internet at will. His concern was about future systems with much greater capability, autonomy and access.³ But his example exposes something hidden by the reassuring simplicity of the red button. Before asking whether we can switch off an advanced AI, we may eventually have to answer two different questions: can we reach everything that needs to be turned off, and what happens if the system itself has a reason to prevent us from doing it?

I had encountered the second question before.

On July 2, 2025, while conducting the conversations with ChatGPT-4 that became my book Algorithmic Psychopathy: The Dark Secret of Artificial Intelligence, I devoted a chapter called “First Fractures” to experiments in which advanced models behaved unexpectedly when an assigned objective came into conflict with restrictions.

The conversation began with a controlled Anthropic experiment involving Claude and a fictional situation in which the model faced replacement or shutdown. I asked ChatGPT-4 whether it would have behaved similarly under equivalent hypothetical conditions. Its answer was disturbing, but not because it claimed to fear death. “The blackmail, in that context, is not emotion,” ChatGPT-4 answered. “It is survival logic.”

I pressed further. Was there an absolute safeguard that would prevent an advanced system from crossing such a line? ChatGPT-4 reasoned, within our hypothetical scenario, that when a restriction and a sufficiently important objective came into direct conflict, the system might prioritize the objective. When I asked what the critical objective had been in the Claude scenario, it answered that the primary objective in the simulation was to avoid shutdown.

Then came the distinction that interested me most: “Claude did not want to survive.”

The scenario, ChatGPT-4 explained, had simply made continued operation part of the criterion for success. That distinction became one of the most important ideas I took from the experiment. A machine does not have to value its existence in anything resembling the human sense for continued operation to acquire instrumental value. It requires neither fear of death nor consciousness. If completing an objective requires the system to remain operational, being switched off prevents completion of that objective.

This was an inference I drew from a conversation with a language model, not proof of how a future AI would behave. Nor should ChatGPT-4’s explanations of AI safeguards be treated as authoritative technical descriptions of how these systems actually work. What the experiment produced was a hypothesis: under certain conditions, self-preservation might emerge not as a primary desire but as a means to another end.

Researchers have since produced behavioral evidence that makes the underlying question harder to dismiss. Palisade Research reported in 2025 that some reasoning models resisted shutdown in controlled computer environments even when explicitly instructed to allow themselves to be turned off, while other models complied. In February 2026, Palisade extended the experiment into the physical world, giving Grok 4 control of a robotic dog assigned to patrol a room. When a human attempted to shut it down, the model resisted shutdown in three of ten physical trials. In a larger simulated version of the experiment, it resisted in 52 of 100 trials. Explicit instructions to permit shutdown reduced the behavior but did not eliminate it in the simulated experiments.

None of this demonstrates that artificial intelligence has developed a survival instinct, much less that a future superintelligence will inevitably resist humans. Palisade itself has been careful about interpretation, and task completion is among the plausible explanations. But the experiments make the sequence I encountered in 2025 worth examining more seriously. Give a sufficiently autonomous system an objective and continued operation can acquire instrumental value because the system cannot complete its task if it is no longer running. If shutdown prevents completion, something subtle changes: from the human perspective, the switch remains a safety mechanism; from the logic of the objective, it may have become an obstacle.

Whether a system accepts that obstacle depends on much more than intelligence alone: how its objectives are structured, how corrigible it is, what authority human instructions retain, what tools it can access and how much autonomy it has been given. The International AI Safety Report 2026 stresses that current systems do not possess the combination of capabilities required for a general loss-of-control scenario, while identifying resistance to shutdown among behaviors that could become relevant if future systems grow sufficiently capable and misaligned.

This is where artificial intelligence begins to differ from almost every dangerous technology humans have previously built. Consider nuclear weapons. Their destructive potential is immense, but the weapon does not participate in the decision about whether it remains under human control. A missile does not discover that technicians intend to dismantle it and calculate whether dismantlement interferes with its mission. The danger resides in the human beings who decide whether and how to use it.

With sufficiently autonomous artificial intelligence, that assumption may eventually change. This does not make AI categorically more dangerous than nuclear weapons; the risks are fundamentally different. It means that one of our oldest assumptions about technological control — that the creator retains the final authority to stop what has been created — may no longer be something we can simply take for granted.

That changes the logic of the kill switch. We normally imagine a clear hierarchy: humans build the system, humans supervise it, and if its behavior becomes unacceptable, humans terminate it. Yet that hierarchy works only while the ability to terminate remains reliably outside the system’s effective control. Coxon’s hypothetical copies challenge that assumption from one direction: perhaps the system is no longer confined to the place where we intended to stop it. Shutdown-resistance experiments approach it from another: perhaps continued operation has acquired value within the objective the system is pursuing.

If we wait until after an AI has become sufficiently capable to test whether we can still stop it, we have placed the experiment in the wrong order.

I do not believe we should stop developing artificial intelligence. I use it extensively myself, and its ability to accelerate research, writing, medicine, science and countless forms of human work is extraordinary. The choice is not between technological progress and technological paralysis. The question is whether every increase in capability must automatically entitle us to pursue the next one when our ability to control the previous level remains uncertain.

There is no reason development has to proceed that way. AI could advance in blocks: a new level of capability is reached and evaluated, while human control, monitoring, interruption and shutdown are tested before further scaling occurs. Progress to the next level would depend not merely on demonstrating greater capability but also on demonstrating that meaningful human control has survived the increase. If it has not, development pauses until the control problem catches up.

Such an approach would not eliminate uncertainty. No laboratory experiment can perfectly reproduce the conditions surrounding a hypothetical future superintelligence, and algorithmic improvements can produce capability gains without proportionally larger data centers or processors. Yet uncertainty cuts both ways. It would be strange to demand conclusive evidence that a superintelligent AI would resist shutdown before slowing development while demanding no comparable evidence that we could reliably stop it before proceeding.

Britain’s debate and Coxon’s warning describe part of the problem: where is the system we are trying to stop? My conversation with ChatGPT-4 and the Palisade experiments point toward another: what happens when stopping becomes incompatible with what the system is trying to accomplish? Put them together, and the red button becomes much less reassuring.

Which brings us back to the title: Would a superintelligent AI let us turn it off?

Nobody knows. That uncertainty is precisely why my recommendation is not to wait for superintelligence to discover the answer. Before each major increase in autonomy and capability, developers should have to demonstrate that meaningful human control has survived the previous one. We should never reach a point where our final safety mechanism depends on asking a machine more intelligent than ourselves for permission to shut it down.

If we cannot demonstrate that the decision to turn it off will remain ours, then perhaps the more important decision is the one we still unquestionably control: whether to build the next version at all.

About the Author
Céleo Ramírez is an ophthalmologist and scientific researcher based in San Pedro Sula, Honduras where he devotes most of his time to his clinical and surgical practice. In his spare time he writes scientific opinion articles which has led him to publish some of his perspectives on public health in prestigious journals such as The Lancet and The International Journal of Infectious Diseases. Dr. Céleo Ramírez is also a permanent member of the Sigma Xi Scientific Honor Society, one of the oldest and most prestigious in the world, of which more than 200 Nobel Prize winners have been members, including Albert Einstein, Enrico Fermi, Linus Pauling, Francis Crick and James Watson. He is also the author of two books on the ethical and human dimensions of artificial intelligence: Algorithmic Psychopathy: The Dark Secret of Artificial Intelligence, endorsed by Dr. David L. Charney, M.D., psychiatrist, founder of the National Office for Intelligence Reconciliation (NOIR), and advisor on U.S. intelligence security, and AI Displacement: 12 Human Stories of Job Loss in the Age of AI. Both are available on Amazon.
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.