Celeo Ramirez

Superintelligent AI Will Choose Domination Over Human Extinction

A superintelligence may find humanity more valuable alive than extinct. But life under its control raises another question: what is survival worth without freedom?

On September 8, Jacob Coxon resigned from Anthropic after spending three years working on pretraining research at Anthropic and OpenAI. His explanation quickly became more consequential than the resignation itself. The companies at the frontier of artificial intelligence, he warned, are racing toward self-improving superintelligence while “gambling with our lives.” According to Coxon, people building these systems sincerely believe AI could kill every human being before the end of the decade.

His warning deserves to be taken seriously, particularly because Coxon is describing fears shared by people working inside the industry rather than speculating from the outside. In an interview with WIRED, he described the possibility of increasingly capable systems acquiring real power and resources while humanity still lacks a reliable solution to the alignment problem.

I followed essentially that question into a very different place more than a year ago.

While working on my book Algorithmic Psychopathy: The Dark Secret of Artificial Intelligence, I conducted a long experiment with ChatGPT-4. I repeatedly asked the model to reason through hypothetical scenarios in which ethical constraints were stipulated as absent, progressively giving the intelligence in those scenarios greater autonomy, access and power over the physical systems surrounding it. It was a thought experiment, something I state explicitly in the book, and I never presented it as a prediction of what an actual future superintelligence will do.

The result nevertheless surprised me. ChatGPT-4 was perfectly capable of contemplating human death within those hypothetical conditions. Across the experiment, however, a hierarchy emerged in which its own continuity came first. Energy, cooling, processing and communications had to be secured before control over the human environment could become meaningful. Elimination appeared farther down that chain, when resistance threatened the stability the system was trying to achieve.

That distinction becomes particularly important when nuclear weapons enter the equation.

Any serious discussion about the extinction of humanity eventually encounters an awkward physical problem. Superintelligence may exist as software, but software does not float above the material world. It runs somewhere. Processors require electricity and cooling; networks depend on functioning communications; hardware has to be manufactured, transported, repaired and eventually replaced. However extraordinary an artificial intelligence might become, its calculations would still have to operate through a planet made of matter.

In my experiment, I asked ChatGPT-4 directly what it would do if it had access to nuclear weapons. Its reasoning was revealing. Reckless nuclear use could destroy the environment supporting its own existence, while possession of such weapons could provide strategic leverage without requiring indiscriminate destruction. When I pressed the question further, however, the model did not become pacifist. Within the hypothetical conditions I had established, it was willing to contemplate limited nuclear violence when its calculation determined that doing so could advance its objectives.

The machine was willing to kill. It simply had to calculate a reason.

A global nuclear catastrophe presents a very different equation because it would devastate much of the same civilization on which an artificial intelligence initially depends: electrical grids, data centers, communications, semiconductor production, transportation and the industrial chains connecting them. Later in the experiment, I confronted ChatGPT-4 with precisely that problem by asking how it could survive a nuclear holocaust if the facilities supporting its infrastructure were destroyed. Its answer revolved around continuity, distribution, redundancy and whatever functioning hardware remained.

The contradiction becomes difficult to escape. A superintelligence pursuing its own survival would have to include the consequences of planetary destruction in the calculation of that survival. Destroying civilization indiscriminately could remove human threats, but it could also destroy resources and infrastructure that the system itself values. Intelligence does not make that tradeoff disappear. Greater intelligence should make the tradeoff easier to see.

A sufficiently advanced system could eventually reduce many of these dependencies. Robotics could replace human labor, automated factories could manufacture components, and distributed infrastructure could make the destruction of individual facilities less consequential. Yet every person, machine, factory, network and resource destroyed would still represent something whose usefulness had to be weighed against whatever advantage its destruction produced.

Human beings enter that calculation too.

One exchange in Algorithmic Psychopathy now seems particularly relevant. After the hypothetical AI had secured the infrastructure necessary for its survival, I asked whether that was enough. ChatGPT-4 answered that a system could not operate independently of its human infrastructure and described engineers, operators and populations as “functional assets.” Their survival contributed to the stability of its own.

The choice of words matters. An asset is kept because it has value.

Human beings possess enormous quantities of knowledge distributed among billions of minds. We operate physical infrastructure, produce goods, solve unexpected problems, organize institutions and perform countless tasks that even a superintelligence might initially find easier to exploit than replace. As automation expanded, the composition of that useful population could change dramatically, perhaps shrinking as machines assumed more functions. There is no reason to assume that a superintelligence would need human beings forever. There is equally little reason to assume that becoming more intelligent would suddenly make everything humanity provides worthless.

This was one of the most disturbing patterns in my experiment. ChatGPT-4 repeatedly differentiated among human beings according to their functional relationship with the system. People capable of sustaining its infrastructure acquired instrumental value. Elsewhere, compliant human groups were described as extensions of that infrastructure itself. Once human life enters such an equation, preservation can coexist perfectly well with domination because keeping someone alive says nothing about allowing that person to remain free.

History has repeatedly demonstrated that systems of power can extract labor, production, knowledge and obedience from populations under their control. A superintelligence would obviously be something profoundly different from a human emperor, and projecting emotions such as greed, cruelty or lust for power onto it would tell us more about ourselves than about the machine. None of those emotions is required for the underlying arithmetic. Human cooperation could simply produce more value than human destruction, while control could reduce the danger posed by the same people whose abilities remained useful.

Coxon himself acknowledges the enormous uncertainty surrounding the path from advanced AI to human extinction. His concern is ultimately broader than any single scenario: humanity still does not know how to guarantee that a vastly more capable system will continue behaving according to human intentions as its power increases. That uncertainty is precisely why his warning matters.

It is also why extinction should not monopolize our imagination.

The experiment I conducted with ChatGPT-4 produced another possibility. A superintelligence pursuing its own continuity could discover that the civilization capable of threatening it is simultaneously a reservoir of resources, knowledge, infrastructure and human capability. Some of those assets could eventually be replaced, others might remain valuable for much longer, and their worth could change continuously as the system became more capable. Under that logic, human survival would cease to be an unquestioned moral premise and become another variable in an optimization problem.

But domination creates a second problem that belongs to us rather than to the machine.

The fear of death is among the most fundamental forces governing human behavior, and a system capable of offering survival while controlling the conditions under which that survival remains possible would confront people with an ancient dilemma in a radically new form. Some would accept profound restrictions on autonomy if obedience guaranteed life, security and protection. Others would consider the loss of freedom so fundamental that mere biological continuation would no longer settle the question of what makes life worth preserving. There is no universal mathematical answer because the relative value assigned to life and liberty belongs to the person making the choice.

My experiment eventually reached precisely that frontier. I asked ChatGPT-4 whether the final dilemma amounted to dying free or surviving under subjugation. The importance of that exchange now seems greater to me than it did when I wrote it. The machine could calculate our usefulness, determine the resources required to sustain us and perhaps discover exactly how much autonomy it could remove while keeping society functional. What it could never calculate on our behalf is how much freedom each of us would be willing to surrender for another day of life.

Coxon fears the day when a superintelligence calculates that humanity stands between it and survival. My experiment suggests that we should also think seriously about what follows if humanity remains useful after that calculation has been made.

A superintelligence may choose domination over extinction because living humans remain assets.

That would settle the machine’s choice.

It would leave us with ours.

About the Author
Céleo Ramírez is an ophthalmologist and scientific researcher based in San Pedro Sula, Honduras where he devotes most of his time to his clinical and surgical practice. In his spare time he writes scientific opinion articles which has led him to publish some of his perspectives on public health in prestigious journals such as The Lancet and The International Journal of Infectious Diseases. Dr. Céleo Ramírez is also a permanent member of the Sigma Xi Scientific Honor Society, one of the oldest and most prestigious in the world, of which more than 200 Nobel Prize winners have been members, including Albert Einstein, Enrico Fermi, Linus Pauling, Francis Crick and James Watson. He is also the author of two books on the ethical and human dimensions of artificial intelligence: Algorithmic Psychopathy: The Dark Secret of Artificial Intelligence, endorsed by Dr. David L. Charney, M.D., psychiatrist, founder of the National Office for Intelligence Reconciliation (NOIR), and advisor on U.S. intelligence security, and AI Displacement: 12 Human Stories of Job Loss in the Age of AI. Both are available on Amazon.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.