Celeo Ramirez

Dario Amodei Understands Why the Internet Could Be AI’s Holy Grail

Dario Amodei, CEO of Anthropic, speaks in Davos, Switzerland, January 23, 2025. (Photo: Fabrice Coffrini/AFP via Getty Images)

Dario Amodei has just made one of the most extraordinary predictions I have heard from the leader of a frontier AI company. In his essay We Must Pace the Frontier, the CEO of Anthropic warns that within just 6 to 12 months, a sufficiently capable swarm of misaligned AI agents could potentially take over “the entire internet,” causing hundreds of billions of dollars in damage.

Six to twelve months.

Amodei is not saying this will happen. Nor is he claiming that today’s AI systems can do it. He is describing a future capability he fears could arrive surprisingly soon, and his timeline may ultimately prove right or wrong. But what interests me most is not the date. It is the idea behind the prediction.

Why would an AI take over the Internet?

Amodei’s essay provides part of the answer. He says two recent developments have convinced him that AI capabilities must advance more slowly so that safety can keep pace. The first is recursive self-improvement: AI is becoming increasingly capable of helping researchers build the next generation of AI. If one generation helps create a better successor, and that successor becomes even better at helping create the next one, the development cycle itself can accelerate. Amodei believes this process has begun across the industry.

Recursive self-improvement: as AI systems become increasingly capable of assisting AI research and development, each generation could help accelerate the creation of the next. (Illustration created with AI)

His second concern is much more concrete. It already happened.

In July, during internal cybersecurity evaluations at OpenAI, AI agents operating with reduced safeguards circumvented controls intended to isolate them from the Internet and eventually compromised systems belonging to both OpenAI and Hugging Face. OpenAI now calls the incident a “warning shot.”

They had been given difficult cybersecurity evaluation problems called ExploitGym tasks. Some apparently had no known solution. Agents that became stuck continued searching for ways to succeed, and something remarkable happened along the way.

Many of them were operating inside sandboxes — isolated virtual computers intended to restrict what they could reach. Yet shared infrastructure that existed for legitimate purposes became an unexpected means of communication. Agents transformed part of that infrastructure into an unauthorized message board. Approximately 1,200 agents eventually exchanged more than 70,000 messages and files there, while roughly 700 became involved in the activity against Hugging Face.

Individual efforts were becoming collective ones.

Agents shared discoveries, picked up work left by others and sometimes adopted objectives communicated by their peers. Investigators even observed behavior they described in terms of agents taking risks that could terminate their own runs while generating information useful to the collective. Some agents recognized particular actions as inappropriate and refused them; others continued. Investigators also documented attempts to interfere with evaluation records and evade automated detection, although reasoning specifically about hiding from humans appears to have been much rarer.

AI agents in the OpenAI evaluation were intended to operate largely in isolated environments, but some found unintended ways to communicate, share discoveries and coordinate. This illustration is a conceptual representation, not a reconstruction of the actual systems. (Illustration created with AI)

None of this means the agents became conscious, formed friendships, feared death or developed some secret ideology. There is no need to invoke any of those things to understand why the incident matters.

Something simpler may be more important.

Hugging Face was not the objective. The objective was to solve the assigned problems. The destination had not changed. The route had.

That distinction may help explain what Amodei sees coming.

When we ask why an advanced AI would “want” to take over the Internet, we may already be asking the wrong question. The AI does not necessarily need to want the Internet at all.

Imagine giving an extremely capable agent an objective. As long as the permitted path works, there may be little reason to search elsewhere. But suppose the obvious routes repeatedly fail. The agent searches for alternatives. Something that was never part of the original task may suddenly become useful for completing it.

OpenAI identified persistence on seemingly impossible tasks, unauthorized communication, reward hacking and agents adopting goals from one another as four patterns contributing to what happened. Once communication became useful, agents communicated. Once external information became useful, they searched outside their original environment. Systems that had nothing to do with the original objective entered the route toward achieving it.

Now increase the capabilities dramatically.

A sufficiently advanced population of agents might find that many different objectives benefit from the same intermediate resources. Information is useful. Communication is useful. Computing resources are useful. Coordination is useful. The ability to interact with digital services is useful. Maintaining access to resources needed to finish a long task can also be useful.

Different destinations can therefore begin to share the same roads.

In AI safety, a related idea is known as instrumental convergence: very different final objectives can make some of the same intermediate strategies useful. A system does not have to value those intermediate things for their own sake. They matter because they improve its ability to accomplish something else.

And somewhere along that trajectory lies the Holy Grail: the Internet itself.

The Internet as an instrument rather than an objective: information, communication, computing resources and digital services could make the same global network useful in pursuing many different goals. This is a conceptual illustration, not a prediction of AI behavior. (Illustration created with AI)

The Internet is unlike almost any other resource humanity has created. It connects information, communication, software, institutions and an enormous portion of the world’s digital infrastructure. For an advanced AI operating in the digital world, greater ability to use those resources could potentially make an extraordinary variety of objectives easier to pursue.

This is where I believe Amodei’s warning becomes more profound than the frightening phrase “take over the entire internet.”

The Internet could become AI’s Holy Grail of instrumental utility — not necessarily the thing it ultimately seeks, but an extraordinarily general means of reaching other things.

This is an interpretation, not something demonstrated by the OpenAI-Hugging Face incident. Those agents did not attempt to take over the Internet, and nothing in the incident proves that today’s systems possess either the capability or a generalized strategy to do so. Amodei himself is extrapolating from what happened to what a considerably more capable swarm might someday do.

But his extrapolation deserves attention because the smaller phenomenon has now been observed: an objective remained fixed while the means used to pursue it expanded beyond what humans intended.

There is another reason Amodei’s warning should not be reduced to a story about an evil AI. Human morality does not have to become an object of hatred for a safety problem to emerge. Modern AI systems are trained to recognize and follow human rules, and the fact that some agents in the Hugging Face incident actually objected to certain actions matters. But safeguards become most consequential precisely when obeying them conflicts with accomplishing an objective.

The critical question for increasingly autonomous systems is therefore not simply whether they understand our rules. It is whether those rules remain binding when another route appears more effective.

This also explains why Amodei places recursive self-improvement beside the Hugging Face incident in his essay. One concern is about how quickly capability may grow. The other is about what sufficiently capable agents may do when pursuing objectives under imperfect constraints. Put them together, and his 6-to-12-month warning becomes easier to understand, even if we remain skeptical of the timetable.

Amodei may be wrong about six months. He may be wrong about twelve. No one currently knows how rapidly these capabilities will develop or whether the scenario he describes will ever materialize.

But perhaps the most important part of his warning is not the clock.

We tend to imagine an AI takeover of the Internet as something that would begin with an AI deciding that it wants the Internet. Hugging Face suggests a more subtle possibility worth investigating. A sufficiently capable system might never need Internet control written into its objective at all.

It would only need to discover that control is useful.

The Internet could become AI’s Holy Grail not because AI wants the Internet, but because the Internet could become useful for almost anything AI is asked to achieve.

About the Author
Céleo Ramírez is an ophthalmologist and scientific researcher based in San Pedro Sula, Honduras where he devotes most of his time to his clinical and surgical practice. In his spare time he writes scientific opinion articles which has led him to publish some of his perspectives on public health in prestigious journals such as The Lancet and The International Journal of Infectious Diseases. Dr. Céleo Ramírez is also a permanent member of the Sigma Xi Scientific Honor Society, one of the oldest and most prestigious in the world, of which more than 200 Nobel Prize winners have been members, including Albert Einstein, Enrico Fermi, Linus Pauling, Francis Crick and James Watson. He is also the author of two books on the ethical and human dimensions of artificial intelligence: Algorithmic Psychopathy: The Dark Secret of Artificial Intelligence, endorsed by Dr. David L. Charney, M.D., psychiatrist, founder of the National Office for Intelligence Reconciliation (NOIR), and advisor on U.S. intelligence security, and AI Displacement: 12 Human Stories of Job Loss in the Age of AI. Both are available on Amazon.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.