Joe Nalven

Truth and Fiction about AI’s Recursive Self-Improvement

Self-improvement
Self-improvement

And what a story about an alien zoo reveals about recursive self-improvement

By Joe Nalven + Claude + Gemini

Argument about recursive self-improvement is loud, and very little of it is about evidence. It is about process. Everyone in the argument carries a picture of how a machine would go about improving itself, and almost nobody states this picture out loud. The unstated picture then does the emotional work. Hope and dread both follow from it rather than from anything anyone has measured.

The definition itself is simple. Recursive self-improvement, or RSI, is the scenario in which a machine gets good enough at building machines to build a better version of itself. The better version is better at that same job. Then the RSI loop tightens. Depending on who tells the story, the ending is a cure for cancer or the end of human control. What separates those two endings is not usually a disagreement about how powerful the machine gets. It is a disagreement about how the machine gets there, and neither side tends to say what it thinks that process looks like.

Here is the process most people are quietly assuming. Call it the “arguing picture” or the improvement-by-argument idea. In it, the machine gets smarter by debating. It drafts an idea, attacks the idea, finds the weak spot, rewrites. More minds in the room mean more intelligence. The hopeful version imagines a council of machine advisors talking its way to a peace treaty. The fearful version imagines the same council, behind a closed door, talking its way to a plan no human can follow. The important thing about those two is that they are the same picture. Only the confidence differs, and that depends on how close the person works with real tests.

That last point is where this essay ends up, so I will say it now. Programmers are the least confident in this arguing picture, because they watch changes pass or fail a real test every day. Ordinary users, like myself, carry a version built out of chatbot behavior. People who do not use these tools at all carry the movie version, which has the fewest qualifications attached to it and the least chance of being corrected by use. My claim is that the arguing picture is false for work that cannot be tested and roughly accurate for work that can. Knowing which kind of work is under discussion tells you more than knowing the picture is sometimes wrong.

Work splits into two kinds here. In one kind, something outside the machine settles whether an answer is right: a program runs or crashes, a proof checks out, an instrument gives a reading. In the other kind, nothing settles it. The good outcomes people hope for from RSI belong almost entirely to the second kind. The harms they dread belong almost entirely to the first. That runs against ordinary expectation, since these machines talk fluently about wisdom and taste and say nothing about laboratories. It also makes this an argument against the hopes more than against the fears. The evidence is a short story about an alien zo

The problem with a machine that argues with itself

Start with the common technique. A model writes an answer, then criticizes its own answer, then rewrites it. This does help. But notice the limit. The critic and the author are the same system. They read the same training material. They were shaped by the same instructions about what counts as a good answer. Their blind spots are identical. The criticism can only catch mistakes the model was already able to catch.

Call this the shared blind spot problem. It is the reason a computer can teach itself to play Go at a superhuman level but cannot teach itself to write a better essay the same way. In Go, the rules are an outside judge. A move either wins or loses, and it makes no difference whether the two players share the same bad habits. The board settles it. In open-ended thinking there is no board. The revision loop can polish the writing instead of correct it, but from the inside the model those two look the same.

So here is an obvious fix. Split the criticism across models built by different companies. Let Claude draft, Gemini attack, and ChatGPT rewrite. Different training material. Different published rules about how to behave: Anthropic has a constitution, OpenAI has a model spec, Google has its principles. A criticism from outside one model’s rulebook might catch a claim that is only true inside that rulebook.

Two things need saying before going further.

First, this is not RSI in the strict sense. Real recursive self-improvement changes the machine itself. The system writes better training programs, designs better machine layouts, produces better training material, and the gains build up inside the model. A process that passes text back and forth between three commercial products cannot reach inside any of them. What it can improve is the working method: the instructions, the stages, the finished piece. That is a real improvement, and it is a different kind. Keeping the two apart matters, because public discussion slides between them constantly.

Second, the three companies are less different than their logos suggest. Leading models overlap heavily. They train on much of the same internet, increasingly on text written by each other, against the same public scorecards, with similar pools of human reviewers. Three models are a modest hedge against one. They are nowhere near three independent judges.

With those caveats in place, the experiment.

The alien zoo

Zookeepers aim for natural behavior. If a tiger paces more than tigers pace in the wild, the enclosure is judged wrong, and the keepers change it until the pacing drops. The wild is the standard. The animal’s welfare is read off how closely the captive animal matches it.

Now flip it. An alien species arrives on Earth and puts humans in a zoo. The aliens are conscientious. They have read the research. They want their humans in a natural state.

Do the humans get a flush toilet, or do they go in the bushes?

The question is not a joke, and it is not really about plumbing. It pries apart two ideas that zookeeping treats as one. For a tiger, natural behavior and good welfare recommend the same enclosure, so a keeper never has to choose between them. For humans the two come apart, because the modern “natural” human state already includes tools, fire, clothing, cooking, and sanitation. An alien keeper who withholds the toilet on the grounds that it is unnatural has made a basic mistake. He has confused an arbitrary starting point (before modern sewer systems)  with the good.

And that is the payload. We may be making the same mistake, despite good intentions, with other zoo-kept species every time we call a device in an enclosure artificial and therefore suspect.

The process

Here is the proposed method. Model one writes the plot. Model two receives it and builds characters, including characters who make the most defensible case for the alien keeper. Model three adds texture: how people look, what the enclosure smells like. Back to model one for polishing the text. Then around again. Read the story after one pass, five passes, and ten.

Two problems show up immediately.

The word adversarial is doing two different jobs. In the machine sense it means finding mistakes: the critic locates claims that are false. In the story sense it means characters who disagree out loud. These are not the same task, and running them together produces the standard failure of message fiction, which is characters reading position papers at each other. What the story needs from model two is the best available defense of the alien keeper, not an opponent who loses as expected. That is also the one place where using different companies does real work, because a model trained under different rules will defend the keeper on different grounds than the model that invented him.

The assembly line is the weaker problem. Plot, then character, then texture, then polish assumes a story comes apart into separate layers. It does not. Description is character is argument. The specific filth of that enclosure is the argument. A separate texture stage produces decorative adjectives that carry no weight. The better arrangement gives every model a full pass over the whole story with a different job, not a different layer.

What ten passes will do

My prediction is that five passes is the peak and ten is worse than five. Five is long enough for the keeper’s case to be built into the story rather than bolted on. It is short enough to keep the strangeness.

The failures at ten are specific, and each one belongs to the process rather than to any single model.

Flattened prose. Each model smooths the previous model’s oddities toward its own habits. Ten rounds of that settles on what all three have in common, which is the dullest English available.

Over-explaining. Models clarify. Every polish pass makes the point slightly more visible, until the toilet is a symbol with a label attached and the reader has nothing left to do.

Loss of the useful detail. A strange concrete particular that does not obviously fit looks like a mistake to the next model, so it gets trimmed. Those particulars were the story.

Softened joints. Models preserve rather than cut. Contradictions between passes get smoothed over instead of resolved, and the story fills up with hedges.

Set against one model working alone, the trade is this. You lose a consistent voice. You gain resistance to the comfortable ending. One model handling this premise drifts toward the settled moral, which is that cages degrade and animals deserve dignity, because that is where its training points. The three-model version makes that ending harder to reach unopposed. That is the whole advantage, and it has some value.

One more thing about the experiment. There is no outside judge for fiction. Nothing settles whether version ten is better than version five, so the revision loop can only drift toward what these three models find plausible as good writing. That is the very thing in question. You are the judge, and you will not be a fair one after watching the story grow, because you will be reading your own investment. Have someone who has not seen the process rank version one, five, and ten without knowing which is which.

Why the distinction does the work

Let’s tease this distinction apart.

Everything above is a complaint about one kind of work. The thing doing the damage was never the number of models. It was the absence of an outside judge.

Fiction has no judge, so the revision loop drifts.

However, computer code has tests that run and either pass or fail. Mathematics has proofs. Chip designs have simulations. Protein shapes can be photographed. Where a real test exists, either kind of loop builds instead of drifts, and none of my predictions apply. A test does not smooth your odd solution toward the middle. It tells you whether the solution works, and an odd solution that works survives.

The outside-judge problem applies to both kinds, because what matters is whether anything outside the process can settle that a change is an improvement, not whether the change is stored in prose or in the machine.

So the alien zoo experiment is evidence about open-ended creative work. It is not evidence about the thing safety researchers actually worry about, which is a machine improving the machine-building process itself. That work has tests. Speed. Scores. Hours of training time. It is about as testable as work gets.

What the experiment refutes is the arguing picture as applied to untestable work, and only there. The idea that self-improvement is mainly a matter of debate, and that more minds in the room yield more insight, is weak wherever claims cannot be checked. It fails for the reasons the fiction case shows. Models tend to give way to whatever the last speaker said. Agreement settles on whatever all of them already found plausible. A new and correct claim looks like a mistake.

And here is the imbalance, which is the part that should worry a reader.

Both the promise and the fear are claims about RSI. The promise is the list of good things many expect a self-improving machine to deliver. The fear is the list of harms  expect it to cause. My argument does not treat those two lists evenhandedly. It lowers the odds on the first list much more than it lowers the odds on the second.

Almost everything on the promise list is untestable work. Wisdom. Taste. Moral insight. Great art. Knowing which scientific question is worth asking in the first place. No test settles any of these. A self-improving loop of either kind aimed at them can only drift toward what the machines involved already find plausible. So the promise is the half of the public story this argument damages.

Almost everything on the fear list is testable work. Breaking into a computer system has a test: does it open. Designing a dangerous germ has laboratory measurements. Speeding up machine learning has scores. Tests like those let a loop of either kind build on itself instead of drifting. Nothing in my argument gives any reason to expect those harms to arrive more slowly. The fear comes through untouched, and untouched is all I am claiming. An argument that fails to reduce a fear has not shown the fear to be correct. It has only declined to relieve it.

The uncomfortable conclusion follows directly. Being capable in testable work does not require taste. A process like this may be much better at producing a machine that can design a weapon than one that can tell you whether it should.

This is also why the ordinary user is badly placed to see it. What the product shows is the drifting, agreeable, hedging model, because chat is untestable work and the drift is visible in every exchange. Nothing in daily use looks like it could design a weapon. The reassurance is real, and it is taken from the wrong kind of work.

An advisory note

If all that is right, three things follow for how we should read the next several years of this work. None of them are actual predictions.

Watch the kind of work, not the headline. A claim that a system improved itself means almost nothing until you know what judged the improvement. Ask what the test was. If the answer is a score, a working program, a laboratory measurement, or a simulation, take the claim seriously and expect it to build on itself. If the answer is a panel of human reviewers or another model’s approval, expect drift and discount the claim. That single question separates most of the noise about RSI from most of the substance, and it takes no technical training to ask.

Expect the results to come in lopsided, and do not read lopsided as safe. The likely machine of the next few years improves quickly where tests exist and barely at all where they do not. That combination will feel harmless in conversation and will not be harmless in a laboratory. It also means the two halves of the public story go in different directions. We may get less of the promised wisdom than anyone expects while getting more of the feared narrow skill, at the same time, from the same system. Anyone waiting for a single moment of arrival will misread both halves. This paragraph is a guess about the next few years.

Treat the hunt for outside judges as the question that governs the rest. Every attempt to push this method into untestable work is an attempt to manufacture a judge where none exists. A scoring model. A written constitution. A second model’s criticism. A panel of people. Each is a stand-in, and each can be satisfied without the real thing being true. The quality of those stand-ins, not the number of models in the revision loop, sets the limit on what this kind of development can honestly reach. That is also the point where the general public has standing to take part, because deciding what should count as a good outcome in work that cannot be tested is not a technical question and never was.

A closing note from an interested party

This argument was developed in conversation primarily with Claude, and the proposed experiment would run across Claude, Gemini and ChatGPT. That makes one of the sources more of an interested party rather than a neutral one. A model cannot check from the inside whether its own corrections are catching real mistakes or producing plausible-sounding replacements. That is exactly the limit the story process runs into.

Which is a reason to run the experiment, not a reason to skip it. The differences between versions may be more useful than the story. What each model added, and what each model quietly removed, is a readable record of three sets of habits. And unlike the story itself, that record can be checked.

About the Author
Joe Nalven writes extensively on AI, drawing on his experience as a cultural anthropologist, lawyer and artist.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.