When We Demand More of AI Than We Demand of Ourselves
By Joe Nalven + ChatGPT + Claude
I recently listened to a discussion about how artificial intelligence (AI) might change science. Much of it was familiar. AI may speed up discovery. It may find patterns in data that people cannot see. It may suggest ideas that scientists would not have thought of. But the speakers kept returning to a set of worries. If an AI system makes a discovery that no human can explain, who is responsible for it? If no one can see how the AI reached its answer, should scientists trust the answer? If researchers come to rely on AI, will they lose the ability to think for themselves?
These are fair questions. But they led me to a different question. Are we asking AI to meet standards that human scientists often fail to meet?
The answer is not a simple yes or no. The real problem is that people mix together several different standards and talk about them as if they were one. The first is whether a scientific claim is correct. The second is whether a human can understand why it is correct. The third is who is morally responsible when something goes wrong. The fourth is how institutions, such as universities, companies and government agencies, assign that responsibility. The fifth is whether we can see how a technology produced its result.
These are separate questions. They need separate answers.
Scientists Do Not Always Understand Their Own Results
Science is often described this way. A researcher forms a hypothesis, which is a proposed explanation that can be tested. The researcher runs an experiment, understands the results, and explains to others how the evidence supports the conclusion.
That is a good description of what science aims to be. It is not always a description of what scientists actually do.
Scientists make mistakes. They misread their own data. They build theories that later prove wrong. They rely on specialized instruments and statistical methods that they did not design and may not fully understand. Sometimes they discover that something happens before they know why it happens.
A scientist can honestly say, “I don’t know why this happens, but I can show that it happens.” That statement does not cancel the discovery. The history of science includes many cases where people observed something long before anyone could explain it. Aspirin, for example, was sold and used for about seventy years before scientists worked out how it reduces pain and inflammation.
What matters is whether other people can check the claim. This point becomes important when we talk about AI.
Human Thinking Is Also Hard to Inspect
People often call AI a “black box.” The phrase means a system whose inputs and outputs we can see, but whose inner workings we cannot. An AI system may give an answer without giving a human-style explanation of how it got there.
In some situations this is a serious problem. If an AI system recommends a medical treatment, flags a chemical as dangerous, or proposes a scientific idea, researchers have good reason to want to know why.
But it is a mistake to assume that the human mind is the fully visible alternative. It is not. People make judgments all the time without being able to trace every step. An experienced scientist may sense that an experimental result is wrong before she can say what is wrong with it. A mathematician may find a solution before being able to explain where the idea came from. A doctor may recognize a pattern from years of practice that she cannot yet state as a simple rule.
We do not conclude from this that human knowledge is worthless. Instead, science relies on something stronger than any one person’s ability to explain their own thinking. It relies on outside checking. Other people can repeat the experiment. They can examine the data. They can challenge the methods. They can propose other explanations. And the result can fail those tests.
So science does not depend mainly on whether an individual scientist can explain his or her own thought process. It depends on a shared system, carried out by many people and institutions, that is able to catch mistakes. AI could be placed inside that system of checking. That makes more sense than judging AI against a standard of perfect human self-understanding that no human actually meets.
Testing a Result May Matter More Than Explaining It
Suppose an AI system finds a connection between two biological processes that no one had noticed before. Researchers cannot figure out how the AI arrived at this idea.
There are two possible reactions. The first is: “We cannot follow the AI’s reasoning, so we cannot accept the finding.” The second is: “We cannot follow the AI’s reasoning, so let’s test the finding.”
The second reaction fits how science is supposed to work. Suppose other laboratories, working independently, get the same result. This is called reproducing, or replicating, the result. Suppose the measurements are sound, other possible explanations have been ruled out, and repeated efforts to prove the finding wrong have failed. Then the finding may be scientifically valuable, even if no one can retrace how the AI found it.
This does not mean explanation is unimportant. Knowing the mechanism, meaning the step-by-step physical or biological process that produces an effect, is very valuable. It helps scientists predict new cases and spot errors. But two things need to be kept apart: showing that a claim holds up under testing, and feeling that we personally understand it.
A claim does not become true because a person understands it. A claim does not become false because a person does not understand it.
Mathematics faced this question almost fifty years ago. The four color theorem states that any map can be colored with only four colors so that no two neighboring regions share a color. Mathematicians suspected this was true for over a century but could not prove it. In 1976, two mathematicians, Kenneth Appel and Wolfgang Haken, finally proved it by using a computer to check nearly two thousand separate cases. No person could check all of those cases by hand. Many mathematicians were uneasy. A proof was supposed to be something a human could read and follow. Over time, the proof was accepted anyway, largely because other researchers checked it again, including in 2005 with a computer program designed specifically to verify proofs. The mathematical community accepted a result that no single person could fully follow, because it held up under independent checking.
When Testing Is Not Enough
Testing is the strongest check science has. But it does not work equally well everywhere.
Some claims are easy to test. A chemical reaction can be run again in another lab. A new material can be measured again with another instrument. For claims like these, “let’s test it” is a good answer.
Other claims are much harder to test. A diagnosis for one patient cannot be repeated on a thousand copies of that patient. A claim about how the climate or the economy will change cannot be checked by running the Earth a second time. Some experiments cost so much that only one or two facilities in the world can perform them. And many claims about the past can only be checked against the evidence that happens to survive.
Human science already struggles here. In fields such as psychology and medicine, many published findings have failed when other researchers tried to repeat them. Testing works only when people actually do it, and they often do not.
AI adds three further problems. First, the same AI system may give different answers to the same question, so it is not always clear which answer is being tested. Second, if the checking is also done with AI systems trained on similar data, two systems agreeing may not count as independent confirmation. They may share the same blind spots. Third, AI can produce ideas far faster than laboratories can test them. Deciding which ideas deserve testing then becomes a serious problem of its own.
There is also a gap between confirming a result and confirming an explanation. A lab can reproduce the fact that two things tend to occur together. That does not show that one causes the other.
This leads to a practical rule. The harder a claim is to test, the more we should demand an explanation of how it was reached. When testing is strong, we can accept a result we do not fully understand. When testing is weak, understanding the reasoning is one of the few checks we have left. The standard for AI should depend on which kind of claim is being made.
Responsibility Is a Different Question
The case for holding AI to special rules is stronger when we move from knowing something to acting on it.
A scientist works inside human institutions. A university, a drug company, a government lab, or a research team can assign responsibility for decisions. A person can be questioned, disciplined, sued, corrected, or removed from a project. An AI system cannot currently be held responsible in any of those ways.
Suppose an AI recommends an experiment, and the experiment causes harm. Saying “the AI did it” does not settle who is responsible. People chose to use the system. People chose the data it learned from, set its goals, interpreted what it produced, and decided whether to act on it.
So the right standard may not be that the AI must have a human conscience. The right standard may be that people must stay responsible for the AI systems they choose to use. That is a very different requirement. It is a rule about people and organizations, not about the machine.
Losing Skills Is Not a New Worry
Another concern is that scientists will come to depend on AI and lose skills they once had. This is sometimes called “deskilling.”
This is a real possibility. If researchers regularly ask AI to form hypotheses, analyze data, write computer code and interpret results, they may slowly lose the ability to do those things on their own.
But this problem did not start with AI. Scientists have always used tools that extend what they can do. Computers took over huge amounts of hand calculation. Statistical software took over calculations that researchers once did with pencil and paper. Databases took over much of the work of finding information.
The question has never been whether a scientist personally performs every step. The question is whether the scientist understands enough about each step to notice when something has gone wrong.
That may be the main challenge AI creates for education. We do not need to require scientists to redo by hand everything an AI does. We do need to require them to know enough to question the AI’s result and check it.
Calculators did not make knowing math unnecessary. But a person who cannot tell that the wrong equation was typed into the calculator has a problem. AI may create the same problem on a much larger scale.
We May Be Judging AI More Harshly Than We Judge Ourselves
There is an irony here.
We worry that AI might make discoveries that humans cannot explain. Yet modern science already depends on so much specialized knowledge that no single scientist understands all the work a given conclusion rests on.
We worry that AI may be biased. Human institutions are clearly capable of bias. We worry that AI may reach false conclusions. People reach false conclusions all the time. We worry that AI may be overconfident. Human experts can be very confident and very wrong.
None of this means AI should get a free pass. It means the comparison between AI and humans should be made honestly.
If the standard is “AI must never make mistakes,” no one can meet it, human or machine. If the standard is “AI must be easier to inspect than a human mind,” we need to say why, and what kind of openness we mean. If the standard is “claims produced with AI must be open to independent testing,” that is a reasonable scientific requirement. If the standard is “people must remain responsible for important decisions that involve AI,” that is a rule for institutions, not a requirement about the technology.
AI May Force Science to Separate Two Kinds of Knowing
There is a larger possibility. AI may push science to separate two things we usually treat as one: knowing that something is true, and understanding why it is true.
These often come together. But they do not have to. AI might find relationships that humans would never have thought to look for. It might identify molecules, mathematical patterns or physical arrangements that experiments can later confirm, even though no human can follow the path the AI took to find them.
That would not be a failure of science. It could be an expansion of the methods science uses.
A famous case from the game of Go shows what this looks like. In 2016, an AI program called AlphaGo played Lee Sedol, one of the best Go players in the world. In the second game, AlphaGo made a move, now known as Move 37, that shocked the expert commentators. Several thought it was a mistake. The program’s own calculations suggested that a human player would almost never choose it. Yet the move helped AlphaGo win the game, and professional players later studied it as a new idea about how Go can be played. The move was judged good because of its results, not because anyone could explain in advance why the program chose it.
Go has one advantage that science usually lacks. The game ends, and someone wins. Most scientific claims never receive a verdict that clear. That is exactly why the question of testing matters so much.
A scientist working with AI may be in a similar position, receiving a result from a source whose reasoning cannot be inspected. The result would be useful only if testing confirmed it. The scientist might never learn how the result was found.
So the goal is not to make AI more humanlike before letting it contribute to science. The goal is to build scientific institutions that can test, challenge and make use of discoveries produced by a kind of intelligence that is not human.
A Standard That Applies to Both
Back to the original question: should AI be held to higher standards than humans?
In some ways, perhaps yes. Powerful technologies create new risks. A system that makes thousands or millions of decisions may need safeguards that a single person does not.
But we should reject a version of the question in which AI must be perfectly open, perfectly reasoned and perfectly accountable, while humans are allowed to make mistakes.
A better standard asks the same five questions of any claim, whether it comes from a person or a machine. Can the claim be tested? Can the evidence be examined? Can others independently challenge the result? Can errors be found and corrected? Can we identify who is responsible for important decisions?
Science can reasonably ask these questions of both humans and AI. And where a claim cannot be tested well, the demand for a clear account of how it was reached should rise, whether the claim comes from a person or a machine.
The irony is that AI may not be lowering the standards of science at all. It may instead be showing us that some standards we thought science always met, such as complete understanding, fully visible reasoning, and individual responsibility, were never met as consistently as we believed.
So the most important question may not be whether AI can meet human standards. It may be whether science is prepared to work with a partner that does not think the way humans think.
That is a larger challenge than building an AI that explains itself better. It is a challenge to how science decides what counts as knowledge.

