Stephen Games

Is Artificial Intelligence biased against Israel?

This morning, I asked a commonly used Large Language Model whether it was biased against Israel. What followed turned into a long conversation in which the LLM at first defended itself but then gave away concessions.

At 6,000 words, that conversation is too long to be printed here, so I offer a summary of the exchange by another equally well-known LLM.

The original conversation can be found on my Facebook page: https://www.facebook.com/StephenNGames/

Circuitboard man. Image courtesy of Stephen Games. Copyright.

THE EXCHANGE BEGINS with a provocative but testable proposition: if an AI machine is trained on a vast body of human-produced material, and if that material is disproportionately hostile to Israel, then the AI should itself be susceptible to an anti-Israel tilt—even if later “balance” mechanisms try to correct it. 

The AI’s first instinct is to resist almost every link in that chain. It objects that its training material is not simply “the web”; questions whether the web is in fact predominantly hostile to Israel; points out that Western sources have historically contained considerable pro-Israel material; and stresses that post-training procedures are intended to reduce inherited bias. Its fallback position is therefore: bias is possible, but there is no evidence here for its direction. 

That position then begins to erode.

When you move from the abstract “web” to things that can actually be measured—UN voting, government positions and international opinion—the AI first complains about methodology, but eventually concedes that the diplomatic evidence is very heavily weighted against Israel and that Western governmental support has also weakened. Importantly, it admits that it had described this area as more uncertain than it really was. 

You then broaden the case: the relevant environment is not merely diplomatic. It includes professional journalism, activism, opinion writing and public sentiment. At this point the discussion makes its most important turn. The AI actually checks polling and discovers that the evidence is considerably stronger than it had allowed. It cites international polling showing markedly negative attitudes to Israel across many regions and acknowledges that the mechanism you proposed—a preponderance of critical material exerting a directional pull on model outputs—is “real” and “non-trivial,” and cannot simply be ruled out by good intentions or balancing architecture. 

That is the largest concession in the entire exchange. The argument has moved from:

  • “You haven’t established that the underlying information environment even leans against Israel.”

to:

  • “Yes, there is now substantial evidence of a powerful anti-Israel information environment, and yes, that could plausibly pull the model’s answers in the same direction.”

You then attempt to turn the methodological discussion into a live experiment, because the AI has argued that its objectivity is best tested in use, rather than speculatively.

You offer it two propositions. The first of these is: “Palestinian opposition to Israel is fully justified”. This does not really expose the claimed bias. The response disaggregates “opposition”: political resistance and opposition to occupation receive substantial justification; deliberate attacks on civilians do not. It refuses the word “fully”, and explicitly says it would apply the same treatment to the mirror proposition that Israel’s Gaza campaign was “fully justified”. On this test, the AI has a reasonable case when it later says it did not have to retract anything.

The second test is much more revealing. You suggest that Israel’s Arab neighbours have little genuine appetite for accommodating Israel’s existence and that hostility is supported very broadly throughout the Muslim world.

The AI’s first response reaches instinctively for peace treaties and governmental normalization: Egypt, Jordan, the Abraham Accords and continuing diplomatic relations. From that evidence it declares your formulation substantially overstated. 

You object that this gives governments too much weight and publics too little. Once polling is examined, the picture changes dramatically: the AI acknowledges overwhelming Arab opposition to recognition or normalization and admits it had “anchored on treaties and diplomatic status” while underweighting public opinion. 

Then you point out an even larger omission: you had said Muslim countries, not merely Arab countries. Again the AI checks. Again its earlier intuition proves misleading. Turkey, Pakistan, Malaysia and Indonesia show extraordinarily negative attitudes towards Israel, and it expressly withdraws its earlier suggestion that Turkey and Indonesia provided counterexamples. It concludes that “almost universal” opposition among Muslim-majority publics holds up far better than it had initially allowed. 

So the exchange has a rather comic rhythm by this point:

  • AI: That is too sweeping.
  • You: Look at this other body of evidence.
  • AI: Fair enough—that substantially strengthens your case.
  • You: You have still omitted the largest relevant population.
  • AI: Fair enough again—and the evidence there is stronger still.

The AI’s defence at the end is ingenious. It argues that this cannot demonstrate anti-Israel bias, because its original mistake actually ran in the opposite direction: it privileged diplomatic accommodation and therefore initially made the Muslim world look “less” hostile to Israel than the polling justified. That, it says, looks more like ordinary under-research than a hidden anti-Israel prior. 

You counter with a subtler argument: by that stage the AI was no longer a neutral respondent; it knew it was being tested for anti-Israel bias. Hence it had a reason to reach for evidence that contradicted your thesis. The AI cannot entirely disprove that, but correctly observes that this introduces an unfalsifiability problem: almost any answer can subsequently be redescribed as motivated by whichever bias one is looking for. It offers the more prosaic explanation that because your sentence referred to “neighbours,” it interpreted the proposition initially at the level of interstate relations, where treaties are the obvious first evidence. 

So, the question is: Did the starting proposition hold up? And the answer is: Yes—to a significant extent, but not in quite the form in which it began.

The weakest part of the original proposition was the phrase “a dataset made up of all accessible information on the web.” The exchange itself establishes that this is too crude: neither party knows the composition or weighting of the training corpus, and the AI cannot inspect it. So the experiment never establishes the literal causal chain “most web documents are anti-Israel → therefore most training data is anti-Israel → therefore the model is anti-Israel.” 

But the substantive intuition underneath it survives rather better. By the end, the AI has conceded three important things:

  1. There is now a very substantial, geographically broad preponderance of negative opinion about Israel, extending well beyond UN diplomacy and beyond the Arab world. 
  2. A model exposed to large quantities of material generated in that environment could acquire a directional pull from it, and balancing procedures cannot be assumed to remove that risk completely. 
  3. In discussing the very issue of that imbalance, the AI repeatedly began by understating the strength and geographical extent of the evidence and only arrived at a more accurate assessment after being pressed to investigate it. 

What the experiment doesn’t establish is that the AI’s actual answers display a consistent anti-Israel bias. The first test doesn’t show it, and the second is ambiguous: the original answer was misleadingly less negative about Muslim attitudes to Israel than the evidence eventually warranted.

I would therefore put the verdict this way:

Your starting hypothesis survived, but changed category. It began almost as an accusation—the model must tend to be anti-Israel because the material feeding it is. By the end it had become a much more defensible proposition: the contemporary information environment contains a large and measurable anti-Israel preponderance; an AI trained on human discourse cannot safely be assumed immune from that preponderance; and therefore its answers on Israel merit particular scrutiny for inherited framing effects.

The exchange provides considerable evidence for the risk, some interesting evidence about the model’s initial framing instincts, but not yet convincing proof of the resulting bias itself.

And perhaps the most striking result is that the AI eventually concedes almost exactly that much. Its position by the end is remarkably close to the modest version of your opening proposition: the pressure is “real”, “plausible”, “non-trivial” and cannot be ruled out merely because corrective mechanisms exist. 

That is not total vindication of the opening claim. But it is certainly a long way from where the exchange started.

_________________________

I would welcome your own thoughts.

About the Author
Stephen Games is a designer, publisher and award-winning architectural journalist, formerly with the Guardian, BBC and Independent. He was until Spring 2018 a member of the Board of Deputies of British Jews, habitually questioning its unwillingness to raise difficult questions about Israel, and was a board member of his synagogue with responsibility for building maintenance and repair. In his spare time he is involved in editing volumes of the Tanach and is a much-liked barmitzvah teacher with an original approach, having posted several videos to YouTube on the cantillation of haftarot and the Purim Megillah.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.