What AI Really Hears When Israelis Call
Yesterday morning I put nine seconds of an ordinary Hebrew phone call in front of eight speech-recognition engines and asked each one the same question: what did this person just say?
The sentence was nothing special. Lo meshane, shel ha-bot, hu kvar yodea — “doesn’t matter, it’s the bot’s, it already knows.” A caller talking to an automated answering service, the way a growing number of Israelis now do every week without thinking about it.
OpenAI’s gpt-4o-mini-transcribe wrote it down as Polish: nam się nazywało, jak to. Not a garbled Hebrew guess. A confident, well-formed sentence in a language nobody on the call speaks, handed downstream to the assistant as though it were what the caller had actually said.
That one line is the cheapest available summary of a problem Israeli business owners are, at this moment, being sold their way into.
The leaderboard and the phone line
There is a public Hebrew speech-recognition leaderboard, maintained by the ivrit.ai project, and it is genuinely good news. The open Hebrew model it ranks, whisper-large-v3-turbo, scores a 4.8% word error rate on a manually transcribed podcast conversation, 7.2% on freestyle WhatsApp recordings, 6.5% on podcast material from the Robo-Shaul corpus. Five percent error is roughly the level at which a machine stops being a novelty and starts being infrastructure. Anyone selling you an AI receptionist this quarter is, directly or indirectly, quoting numbers from that side of the world.
Now look at where those recordings come from: podcast episodes, volunteer voice notes, crowdsourced read speech, broadcast archives. All of it wideband audio, recorded into a decent microphone in a room, at 16,000 samples a second or better.
An Israeli business phone call does not arrive that way. It arrives over the ordinary telephone network as G.711 μ-law at 8,000 samples a second, the format every landline and every carrier hand-off in the country has used for decades. Half the audio bandwidth a modern model expects simply is not present. It was never transmitted. And Hebrew, unluckily, keeps a great deal of its meaning in exactly the sounds that narrow band treats worst: the soft consonants, the difference between shin and sin, the short unstressed vowels that separate one verb form from another.
So I ran the measurement the leaderboards cannot run for me. I took real recorded calls from a Hebrew answering service, cut out the caller’s turns, wrote down by hand what each caller actually said, and scored eight engines against that reference.
The numbers
Against nine hand-checked caller turns, forty-four words of real Israeli phone Hebrew, the results were not close. Soniox v5, given the conversation’s context, got six of the nine turns word-perfect and finished at a 13.6% error rate. Without that context it slipped to 18.2%. The ivrit.ai model, the same whisper-large-v3-turbo the leaderboard ranks, came in at 25.0%. Then the drop: OpenAI’s gpt-4o-transcribe at 61.4%, its mini version at 63.6%, Gemini 3.5 Flash at 70.5%, Gemini 3.5 Flash-Lite at 93.2%.
Forty-four words is a small reference set and I will not pretend otherwise, so I scored the same engines again against a larger, independently transcribed set of twenty turns and seventy-five words. The absolute figures moved, as they should. The order did not budge: Soniox 24%, OpenAI 53%, Gemini 65%, Flash-Lite 77%.
Look again at the third figure in that first list. The ivrit.ai model that scores 4.8% on podcast audio scored 25% on my phone line. Same model, same language, roughly five times the error, purely from the channel the audio travelled through.
And the two engines whose names appear in almost every AI product announcement of the past year lost more than half of every sentence. Gemini 3.5 Flash took a caller giving a name, al shem Maurine, “under the name Maurine,” and produced an English aside instead. At 93.2%, Flash-Lite is not transcribing Hebrew at all. It is generating plausible Hebrew-shaped text next to an audio file.
Why the famous models lose
This is not incompetence on anyone’s part. The large multimodal models are attempting something harder: taking audio and producing meaning in a single pass, across a hundred languages, with Hebrew occupying a rounding error of the training diet. A narrow streaming recogniser that does one job, and can be locked to Hebrew so hard it is forbidden to answer in any other language, has an unfair advantage at that job. It should win. What surprised me was the size of the gap once the audio got bad.
It also explains a failure mode Israeli owners are quietly living with. When a general model is unsure, it does not fall silent. It produces its best guess in whatever language the acoustics happen to resemble, with full confidence and no flag attached. The Polish sentence was not an error message. It looked exactly like a successful transcription.
What this actually means if you are buying
The practical conclusions are cheap enough to act on this week.
A demo proves nothing. Every AI voice demo you will be shown runs on the clean side of that 8,000 Hz wall: a laptop microphone, a quiet room, a vendor who speaks clearly because he knows what the system expects. Insist instead that the system answer your own line, and that somebody call it from the oldest handset in the office, on speaker, from a moving car, with a customer’s accent rather than yours.
Then ask what the system does when it is unsure, and treat a vendor with no answer as having answered. The most useful change I shipped this week was not a better model. It was a confidence gate: when the average certainty across a caller’s turn drops below a set threshold, the assistant stops guessing and says, in Hebrew, “sorry, I didn’t catch that, could you say it again?” A machine that admits it misheard is worth more than one that is fluent and wrong, because the first costs a few seconds and the second costs the customer.
And be sceptical of any Hebrew accuracy figure quoted without the channel it was measured on. It is the single most important number in the sentence and it is almost always missing.
None of this is an argument against putting AI on the phone. The service these recordings come from answers calls that would otherwise ring out, and I should disclose that I build and run it, under the name Onim Li (onimli.com), which is also why I have the recordings to measure at all. Hundreds of calls a week now get answered instead of lost, and the customers on the other end mostly cannot tell.
But the reason it works is not that a frontier model solved Hebrew. It is that somebody sat down with a hand-written reference and a pile of ugly 8 kHz audio and picked the boring engine that happened to win. Most of the value in this market comes from that kind of measurement rather than from anything a demo can show, which is roughly what boutique automation studios such as Achiya Automation spend their days doing.
The frontier labs will close this gap eventually. Hebrew will stop being a rounding error in the training data. Until then, the distance between a vendor’s demo and a Tuesday morning in your shop is the part of the product nobody is quoting you a number for.
So here is the number I would ask for, and the one I would genuinely like to hear from other Israeli owners running these systems: of the calls your AI answers in a week, what percentage end with the caller repeating themselves? Not your vendor’s accuracy claim. Your repeat rate. If you know yours, tell me. I have not yet met an owner who has measured it.

