We have written that numbers are the hardest thing a voice agent hears in Norwegian. That holds – the Norwegian way of saying numbers is a local problem no international benchmark picks up. But on the measurements that exist of real agent conversations, numbers are not the worst-performing category. Names are.
Where the errors actually sit
AssemblyAI publishes error rates broken out by entity type, measured against Pipecat's open benchmark. That benchmark is built from real conversations between callers and voice agents rather than read speech. Four models, same audio:
| Model | Word error rate | Names | Phone numbers |
|---|---|---|---|
| Universal-3.5 Pro Realtime | 6.99% | 16.92% | 3.55% |
| Google Chirp3 | 9.04% | 22.10% | 4.95% |
| ElevenLabs Scribe v2 | 9.76% | 38.03% | 4.78% |
| Deepgram Flux | 15.58% | 39.21% | 10.41% |
Three things are worth stopping on. Names are the weakest category in all four rows. Names run between roughly four and eight times worse than phone numbers – same model, same audio. And the headline figure does not predict it: Chirp3 and Scribe v2 sit less than a percentage point apart on word error rate, and sixteen percentage points apart on names.
The caveat belongs here as well. These are a vendor's own published numbers, and the audio is English-language conversations. The levels do not transfer to a Nordic phone line, and a leaderboard position is never measured on your own line. The ordering and the mechanism do.
An independent measurement points the same way. In a paper presented at EDM 2026, Trinh, He and Whitehill measure Whisper large-v3 on lecture audio and get a 32.3% word error rate on named entities against 7.0% on everything else in the same recordings. That is a factor of 4.6 – on exactly the words that carry the meaning. The same authors report Whisper large-v3 at 28.9% for person names on ConEC, a benchmark built from earnings calls.
Why a name is harder than a number
A number has a check built in. A name has none. Ten possible values per position, a known length, and often a check digit that gives away that something is wrong. A surname has a vocabulary of hundreds of thousands and no redundancy at all. If the model hears "Bakke" where the caller said "pakke", nothing in the string says it is wrong.
The letters that distinguish names are the ones the line throws away. B and D, P and T, V and Z, F and S differ in frequency bands that ordinary telephony strips out on the way in. Those are the classic confusable pairs in English, and they are the same pairs that turn a Nordic surname into an ordinary word: Bakke into pakke, Dahl into tall, Vold into fold.
Nordic names add a layer English does not have. Æ, Ø and Å in Norwegian and Danish, Ä and Ö in Swedish, are spelling choices rather than sounds a model can hear as separate letters. Sæther, Sether and Seter can sound alike on a poor line. Bjørn and Bjorn are the same sound. Løvås, Lovaas and Lovås are three ways of writing the same family. A model trained mostly on English has no strong reason to put those characters where they belong.
The same name has several lawful spellings. Olsen and Olsson. Andersen and Anderssen. Christiansen and Kristiansen. Even a phonetically perfect transcription does not point unambiguously at one row in your customer records.
And a name is the one thing a caller cannot say a different way. "Could you repeat that?" works for a question. For "Kjærstad" it hands the agent the same signal again, slightly louder.





