The demo sounded perfect. The agent understood everything, answered quickly, read the email address back without a single error.
Then it went onto a phone number. Now it hears "F" as "S", and it can't tell a Norwegian caller's "Kjell" from "skjell".
The obvious conclusion is that the model isn't good enough. Usually the model is exactly the one that was in the demo. What changed is the audio going into it – and neither of those errors is random. They sit at opposite ends of what the phone line throws away.
The line is a filter, not a pipe
Ordinary telephony encodes audio as what is called narrowband: 8,000 samples per second, across a frequency band running from roughly 300 to roughly 3,400 hertz. Everything above and below that is gone.
The standard is inherited from the analogue era, and it was set low deliberately. It had to carry just enough for one person to understand another, and no more, because bandwidth cost money.
For a human, it is plenty. The ear guesses the rest. This is not a new observation: the reason spelling alphabets exist at all – "S as in Sierra, F as in Foxtrot" – is that s and f become genuinely hard to tell apart once a phone line has removed the high frequencies that separate them.
A speech recognition model gets no such help. It gets a signal. And most models are trained on audio sampled 16,000 times per second – wideband. Give one half of what it is used to and it guesses too, but it guesses worse, and it guesses at things a person never would.
| What the sound carries | Roughly where | Survives narrowband? |
|---|---|---|
| The fundamental of the voice – intonation, pitch accent | 80–250 Hz | No, it sits below the 300 Hz cut |
| Vowels | 300–2,500 Hz | Yes |
| s, sh and other sibilants | 3,000–8,000 Hz | Barely, or not at all |
| The burst in t, k and p | 2,000–6,000 Hz | Partly |
The Nordic languages use both ends at once
English loses something to narrowband. Norwegian, Swedish and Danish lose something at both ends of the band simultaneously, and that is a harder problem.
At the top of the band sit the sibilants. Norwegian separates kj from skj – "kjenner" from "skjenner", the name "Kjell" from "skjell" – and Swedish separates its tje-sound from its sje-sound, as in "tjära" and "skära". Both contrasts are carried mainly by energy above 3,000 hertz. So is the release burst on a final t or k, which is what Danish leans on to keep "tak" apart from "tag". Narrowband cuts straight through that region. For a proportion of callers there is a second filter stacked on top of that one, in exactly the same frequency range.
At the bottom of the band sits the fundamental frequency of the voice. Norwegian and Swedish are pitch-accent languages: "bønner" (beans) and "bønder" (farmers) are told apart by tone alone, as are Swedish "anden" (the duck) and "anden" (the spirit). Danish does the equivalent job with stød, a brief creak in the voice that separates "hun" from "hund". All of it lives in the fundamental, which for most adults sits between 80 and 250 hertz – below the 300 Hz cut. The human ear reconstructs the fundamental from its overtones, so a person hears these distinctions on the phone without difficulty. A model asked to track pitch in a signal whose fundamental has been filtered out has a substantially harder job.
The Nordic languages are therefore unlucky at both ends at once: they use the top of the spectrum to separate consonants and the bottom to separate words, and the phone line takes some of each. It is a related reason why numbers are the hardest thing a voice agent hears – there the problem is how the numbers are counted, here it is acoustics, but they hit the same call.





