A man of eighty-two calls a service number one morning. His hearing is roughly where most people's is at that age: he can hear perfectly well that someone is speaking, and he loses every other consonant.
He hangs up after fifty seconds. In the log the call is recorded as abandoned by the caller, and there it stays.
The call below is constructed. None of the five moments in it are. They are the places where something has to have been decided in advance, and not one of them is solved by a better model.
How many people this covers
The Norwegian Institute of Public Health updated its figures in June. Around 19 percent of adults in Norway have some hearing loss when mild and one-sided losses are counted – roughly 800,000 people. About 6 percent have a moderate or greater loss, which usually means clear difficulty making out speech, especially with background noise.
The distribution is what matters for a phone number. Among those over 65, 59 percent have some hearing loss and 23 percent a moderate or greater one. Around 80 percent of all hearing loss in Norway is found in that age group.
If your callers skew over 65 – a municipality, a doctor's surgery, insurance, energy, property management – this is not an edge case. It is one of the largest single groups on the line.
Two filters in series
Age-related hearing loss starts at the top of the spectrum. The high frequencies weaken first, which is why the unvoiced consonants – p, k, t, s, f, h – are the first to disappear.
That is exactly the region the phone line has already thrown away. Ordinary telephony carries roughly 300 to 3,400 hertz, and the sibilants sit partly above that ceiling.
The two filters sit in series. The network takes the top of the signal, his hearing takes much of what is left. A caller of thirty barely notices the first one. He gets the sum of both.
And that is where the asymmetry lies that makes this a design question rather than an audio question: the agent can guess from context, look things up, ask for a repeat. He has only his ears.
The opening
The agent says who it is, who it answers for, and that it is artificial intelligence.
He catches "Hello, you're speaking with", and then a run of syllables.
The sentence isn't bad. The problem is that the two pieces of information that matter most sit in the part of the line where the pace is fastest and he is still working out what kind of call this is. The opening line is the only part of the setup every caller hears – provided it is built so that it can be made out.
What has to be decided: how fast the agent speaks, and whether there is a pause after the greeting. Half a second of silence costs nothing, and it moves the important part out of the busiest stretch of the line.
"Sorry, what?"
This is where most setups part company with a human being.
A person asked to repeat says it a different way: shorter sentence, different words, a little slower. A voice agent usually repeats the same sentence at the same pace, because that is the sentence it has.
It doesn't help. If "confirmation code" was inaudible the first time, it is inaudible the second time.
What has to be decided: that the second attempt is a different phrasing, not the same one louder. And that there is no third attempt – two repeats are a signal, not a state the agent should be left sitting in.
The number
The agent reads out a date and asks for confirmation.
"Fifteen" and "sixteen" differ chiefly on f and s – two unvoiced sounds, both in the part of the spectrum that is thinnest for him and narrowest on the line. He hears that a number arrived. He does not hear which one.
So he confirms. People do. Asking someone to confirm something they did not hear is a politeness trap, not a check.
What has to be decided: that the confirmation is something he can answer, not something he can say yes to. "Could you read the date back to me?" exposes the error. "Is that correct?" hides it.

The handover
He asks to speak to someone. The agent transfers the call, and a summary of the case goes with it.
The human therefore learns what the call is about, but not the thing that decides how the next thirty seconds go: that this caller could not make out the agent. Escalation is the part of the setup that is most expensive to do halfway, and halfway costs here means the conversation begins again at the same pace as the one that just failed.
Fifty seconds
He hangs up before anyone picks up.
In next month's report this is one abandoned call among several hundred, and it looks like noise. The metrics that flatter a voice agent most are the ones that do not distinguish a short call from a failed one.
The remedy requires no new feature. Listen to ten short calls a month – not the longest ones, not the escalated ones, but the short ones where somebody said "sorry, what?" more than once. They are already in the log.
What the five moments have in common
None of them is solved by a better model, and none of them is solved by more volume.
They are solved by five decisions: how fast the agent speaks, what happens on the second attempt, whether a confirmation can be answered without having been heard, what the human is told during the handover, and whether anyone ever listens to the short calls.
Threll.ai builds voice agents in Norwegian, Swedish and Danish. The caller who has the hardest time hearing the agent is usually also the one with the fewest other ways of getting in touch.




