Accessibility

“Voice is the most inclusive interface.” Half of that is true.

The argument for letting an agent answer the phone is often that the phone is the one channel everyone can use. That holds for more people than you would think – and not for a group that almost never makes it into the calculation.

A woman in her sixties in a light cardigan sits alone at a small oak desk in the corner of a living room, half turned towards the window, holding a mobile phone away from her ear while her other hand rests beside an open notebook; flat daylight from the window on the left and a low bookshelf along the wall behind

The claim turns up early in almost every project, and it rarely gets challenged: an app asks you to find it, log in, hit the right button and read the small print. The phone only asks you to speak. Everyone can speak.

The first half of that is true, and it matters more than critics of voice agents tend to concede.

Who the phone genuinely opens the door for

For someone who is blind or partially sighted, a phone line is an interface with no screen to navigate. For someone with dyslexia it is an interface with no form. For someone without a smartphone, or whose national e-ID has never quite worked, it is the only way into a business that has moved everything else online. And for someone speaking a second language, it is often easier to explain a problem in their own words than to find the right category in a dropdown.

This is not a small group. In practice the phone is the fallback for every case self-service does not fit – and an agent that answers on the first ring, around the clock, makes that fallback markedly better than a queue with hold music.

So far the claim holds.

Where it stops holding

“Everyone can speak” is a claim about the person. It says nothing about the machine.

A speech recognition model hears well whatever resembles what it was trained on, and worse everything further away. We have written before about the mild cases of that: the phone line throws away a good deal of the signal before the model ever gets it. Further along the scale it looks entirely different.

A research team at Stony Brook University tested Whisper against the TORGO dataset last autumn – recordings of eight people with dysarthria, a motor speech disorder caused by damage to the nervous system, after a stroke, in cerebral palsy or in ALS. With no adaptation at all, Whisper-small got 71 per cent of the words wrong on average.

The average is not the interesting part. The spread is. Of the eight, the one least affected had an error rate of 7 per cent – roughly what anyone else would get. The one most affected was above 100 per cent, meaning more errors than there were words to get wrong. Same model, same task, same diagnosis, and two completely different services.

The study is small, and the authors say so themselves: eight people, English speech, one model family. But the mechanism is not English. It applies equally to stammering, to aphasia after a stroke, to speech after laryngeal surgery, to very slow speech – and, more mildly, to a strong accent.

One channel, two groups

The phone is easier The phone is harder
Why No screen, no form, no login The speech resembles little of the training data
For whom Blind and partially sighted people, people with dyslexia, older callers without a smartphone, people without a working e-ID Dysarthria, aphasia, stammering, speech after laryngeal surgery, very slow speech
What happens They get through without asking anyone for help They are not understood, try again, give up

That is the uncomfortable part of the claim: the same channel is at once the most and the least accessible one. Which of the two it is for you does not depend on whether you have a disability, but on which one.

Want to talk it through with someone who has done this before?

Fifteen minutes on a call. We will tell you honestly whether a voice agent fits your setup – and what has to be settled before it goes live.

15 minutes · no obligation · pick your own time

What makes it worse than the error rate suggests

A form waits for you. An agent does not. Three entirely ordinary settings penalise anyone who needs longer.

The timeout. The agent reads a pause as the caller having finished, and starts talking. For someone who stammers, the pause is in the middle of the word.

The confirmation loop. The agent mishears, asks for a repeat, mishears again. On the third attempt people hang up. It is not a technical failure, and it shows up in no error statistic.

The way out requires speech. Being told to say “agent” to be put through does not work for the person who has just failed to be understood. The exit is closed for the same reason the entrance is.

Four decisions

The answer is not to stop letting an agent answer the phone. That would hurt the first group in order to spare the second.

  1. A way out that does not require the agent to understand anything. A keypress is the simplest: “press zero and I will put you through.” It works whatever the caller says or does not say, and costs nothing to add. Where that keypress goes, and what happens if nobody picks up, is the same decision as any other escalation.
  2. A longer threshold before the agent takes the floor. The defaults are set to make the conversation quick. They are not set for callers who need time to finish a sentence.
  3. Automatic transfer after two failed attempts. Not three, not five. If the agent has asked the same thing twice, the chance that a third attempt resolves it is small and the cost to the caller is high.
  4. Measure the drop-off, not the average. Recognition can be excellent on average while a small group never gets through at all. Numbers that look good on a scorecard rarely say anything about the individual call. Find the calls where the caller hung up within the first ninety seconds, and listen to ten of them.

It is not the caller's problem to solve

We have previously walked through the call where the agent heard everything and the caller did not. This is the opposite direction, and it is easier to miss, because it does not look like a fault in the log. It looks like a customer who hung up.

Voice is the most inclusive interface for the great majority. That is not the same as for everyone. The difference does not lie with the person calling, but in what you have decided happens the third time the agent fails to understand.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. The test takes five minutes: call your own number, say one sentence deliberately indistinctly, and count how many attempts you get before you are out of the conversation.