The claim turns up early in almost every project, and it rarely gets challenged: an app asks you to find it, log in, hit the right button and read the small print. The phone only asks you to speak. Everyone can speak.
The first half of that is true, and it matters more than critics of voice agents tend to concede.
Who the phone genuinely opens the door for
For someone who is blind or partially sighted, a phone line is an interface with no screen to navigate. For someone with dyslexia it is an interface with no form. For someone without a smartphone, or whose national e-ID has never quite worked, it is the only way into a business that has moved everything else online. And for someone speaking a second language, it is often easier to explain a problem in their own words than to find the right category in a dropdown.
This is not a small group. In practice the phone is the fallback for every case self-service does not fit – and an agent that answers on the first ring, around the clock, makes that fallback markedly better than a queue with hold music.
So far the claim holds.
Where it stops holding
“Everyone can speak” is a claim about the person. It says nothing about the machine.
A speech recognition model hears well whatever resembles what it was trained on, and worse everything further away. We have written before about the mild cases of that: the phone line throws away a good deal of the signal before the model ever gets it. Further along the scale it looks entirely different.
A research team at Stony Brook University tested Whisper against the TORGO dataset last autumn – recordings of eight people with dysarthria, a motor speech disorder caused by damage to the nervous system, after a stroke, in cerebral palsy or in ALS. With no adaptation at all, Whisper-small got 71 per cent of the words wrong on average.
The average is not the interesting part. The spread is. Of the eight, the one least affected had an error rate of 7 per cent – roughly what anyone else would get. The one most affected was above 100 per cent, meaning more errors than there were words to get wrong. Same model, same task, same diagnosis, and two completely different services.
The study is small, and the authors say so themselves: eight people, English speech, one model family. But the mechanism is not English. It applies equally to stammering, to aphasia after a stroke, to speech after laryngeal surgery, to very slow speech – and, more mildly, to a strong accent.
One channel, two groups
| The phone is easier | The phone is harder | |
|---|---|---|
| Why | No screen, no form, no login | The speech resembles little of the training data |
| For whom | Blind and partially sighted people, people with dyslexia, older callers without a smartphone, people without a working e-ID | Dysarthria, aphasia, stammering, speech after laryngeal surgery, very slow speech |
| What happens | They get through without asking anyone for help | They are not understood, try again, give up |
That is the uncomfortable part of the claim: the same channel is at once the most and the least accessible one. Which of the two it is for you does not depend on whether you have a disability, but on which one.





