Model selection and testing

Two new speech models in three days. Neither score was measured on your phone line.

Meta and Microsoft each released a speech recognition model in the first week of September, with leaderboard places and error rates to show for it. The numbers are real enough. They were just measured on something quite unlike a Nordic phone line.

5 min read
Two colleagues sit on either side of a pale wooden desk in a back office at a clinic; the woman on the left holds the handset of a black desk phone to her ear, listening, while the man on the right leans over a printed form with a pen in his hand and a mobile lies face down on the desk between them

On 1 September, Meta launched a speech model called Muse Voice Transcribe. Two days later Microsoft followed with MAI-Transcribe-2. Both arrived with numbers attached: Meta reports a 3.1 percent word error rate and first place on Artificial Analysis' streaming speech-to-text ranking, while Microsoft reports a 5.2 percent average word error rate on FLEURS across 60 languages and second place on the same organisation's word error rate leaderboard.

The numbers are not invented, and the models are not uninteresting. The question is what the numbers were measured on – and whether that measurement says anything about your customer line.

What an average across 60 languages was measured on

FLEURS is an evaluation set built on the FLoRes-101 translation set: sentences taken from Wikipedia, read aloud by people, roughly twelve hours of audio per language, sampled at 16 kilohertz.

Read that description once more, because every clause in it is a distance from your phone.

It is read text, not spontaneous speech. Nobody changes their mind mid-sentence, nobody interrupts, nobody says "er, no, hang on – that was last year". It is Wikipedia prose, not "hi, I'm calling about the invoice that arrived yesterday". And it is 16 kilohertz, while an ordinary phone call carries roughly 300 to 3,400 hertz and throws away both the top and the bottom of the voice on the way in.

It is worth being precise here, because the bandwidth is not fixed. Threll.ai delivers telephony over a SIP trunk with Telenor, and the call runs in HD Voice at 16 kilohertz wherever the handset and the network at the other end support it. Where that holds, the evaluation set's sampling rate is no distance at all. The condition is simply that the support runs the whole way – if one link falls back to narrowband, it is narrowband for the rest of the call. So the point stands either way: you cannot assume which of the two a given call ran in, and that is precisely why the measurement has to be made on your own calls.

An average across 60 languages also hides every individual one of them. Meta reports that Muse was trained on more than 70 languages, and that 25 of them were extensively verified for the launch. Which 25 is a fair question to put to a vendor. For a buyer in Norway, Sweden or Denmark, the answer matters rather more than the first place does.

The research that landed three weeks before the launches

On 20 August, Lebryk and co-authors published a preprint titled "Towards Quantifying Benchmark Optimization in ASR Models", followed a day later by an article from Hume AI authors on Hugging Face. The question they ask is uncomfortably simple: does a speech model always transcribe the audio it actually receives, or can it sometimes reproduce a reference text it already knows?

They tested eleven widely used open speech models with three families of probe:

  1. The audio and the reference disagree. The audio is altered so that it no longer matches the reference text the evaluation set uses.
  2. Part of the audio is removed. Numbers and names are masked out, so the basis for writing them is not present in the signal at all.
  3. Several spellings are possible. The probe looks at what the model does where the orthography could have gone either way.

The finding from the second probe is the one that matters here. In the LibriSpeech experiment, where numbers had been masked out of the audio, several of the models that score best on the leaderboards recovered the masked-out number in roughly 30 to 40 percent of cases. The effect was consistently weaker on audio that had been collected recently.

Two qualifications, because they matter: the study tested neither Muse nor MAI-Transcribe-2, and it contradicts none of the launch claims. It says something about the instrument, not about the two new models.

Do you measure model switches on your own calls?

Fifteen minutes in which we set up the simplest version of the test: thirty calls, your own ground truth, and the fields that actually carry the case.

15 minutes · no obligation · pick your own time

A model that fills in a number is worse than one that says it did not hear

This is where it stops being academic.

On a customer line, the numbers are often the entire payload. The order number, the account number, the date, the amount, the house number. Digits are the weakest point in Nordic speech recognition to begin with – narrowband audio, competing ways of saying a figure and pronunciation that shifts from one valley to the next all pull in the same direction.

A model that writes "unclear", or that flags its own uncertainty, gives your setup a chance: the agent can ask again, read the number back, or transfer the call. A model that writes a plausible number gives you no such chance. It delivers an answer that looks exactly as correct as all the other answers, and the error surfaces in the accounts or on the delivery date.

Word error rate does not distinguish between the two. A wrong digit and a superfluous "you know" count the same. That is why the number is a poor purchasing criterion even when it has been measured correctly.

The test you can run without a laboratory

This needs no new vendor and no new feature.

  1. Take thirty of your own calls. From the number that is actually in use, through the call path that is actually in use. Not recordings made on a good microphone in a meeting room.
  2. Write the ground truth yourself, afterwards. Someone who knows your customers listens through and writes down what was said. The whole point is that no model can have seen this ground truth before.
  3. Count only what carries the case. Numbers, names, addresses, times – and negations. A "not" that disappears reverses the meaning of an entire sentence. Everything else is noise in this measurement.
  4. Separate errors from fabrication. Note specifically where the model wrote something concrete in a place where the audio was unclear. That is the error type that does not announce itself until it is expensive.

Thirty calls is enough to see a difference that matters. It is also small enough that nobody has to apply for a budget to do it.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. New models are mostly better than the old ones, and switching to them is worth doing. We just tend to suggest that the switch be measured on your own calls – because a switch that happens by itself, with nobody measuring anything, is not an upgrade.

Frequently asked questions

That cannot be answered from the outside, and that is the whole point. Use the leaderboards to shortlist candidates, and let your own calls decide. While you are at it, check which languages the vendor has actually verified, which region the audio is processed in, and what the price becomes after any launch offer expires.

No, but it is a screening tool. It treats every error as equally large, and on a customer line they are not: a wrong digit in an account number and a superfluous filler word count the same. Measure word error rate if you like, but do not decide anything on it alone.

No. The study tested eleven open models, and neither Muse nor MAI-Transcribe-2 was among them. It says that a high score on a well-known evaluation set may come partly from recognising the set itself – that the instrument is imprecise, not that anyone has done anything wrong.

Thirty is enough to see differences that matter, provided they come from the number and the call path you actually use. Thirty real calls tell you more than three hundred studio recordings. And the ground truth has to be written by you afterwards, so that no model can have seen it before.