On 1 September, Meta launched a speech model called Muse Voice Transcribe. Two days later Microsoft followed with MAI-Transcribe-2. Both arrived with numbers attached: Meta reports a 3.1 percent word error rate and first place on Artificial Analysis' streaming speech-to-text ranking, while Microsoft reports a 5.2 percent average word error rate on FLEURS across 60 languages and second place on the same organisation's word error rate leaderboard.
The numbers are not invented, and the models are not uninteresting. The question is what the numbers were measured on – and whether that measurement says anything about your customer line.
What an average across 60 languages was measured on
FLEURS is an evaluation set built on the FLoRes-101 translation set: sentences taken from Wikipedia, read aloud by people, roughly twelve hours of audio per language, sampled at 16 kilohertz.
Read that description once more, because every clause in it is a distance from your phone.
It is read text, not spontaneous speech. Nobody changes their mind mid-sentence, nobody interrupts, nobody says "er, no, hang on – that was last year". It is Wikipedia prose, not "hi, I'm calling about the invoice that arrived yesterday". And it is 16 kilohertz, while an ordinary phone call carries roughly 300 to 3,400 hertz and throws away both the top and the bottom of the voice on the way in.
It is worth being precise here, because the bandwidth is not fixed. Threll.ai delivers telephony over a SIP trunk with Telenor, and the call runs in HD Voice at 16 kilohertz wherever the handset and the network at the other end support it. Where that holds, the evaluation set's sampling rate is no distance at all. The condition is simply that the support runs the whole way – if one link falls back to narrowband, it is narrowband for the rest of the call. So the point stands either way: you cannot assume which of the two a given call ran in, and that is precisely why the measurement has to be made on your own calls.
An average across 60 languages also hides every individual one of them. Meta reports that Muse was trained on more than 70 languages, and that 25 of them were extensively verified for the launch. Which 25 is a fair question to put to a vendor. For a buyer in Norway, Sweden or Denmark, the answer matters rather more than the first place does.
The research that landed three weeks before the launches
On 20 August, Lebryk and co-authors published a preprint titled "Towards Quantifying Benchmark Optimization in ASR Models", followed a day later by an article from Hume AI authors on Hugging Face. The question they ask is uncomfortably simple: does a speech model always transcribe the audio it actually receives, or can it sometimes reproduce a reference text it already knows?
They tested eleven widely used open speech models with three families of probe:
- The audio and the reference disagree. The audio is altered so that it no longer matches the reference text the evaluation set uses.
- Part of the audio is removed. Numbers and names are masked out, so the basis for writing them is not present in the signal at all.
- Several spellings are possible. The probe looks at what the model does where the orthography could have gone either way.
The finding from the second probe is the one that matters here. In the LibriSpeech experiment, where numbers had been masked out of the audio, several of the models that score best on the leaderboards recovered the masked-out number in roughly 30 to 40 percent of cases. The effect was consistently weaker on audio that had been collected recently.
Two qualifications, because they matter: the study tested neither Muse nor MAI-Transcribe-2, and it contradicts none of the launch claims. It says something about the instrument, not about the two new models.





