The customer calls to move an appointment. She is standing at the kitchen counter with the phone on speaker. The agent offers Thursday at 2 pm. She says "yes, that …", and at the same moment her husband at the table says "no, Thursday we're going to my mother's".
The agent has received two answers to one question, from two people, in the same audio stream. It doesn't know there are two.
On Monday, 28 September, Fish Audio launched transcribe-1-pro, a transcription model that labels who is speaking in a recording. It inserts numbered speaker tags and markers such as "[laughter]" straight into the text, at $0.36 per audio hour. That is useful. It is also a good way into explaining why the problem is so much harder while the call is happening than afterwards.
One channel, every voice
An ordinary phone call is a single audio channel. The phone captures what its microphone hears and sends it on as one stream. There is no separate track for the customer and another for the rest of the room. The line also throws away a good part of the sound on the way, so what reaches the agent is both narrower and more mixed than what the customer hears herself.
Speech recognition does what it was built to do: it writes down speech. It does not distinguish between speech meant for the agent and speech that just happens nearby. Four sources come up again and again:
| What is in the room | What the agent gets wrong | What helps |
|---|---|---|
| Another person talking | Takes a remark to the customer as an answer | Reading important details back and waiting for a yes |
| Radio, TV or music in the background | Transcribes words nobody said to it, or thinks the customer is interrupting | Requiring more than a couple of syllables before the agent stops |
| The agent's own voice coming back | Hears itself and interrupts itself | Echo cancellation, tested on speakerphone and in a car |
| An open-plan office | Mixes in colleagues' conversations | Coping when the customer says "hang on" |
The agent that interrupts itself
The third row surprises people most. When the customer has the phone on speaker or talks through the car's hands-free system, the agent's voice comes out of a loudspeaker a few centimetres from the microphone. The phone tries to subtract that sound again – this is called echo cancellation – and it is good, but not flawless. Whatever slips through goes back to the agent.
An agent configured to stop talking the moment it hears speech then does exactly that: it stops mid-sentence because it heard itself. To the customer it sounds like an agent that stutters, or one that never lets her hear a complete answer.
The same setting decides what happens with the radio. Set the threshold too low and the agent stops for every voice in the background. Set it too high and it talks over the customer when she really does want to interrupt. There is no single right value. It has to be tried on the calls you actually get.





