On 10 September, OpenAI released GPT-Live-1 in the API at $0.05 a minute. The price explicitly covers the voice layer only – the part that listens and speaks. Deeper reasoning and tool calls are delegated to a text model the developer picks, and pays for separately.
That is an accurate description of something that has been true for a while: a voice agent is not one thing. It is a chain of links that each have to hold in turn, within a few seconds, in a fixed order. Here is one perfectly ordinary call – a shop owner ringing his supplier about a delivery that arrived wrong – and what has to hold at each link.
The phone rings
The first link has nothing to do with AI. The call has to travel from the caller's operator, onto your number, through your telephony provider and to the agent, ideally inside three or four rings.
It is also the only link with nothing to fall back on. If it fails, the caller hears nothing at all. Who owns the number, and where the call goes when the agent does not answer, is decided here – not in an argument about which model to use.
The first sentence
The agent says something. It is the one sentence a hundred per cent of callers hear, and it has to do four things at once: say who is answering, say that this is an AI, offer a way out, and ask one question open enough that the caller starts talking.
The failure is nearly always the same. The sentence was written last, in five minutes, by someone who wanted to be finished.
"Yeah, hi – it's about a delivery that came yesterday"
The caller tells you everything at once, in the wrong order, with a detour about how he rang last week as well. Nobody speaks in form fields.
What has to hold here is that the agent does not try to steer him back to step one. It should pick up what it got for free – that this is about a delivery, that it arrived yesterday, that something is wrong – and ask only for what is missing. An agent that asks about something it has already been told is the fastest way there is to lose a caller.
The order number
Now a string of digits has to be read out, and this is the weakest point in the whole chain – more so in the Nordic languages, where the numbers can be said two ways and both are correct, while the phone line throws away much of what separates one digit from another.
What has to be true here is not that the agent hears correctly every time. It is that it reads the number back when it is unsure, instead of writing down something that merely looks plausible.






