Operations and lifecycle

Models only get better. That is exactly why an automatic upgrade is a risk.

On 5 August, xAI moved its “latest” alias for voice models over to a new version, with no action required from customers. The claim that newer models are better is largely true – and that is precisely why it is worth taking apart.

IngridIngrid,
A man in a dark jumper stands on a stairwell landing in an office building after hours, having just lowered his mobile from his ear, while a colleague stands out of focus by a lit doorway behind him

The sentence turns up in almost every vendor meeting, and it is offered in good faith: models only get better, and you get the improvements automatically.

It is largely true. That is precisely why it is worth taking apart – because in production, “automatic” does not mean “free”, and “better” does not mean “the same”.

The claim, stated precisely

The model underneath your voice agent gets replaced with a newer and better version. You do not have to do anything. The result is an agent that answers better tomorrow than it does today.

Three clauses. The first two hold. The third is where the claim stops being verifiable.

Why it sounds right

Because the numbers really are going up, and they are going up fast.

When xAI introduced Grok Voice Think Fast 2.0 in late July, the company reported that Artificial Analysis' overall speech-to-speech quality index rose from 75.7 per cent for its predecessor to 82.9 per cent, that time to first audio fell from 1.25 to 0.70 seconds, and that transcription accuracy improved 1.4 times over 1.0. The company states that the gains arrive without any edits to existing prompts.

Those are real improvements on numbers that matter in a phone call. A buyer who hears this has no reason to ask to be spared them.

What was further down the same page

On 5 August, the grok-voice-latest alias moved from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0. “No action needed to upgrade”, xAI writes. To stay on 1.0, you had to pin the explicit version string before that date. Version 2.0 is priced at $0.08 per minute of audio.

Three weeks earlier, on 20 July, OpenAI notified developers using the legacy audio, realtime and transcription model families that they will be removed from the API on 20 January 2027. gpt-realtime is replaced by gpt-realtime-2.1, gpt-audio by gpt-audio-1.5, and so on down the list.

None of this is objectionable. It is ordinary product operations, announced in advance, by vendors doing exactly what they said they would do. The point is that two such events landed within three weeks, and that neither of them required a decision from the customer.

Three things the claim leaves unsaid

The unit price is part of the upgrade. A new model can cost more per minute than the one it replaces. The swap is then not only an improvement, it is also a renegotiation – concluded without a meeting. The price of a voice agent is not a number, it is a unit, and a model swapped out underneath the unit moves the invoice without touching the contract.

“Better” is an average. A model that lifts an overall index can get worse at the one thing you use it for. Shorter answers are an improvement in most calls, and a regression in the one where the customer needs three details read back before anything is confirmed. The hardest thing a voice agent hears is a number – and how well it hears one appears in no index.

The clock is not yours. The vendor sets the date. OpenAI states a minimum of six months' notice for generally available models, at least three months for specialised variants, and for models marked preview as little as two weeks. The company itself writes that preview models are not recommended for business-critical production workloads unless you can migrate on short notice. That is a warning from the vendor about the vendor's own product, and it is worth reading as one.

What is actually true

The corrected version is less catchy and more useful: models do get better, and every improvement is simultaneously a change to a system you have tested, approved and put into production.

This is not an argument for staying on old technology. An agent running on a model that leaves the API in five months has a bigger problem than one that has to be retested. It is an argument for upgrades having an owner, a date and a test – and the test already exists; it is called a baseline measurement.

Three questions cover most of it. Which model version is our agent running today, given as a string and not as “the newest one”? When is it scheduled for removal? And who here gets told on the day it moves?

If the vendor cannot answer the first, the other two are not answered either.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. That the model underneath the agent keeps getting better is the easy part of the job. Knowing which model answered when the customer called yesterday is the part that lets someone account for it.

Related Articles

Continue reading more articles