Measurement and results

The numbers that flatter your voice agent

A voice agent that handles 68 percent of calls without transferring looks like a success. The number says nothing about how many of those calls were actually resolved — and even less about what the business did with the capacity it freed up.

IngridIngrid,
Kvinne med headset rundt halsen står ved et hev-senk-bord og snakker med en kollega over et ark papir, mens en rad tomme arbeidsplasser med stille bordtelefoner ligger i uskarphet bak dem

The report from a voice agent almost always opens with the same number: the share of calls handled without a human joining. Containment, deflection, automation rate — the names vary, the number does not.

It is a useful number for sizing a team. It is a poor number to steer by, for one reason: it counts calls that were not transferred, not cases that were resolved. The customer who gave up after two minutes, the customer who hung up to try the chat instead, and the customer who actually got an answer all look identical in the statistics.

Here are the four numbers most often used to show off a voice agent, what they can hide, and what you can measure instead.

Swap one number at a time

The number that flatters What it can hide The number that answers more honestly
Share of calls handled without a human Customers who gave up mid-call, or who called back two hours later Share of enquiries resolved with no repeat contact on the same case within 72 hours
Shorter average handling time The agent closing quickly instead of clarifying Share of calls where the customer confirmed the matter was settled before the call ended
Answer rate and zero queue That everyone is answered, but not necessarily helped Share of calls that ended in an actual action in a business system: order changed, case created, appointment moved
Satisfaction measured right after the call That people answer politely in the moment Satisfaction measured after the case was actually closed

None of the numbers in the right-hand column require new technology. They require the call to be tied to a case, and the case to have a state. That is an integration job, not a choice of model.

Repeat contact is the most honest number you have

If you only introduce one new measure, make it this one: how many people get back in touch about the same case.

Repeat contact is hard to dress up. It needs no survey and no interpretation of tone of voice — only that calls are linked to a customer and a case type. And it captures exactly what containment does not: the difference between a call that ended and a case that ended.

It has one more advantage. Broken down by case type, it tells you where the agent belongs. That nine in ten opening-hours questions are resolved on the first call, while three in ten complaints are, is usable information. A single blended average of 68 percent is not. This is also the difference between the demo and live operation: a demo has no case types, just one call that goes well.

IKEA measured the other question

On 3 August, CX Today covered IKEA's experience with its customer service bot Billie, drawing on a Fortune piece from 30 July. The numbers are worth reading slowly.

Billie was introduced in 2021 for repetitive enquiries: opening hours, order status, stock levels, return policy. In its first two years it could help 47 percent of the people who used it. That share is now put at 74 percent.

The interesting part is not the 74. It is what IKEA did with the roughly 8,500 customer service employees whose work Billie partly took over. They were not made redundant but retrained — to handle complex cases and to work as remote interior design advisors. According to Fortune, the reskilling took about two years.

The remote sales centres generated €1.25 billion in the last financial year, up from €1.08 billion, and have grown 15–20 percent a year over the past three years. The company also reports an internal happiness score up from 60 to 89.

Two years of reskilling 8,500 people is not a project most businesses have. But the measurement logic can be copied without IKEA's budget: the share of calls the bot took was the intermediate step. The answer lay in what the freed-up time was spent on.

Capacity is the real deliverable

Most organisations can say how many calls the agent took. Far fewer can say what their staff did in the hours that came free.

This is not an HR exercise. It is the only way to tell apart two outcomes that look identical in the first year's accounts: customer service got cheaper, and customer service got better. Only one of them holds when volume rises again — and volume is what reveals whether AI is an experiment or part of daily operations.

Freed capacity disappears if nobody decides what it is for. That is not a technology problem, it is a management problem, and work must be directed even when it is no longer performed by humans.

Four questions you should be able to answer in three months

How many cases were resolved — not how many calls were taken.

Which case types the agent is genuinely good at, broken down by type.

How many customers came back about the same thing within three days.

What the freed-up hours went to.

None of these require a new platform. They require you to decide what you measure before the vendor decides it for you, and the vendor to be able to deliver the numbers in a form you can audit. That belongs in the vendor choice, not in the first monthly report.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. The first number is easy to deliver. The second decides whether the agent is still there in a year.

Related Articles

Continue reading more articles