Rollout and piloting

How to run a two-week pilot that actually settles the question

Most voice agent pilots are set up so that they cannot fail. Which means they cannot settle anything either. Here are the eight steps that turn two weeks into a decision.

IngridIngrid,
Two colleagues sit on either side of a desk in a small warehouse office after hours, lit by desk lamps; the man on the left has just put down the desk phone standing between them, while the woman on the right holds a printed sheet and asks him something

A pilot ends remarkably often with the sentence “it went fine”. Nobody quite knows what went fine, and nobody can say what would have had to happen for it not to.

That is rarely because the technology is unclear. It is because the pilot was set up as a demonstration rather than a trial. A demonstration exists to show that something can work, and it almost always manages that. A trial exists to find out whether it does – and it therefore has to be able to fail.

Two weeks is enough. The most common mistake is spending three months on something that was settled after nine working days. What follows is the order that turns those two weeks into a decision rather than a recommendation.

1. Write down the decision before you write the conversation

The first sentence of the pilot document should not be about the agent. It should be about what you are going to do on the fifteenth of next month.

“We will decide whether the agent takes every incoming appointment booking outside opening hours from 1 October” is a decision. “We will test voice AI” is not.

Then write down the harder half: which result means no. One number, agreed before you have seen any data. If nobody can articulate what would make you call it off, you are not going to call it off – whatever the two weeks show.

2. Pick one call type, not one department

“Customer service” is not a call type. It is twenty.

Pick one with high volume, a short life and a clear end point: confirming an appointment, giving the status of a case, taking a cancellation, logging a fault. Something you can count before and after.

The objection arrives immediately: surely the easy calls are not worth automating. It is wrong in two ways. The easy ones are most of them, and they are the only ones that give a clean answer in two weeks. A pilot placed on the hard calls does not measure the agent. It measures the escalation.

3. Measure the baseline for five days before the agent goes live

Almost nobody does this, and it is why almost no pilot proves anything.

Five ordinary working days, agent switched off. Count how many people called about the chosen call type, how many got through, how long they waited, how many hung up before anyone picked up – and how many called back about the same thing within a day.

That last number is the important one, and the only one that is slightly awkward to extract. Do that work now. In two weeks it is the only line you will be arguing about.

4. Agree the way out before you agree the way in

What decides whether a pilot feels successful is not the opening line. It is what happens when the agent cannot get any further.

Settle three things before the first call: what triggers a transfer, what the human is told when the call arrives, and what happens when nobody picks up. Escalation is the part that is most expensive to do halfway, and in a pilot it is also the part everybody notices.

Set the threshold low for the first few days. An agent that transfers too early hands you a list of what it could not do. An agent that hangs on too long hands you irritated customers and no list.

5. Week 1: limit the volume, not the ambition

Do not let the agent do ten per cent of the job on all of the calls. Let it do the whole job on ten per cent of the calls.

An agent that resolves one case type completely, on a small volume, is measurable. An agent doing a bit of everything across the whole volume is not – you end up with an impression rather than a result.

The agent must say that it is AI in its opening line. That is an obligation, not a tactic, and it applies in a pilot too, however small the volume. It is also bad measurement practice to leave it out: a pilot that conceals what the agent is measures something you will not be allowed to run afterwards.

6. Week 2: take the help away

In the first week somebody will be watching closely. Someone adjusts a phrase on Tuesday, fixes an error on Wednesday, sits alongside on Thursday.

That is right in week 1 and ruinous in week 2. A system that changes every day cannot be measured – you end up measuring your own attention rather than the agent.

Freeze the configuration on the Monday morning of week 2. Write down everything you want to change and leave it until Friday. The list will be shorter than you expect.

7. Listen to twenty calls you did not choose

Draw twenty at random. Not the twenty shortest, not the twenty the vendor suggests, not the twenty that were escalated.

The numbers tell you how often something happened. The twenty calls tell you what happened – and they are the only source for the errors that have no counter behind them: the agent confirming the wrong date in exactly the right format, the customer who gave up without hanging up, the word nobody had thought of that people actually use for your product.

Set aside two hours. Have the person who will run this day to day sit in.

8. Decide on the number from step 1

On the Friday of week 2, put the baseline from step 3 next to the two weeks and read the one line you agreed on in advance.

Resist the urge to swap numbers. The urge will come, and it will come nicely wrapped: average handling time is down, the answer rate is higher, the agent handled 340 calls without complaining. Those are numbers that flatter, and they are usually true. They are just not what you asked.

What two weeks cannot settle

Ten working days say nothing about running this over time. Nothing about what happens in the week between Christmas and New Year, about how the agent behaves when an integration is down, or about what becomes of the setup when the person who built it leaves.

The pilot can settle one thing, and that one thing is worth two weeks: whether this call type gets better or worse for the customer when the agent takes it.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. A pilot that could not have ended in no has answered nothing. It has spent two weeks confirming that somebody had already made up their mind.

Related Articles

Continue reading more articles