Conversation design

A manuscript is not a script. How to write one that survives a real call.

The manuscript a customer sends over was written by someone picturing the call going well. A script also has to cover the times it doesn't. Ten moves that separate the two.

A woman in her late thirties sits at a wooden desk in a quiet office marking up a printed manuscript page in pen, arrows and boxes drawn in the margin, a headset and a face-down phone beside her

The customer sends over a manuscript. It is written in prose, it reads well, and it was written by someone picturing the call going well.

A call script is a different object. It is a state machine with defined exits, tool contracts and a voice that has to survive being said out loud to someone who is doing something else while they talk. The difference between the two is almost always what the manuscript does not say: what the agent must never say, where the phone number it collects is supposed to end up, and what happens when the customer answers something else entirely.

Those gaps are what breaks in production. Here are ten moves that close them.

1. Don't write the script until you know what is missing

This is the rule it is most tempting to break, because a manuscript plus ordinary common sense is always enough to produce something that reads finished. A script written over unresolved gaps is worse than nothing: it looks reviewable, so it gets reviewed for tone instead of for the missing exit, and the defect turns up on a live call rather than on the page.

Sort everything you were handed into three piles: what the manuscript actually says, what is reasonable to assume but should be confirmed, and what is not there at all. Show the list before you ask anything. Then the customer answers only what is genuinely missing instead of repeating what they already sent — and the list shows them how thin a normal brief is.

Never fill a gap with something plausible. If it stays open because nobody knows the answer, it goes into the delivery as an open question.

2. Start with the endings

Before you write a single line, list every way the call can legitimately end. Then work backwards to the branches that lead there.

The order is the whole trick. Endings written last end up unreachable, because by then the flow is already laid out and nothing routes to the case you just thought of. An outbound call almost always has more endings than the brief implies:

  • goal reached
  • customer declines
  • wrong person
  • no conversation possible
  • customer wants another channel
  • callback agreed
  • transferred to a human

Each needs one line and a hangup, and each is owned by exactly one place in the script. Two definitions of the same ending means two different goodbyes for the same situation.

3. One mandate, and only one

The mandate should fit in a single sentence, with a defined success state and a clear boundary for what is out of scope.

Two mandates in one script gets you two jobs done halfway. If the agent is meant to both follow up an offer and map demand for something new, that is two calls — and usually two agents. If you can't read the script start to finish in one sitting, the mandate is too broad, not the script too long.

4. Write states, not paragraphs

The most common weakness in an otherwise well-written script is that it is one long paragraph. Give every step in the flow five fixed fields: id, purpose, instructions, example lines, and exits with conditions.

map_need

Purpose: find out whether the customer has an active policy with a competitor. Instructions: Ask one question, wait for the answer, confirm briefly, move on. Do not comment on the provider the customer names. Example lines: "Do you have insurance with anyone else at the moment?" Exits: Has a competitor → next step. Has nothing → skip the offer. Won't answer → exit "declines".

Two things you get for it. You can change one state without touching the rest, and a missing exit becomes visible on the page instead of a place where the call quietly stalls.

5. Anything in quotation marks is said out loud

Keep one convention throughout: what sits inside quotation marks is spoken verbatim. Everything outside is instruction to the model. Never mix the two in one sentence.

That means no markdown, no bullets, no emoji and no parentheses inside the lines. A stage direction in brackets inside the quotes gets read aloud. Numbers, dates, amounts, phone numbers and opening hours are written the way they are pronounced — "twelve ninety-nine", "Friday the fourth of September", "nine to four on weekdays". Numbers are the hardest thing a voice agent handles anyway, so they deserve to be spelled out.

And sweep for placeholders before you deliver. A [company name] or {{customer_name}} that never got filled in is read out exactly as it stands.

6. One sentence at a time, one question per turn

Customers interrupt. That isn't rudeness, it is how phone calls work — and when they do it, whatever the agent hadn't said yet disappears. What never reached the caller is not in the conversation afterwards, not for the model either.

So: one or two sentences per turn, and one question. A double question gets one answer, and you won't know which half it belongs to. Write an explicit interruption rule as well: stop, listen, answer what came, don't resume the sentence.

The opening line is where this costs most, because it is the one sentence every caller hears.

7. Give the example lines a variation rule

Example lines without a rule get read as a manuscript. Anyone who hears two calls notices immediately, and it is the fastest route to an agent that sounds like an answering machine.

Write the rule in: the example lines are inspiration, not a script — vary the wording, and never repeat the same sentence twice in one call.

Does your script have an exit for every way the call can go?

We're happy to read through an existing script and show you where the branches stall — with nothing rewritten first.

15 minutes · no obligation · pick your own time

8. Three fields per tool — and none left undocumented

Anything that fetches or writes data sits behind a tool. Every tool needs three things, and missing one of them means the script isn't finished:

  • An invocation condition. Written as its own sentence, not chained on with "and then". "Call this only once the customer has named a specific time. Don't call it on 'maybe Thursday'."
  • A waiting line. The call takes time, and the gap is silent. Without "Let me check that, one moment" the customer starts talking into the pause.
  • A failure line. Without it, the agent invents a confirmation that never happened while the backend was down.

The list has to cover every tool the agent can call, not just the ones the flow plans to use. The classic case is a knowledge-base search switched on during setup and never mentioned in the script: the model reaches for it the moment someone asks something unusual, and the silence arrives with no waiting line. A tool that should never be used gets one line saying exactly that.

The same logic applies to promises. "I'll send you a text" requires an SMS channel. "An adviser will call you tomorrow" requires someone who actually gets the message. A promise without a tool behind it is a defect, not a courtesy.

Close-up of a printed script page where one long line has been split in two in pen and a single word further down is underlined
One long line split in two. The full stop is the pause that behaves the same way every time.

9. Pauses, pace and stress are written in — they can't be marked up

Delivery is steered with plain language, not with markup. Three to five stable directions covering the whole call beat twenty detailed ones per line — and detailed emotional direction is the most common reason an agent sounds like an actor. All direction goes outside the quotation marks; inside, it gets read aloud.

There is no tag syntax for prosody. "[pause]", asterisks around a word or angle brackets are either ignored or spoken. What you have is the words, the punctuation and a few directions in plain language — and used deliberately, that goes further than people expect:

  • Pauses come from punctuation and sentence length. A full stop gives a clearer pause than a comma, and a short sentence on its own gives more air than a long one with a subclause. If you need a marked break — after the company name, before an amount, after a difficult question — write it as two sentences instead of one. Dashes and ellipses don't behave the same way twice. A full stop does.
  • Pace is set in plain language, and repeated where it matters. "Speak calmly, a little slower than normal conversational pace" belongs in the voice section and covers the whole call. If one step needs to be slower — reading back an account number, a confirmation the customer is writing down — put the direction on its own line in that step, outside the quotation marks. Pace is also contagious: the agent tends to match the customer, so a stressed caller pulls the tempo up unless the script says otherwise.
  • Stress comes from word order, not from formatting. The word that carries the weight goes last in the sentence, or stands alone in a short one. "There is no lock-in." lands harder than "We don't have any lock-in on this product, just so you know." If you want the customer to hear the day, "Let's say Thursday." beats anything you could think to mark up.
  • Pronunciation is written phonetically, as a rule. There is no pronunciation markup to lean on, so put it in the spoken-style section: "Threll is pronounced trell." Same move for abbreviations, foreign product names and surnames the agent will say often. Names are what an agent mishears most, and they are also what it most often says wrong.
  • If a sentence has to sound identical every time, don't let the model make it. A required disclosure, or a fixed welcome with a particular stress, is safer as a pre-rendered audio file from a speech synthesiser than as something generated afresh on every call with slightly different rhythm.

One last thing that surprises people: the directions drift. After a few interruptions the agent is back at its own pace, whatever was written at the top of the script. If tone matters at the close, repeat the direction in the final step. And the same wording lands differently on different voices — test per voice, not just per script.

10. Robustness is its own section — and the absolute rules come last

Give everything that isn't the main path a section of its own: silence, an unclear answer, voicemail, an angry customer, and an outer limit on how long the call can run. A real-time session has a ceiling, and a script that assumes twenty minutes is a design error, not a long script. Build in a summarise-and-close move before the ceiling is reached.

This is also where escalation to a human belongs, along with what happens if the customer switches to a language the agent didn't open in.

The absolute rules — what the agent must never do — belong at the very end. Moved to the top they read as background, and lose against a helpful-sounding objection line further down. Write them as prohibitions, not preferences: "avoid where possible" is not a rule. Use capitals on three to five things, maximum. Used everywhere, they mean nothing.

The order that makes a script readable

Put together, a script you can live with runs in this order:

role → voice → spoken style → context → mandate → opening → flow → modules → tools → objections → out of scope → robustness → terminal states → absolute rules.

Persona and framing first, the flow in the middle, the prohibitions last. That isn't a matter of taste — the order is part of how the instruction is read.

Before you send it out

Read the script out loud. Half the defects you hear before you see them: sentences too long to be spoken, numbers written in digits, a parenthesis that should have been an instruction.

Then work through four questions. Does every branch lead to an exit that exists? Does every promise have a tool behind it? Does every tool have all three fields? Is there an objection line that quietly cancels one of the absolute rules?

And don't deliver a draft with known defects described as "something we'll fix later". A defect in a delivery note is a defect the customer approves by accident.

Finally: put a version number on the script and keep it in the document. When someone asks in three weeks why the agent said what it said on a particular call, the version number is the difference between an answer and a guess. Remember too that the analysis schema for call outcomes is a separate job. Script and schema drift apart without anyone noticing, and the drift is only visible once both exist.

Threll.ai builds voice agents in Norwegian, Swedish and Danish. The script is the part of the setup that is cheapest to change and most expensive to get wrong — and almost always the part that gets the least time.

Frequently asked questions

The test isn't word count, it's whether you can read it start to finish in one sitting and still remember what was at the top. If you can't, the mandate is usually too broad rather than the script too long. The answer then is to split the job across two agents with one mandate each, not to cut robustness and exits to make room.

Design as if you can't. The instruction is set when the call connects, and a design that assumes it will be swapped out when the customer says yes is fragile in a way that only shows up in production. Modularity belongs in the authoring phase: build reusable blocks for booking, handover and identification, and compile them into one script before it goes live.

With the same set of calls every time, so two versions can be compared. Run the main path, the wrong person, the caller who says nothing, the caller who interrupts mid-sentence, the most common objection, a question outside the mandate, and voicemail. Read the transcripts afterwards rather than only listening to recordings — that is where you see over-long turns, repeated phrasing and language drift.

The exits and the tools. Count how many ways the call can end and check that every branch actually points at one of them. Then go through everything the agent is able to call and look for tools that are switched on but never mentioned in the script. Those two things explain most of the calls that stall for no visible reason.

You can, but it is audible. A sentence translated word for word keeps the rhythm of the language it came from, and in an outbound call that is precisely the first impression you don't want. Rewrite the lines in the target language with the same flow and the same exits, and have a native speaker read them through before they go live.