From keyboard to conversation
Most people associate artificial intelligence with text: ChatGPT, email assistants, text writers, and analytical language models. But it's easy to forget one fundamental truth:
Humans are not made for the keyboard. We are made for the voice.
We speak before we write. We use intonation, pauses, laughter, and hesitation to convey meaning. We interrupt, think aloud, and don't always finish our sentences. All of this is natural to us – but extremely difficult for machines.
Therefore, Voice AI is the next big step in artificial intelligence. It's not just about getting machines to read aloud text. It's about getting them to understand, respond, and act in natural conversations – in a truly human way.
And that is precisely what Threll Voice Lab is designed to master.
What is Voice AI – really?
Voice AI is technology that allows people to speak directly with artificial intelligence – without a keyboard, screen, or menus. When you talk to a voice agent from Threll, a lightning-fast chain of processes occurs within a few hundredths of a second.
This process is called the speech process, and it consists of five steps:
1. ASR – Automatic Speech Recognition
Speech in → Text out
ASR is the ears of the system. It is here that your voice is translated into text that the AI can understand. This may sound simple, but it is one of the most challenging aspects in Norwegian Voice AI. ASR must handle:
- Norwegian dialects, Nynorsk, and regional expressions.
- Telephone lines that cut off large parts of the sound spectrum (only 300–3400 Hz).
- Bluetooth delays and unstable bandwidth.
- Background noise, echo, and overlapping speech.
At Threll Voice Lab, we train ASR on real Norwegian voices – from all over the country – so that the system learns how people actually speak. Not how they write.
2. NLU / LLM – Natural Language Understanding
Text → Understanding
When words are transcribed, the AI must understand what you mean. NLU is about semantics – that is, interpreting the meaning behind what you say. Our language is full of short, incomplete, and ambiguous expressions:
"Yes, you can do that." "Maybe. We'll see." "No, wait a minute – I meant Tuesday, not Thursday!"
NLU (often driven by large language models, so-called LLMs) needs to figure out:
- What is the intention (order, cancel, ask, confirm)?
- What key information is mentioned (name, date, product, location)?
- What does the tone mean – is the user satisfied, irritated, unsure?
3. Dialog Manager
Controls the flow of conversation
Dialog Manager (DM) is the brain that keeps track of the context in the conversation. It determines:
- When the AI should speak, and when it should listen.
- What has been said earlier, and what remains.
- How to handle interruptions, overlapping speech, and questions outside the topic.
Without a good Dialog Manager, conversations become chaotic. A professional DM ensures that the conversation flows naturally, even when you're talking over the AI – just like humans do.
4. Action Layer
Understanding → Action
Here lies the core of Threll's approach: AI should not just talk – it should do something.
The Action layer connects to the customer's own systems, such as:
- CRM (HubSpot, Salesforce, SuperOffice)
- ERP (Visma, SAP, Tripletex)
- Order and booking systems
- Customer service tools
When you say:
"Can you send me a copy of the invoice from June?" … then Voice AI does exactly that. It finds the invoice, sends it, and tells you when it's done.
This is the difference between a talking robot and a voice agent that actually creates value in the company's value chain.





