Voice Lab

From text to speech – why the future of AI is about the voice

Threll Voice Lab builds voice agents that understand Norwegian, all dialects – and that don't just speak, but act.

IngridIngrid,
Class room

From keyboard to conversation

Most people associate artificial intelligence with text: ChatGPT, email assistants, text writers, and analytical language models. But it's easy to forget one fundamental truth:

Humans are not made for the keyboard. We are made for the voice.

We speak before we write. We use intonation, pauses, laughter, and hesitation to convey meaning. We interrupt, think aloud, and don't always finish our sentences. All of this is natural to us – but extremely difficult for machines.

Therefore, Voice AI is the next big step in artificial intelligence. It's not just about getting machines to read aloud text. It's about getting them to understand, respond, and act in natural conversations – in a truly human way.

And that is precisely what Threll Voice Lab is designed to master.

What is Voice AI – really?

Voice AI is technology that allows people to speak directly with artificial intelligence – without a keyboard, screen, or menus. When you talk to a voice agent from Threll, a lightning-fast chain of processes occurs within a few hundredths of a second.

This process is called the speech process, and it consists of five steps:

1. ASR – Automatic Speech Recognition

Speech in → Text out

ASR is the ears of the system. It is here that your voice is translated into text that the AI can understand. This may sound simple, but it is one of the most challenging aspects in Norwegian Voice AI. ASR must handle:

  • Norwegian dialects, Nynorsk, and regional expressions.
  • Telephone lines that cut off large parts of the sound spectrum (only 300–3400 Hz).
  • Bluetooth delays and unstable bandwidth.
  • Background noise, echo, and overlapping speech.

At Threll Voice Lab, we train ASR on real Norwegian voices – from all over the country – so that the system learns how people actually speak. Not how they write.

2. NLU / LLM – Natural Language Understanding

Text → Understanding

When words are transcribed, the AI must understand what you mean. NLU is about semantics – that is, interpreting the meaning behind what you say. Our language is full of short, incomplete, and ambiguous expressions:

"Yes, you can do that." "Maybe. We'll see." "No, wait a minute – I meant Tuesday, not Thursday!"

NLU (often driven by large language models, so-called LLMs) needs to figure out:

  • What is the intention (order, cancel, ask, confirm)?
  • What key information is mentioned (name, date, product, location)?
  • What does the tone mean – is the user satisfied, irritated, unsure?

3. Dialog Manager

Controls the flow of conversation

Dialog Manager (DM) is the brain that keeps track of the context in the conversation. It determines:

  • When the AI should speak, and when it should listen.
  • What has been said earlier, and what remains.
  • How to handle interruptions, overlapping speech, and questions outside the topic.

Without a good Dialog Manager, conversations become chaotic. A professional DM ensures that the conversation flows naturally, even when you're talking over the AI – just like humans do.

4. Action Layer

Understanding → Action

Here lies the core of Threll's approach: AI should not just talk – it should do something.

The Action layer connects to the customer's own systems, such as:

  • CRM (HubSpot, Salesforce, SuperOffice)
  • ERP (Visma, SAP, Tripletex)
  • Order and booking systems
  • Customer service tools

When you say:

"Can you send me a copy of the invoice from June?" … then Voice AI does exactly that. It finds the invoice, sends it, and tells you when it's done.

This is the difference between a talking robot and a voice agent that actually creates value in the company's value chain.

5. TTS – Text-to-Speech

Action → Voice response

Once the action is performed, the AI must respond to you naturally – with the right tone, pace, and feeling. Here comes Text-to-Speech (TTS).

TTS converts the text into human speech, but not just as a reading aloud. It should:

  • Adapt to the context (calm voice when complaining, energetic when selling).
  • Vary tone and rhythm.
  • Maintain low latency – respond within half a second to feel natural.

The final touch is about emotional realism: small pauses, laughter, surprise – everything that makes the AI voice feel lively.

When everything works together

When all five layers – ASR → NLU → Dialog → Action → TTS – work seamlessly together, magic happens:

You speak. The AI understands. It acts. It responds. All in under a second.

This is the difference between a digital assistant – and an autonomous voice agent.

Why this matters for businesses

For companies, Voice AI is not just about technology. It’s about efficiency, customer satisfaction, and scaling.

  • Customer Service: Voice can handle inquiries before they reach the queue.
  • Sales: AI can manage qualification calls and meeting bookings 24/7.
  • Support: Voice can provide guidance and updates without waiting.
  • Operations: Integrated Voice AI cuts manual steps and frees up staff.

When the voice agent is connected to the company's systems, it can make real decisions – not just answer questions

Why this is exciting for investors

Voice AI has three characteristics that make it particularly interesting as an investment area:

  1. **High technical barrier.**The system requires real-time interaction between multiple advanced models (ASR, NLU, TTS) – something few companies master.
  2. **Local language moats.**Threll Voice Lab trains in Norwegian – including all dialects. The data and model are unique and difficult to copy.
  3. **Strong customer lock-in (stickiness).**When the Action layer is integrated into CRM and ERP, Threll's solutions become part of the customer's core process. Switching providers becomes difficult.

In short: Voice AI combines technological complexity with direct, measurable business value.

Why developers love Voice AI

For technologists, Voice AI is where the forefront of artificial intelligence is happening right now. Here, multiple disciplines come together:

  • Signal processing – filtering, analyzing, and improving audio signals.
  • Machine learning and language models – understanding meaning and context.
  • Real-time systems – delivering responses within 500 milliseconds.
  • Integrations – orchestrating API calls, events, and data flows across systems.

Voice AI is where text AI was five years ago – just much more technically demanding and infinitely more human.

Where Threll Voice Lab begins

We start with the most challenging but most valuable part: understanding Norwegian speech. That means we train models that:

  • Understand Bokmål, Nynorsk, and dialects.
  • Tolerate background noise, echo, and overlapping speech.
  • Handle low bandwidth and Bluetooth delays.
  • Provide real-time transcription (streaming ASR) with low latency.

We are also building:

  • Dialog Manager that can withstand human interruptions.
  • Action Layer that connects to the company's systems.
  • TTS that responds naturally and confirms actual actions.

The future speaks – literally

The AI revolution started with text. But it will be completed with voice.

The text-based AI helps us write faster, while Voice AI helps us act faster – by speaking directly with our systems.

This is the future Threll Voice Lab is building. Voices that understand Norwegian – all dialects, all intonations – and do what you ask. Voices that create value, not just sound.

**Want to hear more?**Whether you are a customer, partner, investor, or developer: Threll Voice Lab invites you to join in shaping how Norway communicates with AI.

👉 Contact us in the chat for a chat.

Related Articles

Continue reading more articles