“Does it understand dialects?” is usually the first question a business asks about a voice agent. It is a good question, the answer has got noticeably better, and it still only covers one direction.
NB-Whisper, the National Library of Norway's Norwegian speech model, is trained on material including the NST corpus and recordings from the Norwegian parliament, and it returns normalised Bokmål or Nynorsk — whoever is speaking. That is not a weakness. It is the whole point: a transcript should be readable and searchable, not a rendering of how someone said it. The dialect is heard, and then it is set aside at the first step.
The resources for the other direction exist too. The same library's pronunciation lexicon, NB Uttale, holds 785,000 words transcribed in five pronunciation variants — East Norwegian, Southwest Norwegian, West Norwegian, Trøndersk and North Norwegian — built precisely for speech technology, synthesis included. NB Tale holds recordings of 380 speakers from 24 dialect areas.
And yet the voices you actually choose between when you set an agent up are, in practice, the standard variety. One option, presented as a list.





