Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
Voice interactions have a strict latency budget: capturing audio, transcribing, running an LLM, retrieving context, and generating speech all consume milliseconds before the user hears a response. TTS is the final step users notice most, and cascaded architectures—purpose-built ASR, TTS, and LLM components—offer independent tuning and deployment control compared to integrated speech models. NVIDIA Magpie Multilingual TTS is built for that approach, with open weights, production-ready NIM, and 12 languages, enabling developers to deploy multilingual speech inside their own infrastructure, optimize latency, and customize the model end to end.
The latest release expands coverage with Modern Standard Arabic, Korean, and Brazilian Portuguese, while improving quality across many existing languages via updated training data and model improvements. The model is a 364M-parameter open-weights model supporting English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, and the three new languages. Each language includes male and female voices through a shared multilingual speaker representation. The release also adds expanded code-switching support for Hindi and Japanese, enabled by IPA grapheme-to-phoneme processing and custom pronunciation dictionaries, making it easier to accurately pronounce names and mixed-language phrases.
These capabilities matter for developers building customer support agents, healthcare assistants, enterprise copilots, translation systems, and conversational AI. Open weights give teams the ability to deploy where data lives, meet enterprise privacy requirements, customize pronunciation and voices, predict latency under production workloads, and scale on their own infrastructure—turning multilingual voice AI into a buildable, tunable component rather than a black box.