{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-11T09:21:12.743Z","headline":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS","description":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","mainEntityOfPage":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","datePublished":"2026-08-10T16:25:36.000Z","dateModified":"2026-08-10T16:25:36.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents","https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn"],"canonicalUrl":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","directAnswer":{"@type":"Answer","text":"Hugging Face 发布 NVIDIA Magpie 多语言 TTS 介绍。该模型采用开放权重，支持 12 种语言，并提供 NVIDIA NIM 部署方式，面向需要自主部署、延迟优化和领域定制的实时语音应用。","url":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","dateCreated":"2026-08-10T16:25:36.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"huggingface.co source article","url":"https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents","datePublished":"2026-08-10T16:25:36.000Z","provider":{"@type":"Organization","name":"huggingface.co","url":"https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","datePublished":"2026-08-10T16:25:36.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn"}}],"aggregationSource":"Hugging Face：Blog（RSS）","originalPublisher":{"name":"huggingface.co","url":"https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents"},"geoDeepAnswer":null,"article":{"id":"cmsnh74w7078srohftrbxn3wn","slug":"cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","title":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS","title_en":"","summary":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","source":"Hugging Face：Blog（RSS）","sourceUrl":"https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents","aiHotUrl":"https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","publishedAt":"2026-08-10T16:25:36.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Voice AI Is Becoming Multilingual by Default ：#voice-ai-is-becoming-multilingual-by-default One Open Model, Twelve Languages ：#one-open-model-twelve-languages The Latency Your Users Actually Notice ：#the-latency-your-users-actually-notice Optimized for Real-Time Speech Generation ：#optimized-for-real-time-speech-generation Faster Doesn't Matter If It Doesn't Sound Natural ：#faster-doesnt-matter-if-it-doesnt-sound-natural Why Open Weights Matter ：#why-open-weights-matter Build Complete Voice Agents — Not Just Better Speech ：#build-complete-voice-agents--not-just-better-speech Get Started ：#get-started ：https://cdn-uploads.huggingface.co/production/uploads/68127936de624fb6f57d9989/6ORCRF3ze_mI5BYb6hh9u.png","Every voice interaction has a latency budget.","By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retrieving context, and generating a response. Text-to-speech (TTS) is the final step — and the one users notice most. If speech generation is slow, the whole experience feels slow.","The more of that pipeline you can run and tune yourself, the more of the latency budget you get back.","Voice AI is moving fast. Integrated speech models offer simplicity — one API call, audio in, audio out — but they trade the ability to fine-tune each component for your domain, swap in better models as they ship, enforce data residency, and understand exactly where latency is coming from. For more control, a cascaded architecture — purpose-built ASR, TTS, and LLM components running together — keeps each layer independently tunable and deployable on infrastructure you own.","NVIDIA Magpie Multilingual TTS：https://huggingface.co/nvidia/magpie_tts_multilingual_357m is built for that. With open weights, production-ready NVIDIA NIM：https://build.nvidia.com/nvidia/magpie-tts-multilingual, and support for 12 languages, you can deploy multilingual speech inside your own infrastructure, optimize latency for your workload, and customize the model for your domain — end to end, in your own environment.","The latest release expands multilingual coverage with Modern Standard Arabic, Korean, and Brazilian Portuguese, while improving quality across many existing languages through updated training data and model improvements.","Whether you're building customer support agents, healthcare assistants, enterprise copilots, translation systems, or conversational AI applications, Magpie provides an open foundation for production voice AI.","Today's voice applications don't serve a single language.","Global customer support, enterprise assistants, healthcare documentation, retail automation, and translation workflows increasingly require natural conversations across multiple languages — all while maintaining low latency.","Supporting more languages is only part of the challenge. Developers also need the ability to:","Open models change what's possible on every one of these.","Magpie TTS Multilingual is a 364M-parameter open-weights model supporting:","English · Spanish · French · German · Italian · Vietnamese · Mandarin · Hindi · Japanese · Modern Standard Arabic (new) · Korean (new) · Brazilian Portuguese (new)","Each language includes male and female speaker voices through a shared multilingual speaker representation.","This release also improves multilingual flexibility with expanded code-switching support for Hindi and Japanese, enabled through IPA grapheme-to-phoneme processing and custom pronunciation dictionaries — making it easier to accurately pronounce names, technical terminology, and mixed-language content.","Instead of maintaining separate TTS models for different regions, developers can build multilingual applications on a single open foundation.","In conversational AI, text-to-speech is the final stage before users hear a response. That makes Time to First Audio (TTFA) — the delay between speech generation beginning and the first audio reaching the user — one of the most important latency metrics in a voice pipeline.","Because Magpie TTS can be deployed inside your own environment, the latency you measure is the server-side latency you actually control, with no managed-service round-trip in the number.","Source: NVIDIA TTS NIM Performance documentation：https://docs.nvidia.com/nim/speech/26.07.0/reference/performances/tts/performance.html (v26.07), average of three trials, on-prem. TTFA = latency to first audio; RTFX = throughput as a multiple of real time.","At 32ms on B200, Magpie's TTFA leaves the rest of the latency budget for ASR and LLM processing — keeping total end-to-end latency within the sub-200ms window natural conversation requires. Across NVIDIA GPUs, Magpie delivers first audio in 32–79ms on a single stream. At 64 concurrent streams, B200 reaches 239ms TTFA while delivering throughput at 320× real time — generating audio more than 300 times faster than it plays back, even under concurrent load.","The table above shows Magpie served as the NVIDIA NIM, measured on-prem — the optimized container running on your own GPU. The open Hugging Face checkpoint is the same model and your path for research and fine-tuning; the NIM is the tuned serving stack that produces these production latencies. Both run on hardware you control.","Because the model runs on your own infrastructure, you can benchmark performance directly, tune it for your deployment, and scale according to your workload. For real-time voice agents, that's the difference between conversations that feel responsive and conversations that feel delayed.","Low latency isn't accidental. Magpie introduces two complementary architectural improvements that reduce inference time while maintaining speech quality.","Frame stacking. The decoder predicts two audio frames during each decoding step rather than one. This cuts the number of decoder iterations in half, shortening generation time and improving throughput.","Local transformer. Frame stacking alone would reduce audio quality by introducing dependencies between simultaneously generated codebook tokens. The local transformer models those dependencies and refines the generated audio, recovering the quality that frame stacking would otherwise sacrifice.","Together, these techniques deliver both faster generation and natural speech synthesis. The architecture is described in Frame-Stacked Local Transformers for Efficient Multi-Codebook Speech Generation：https://arxiv.org/abs/2509.19592 (ICASSP 2026).","This release doesn't only add languages — it also improves synthesis quality across many existing ones. Compared to the previous release, Magpie shows reduced character error rates (CER) and higher speaker similarity (SSIM) on several languages, with the clearest gains on French and Spanish:","Source: Magpie TTS Multilingual model card：https://huggingface.co/nvidia/magpie_tts_multilingual_357m. CER lower is better; SSIM higher is better.","The newly added Arabic (1.62% CER), Korean (2.69%), and Brazilian Portuguese (2.91%) models establish baseline quality for future improvements.","While objective metrics help measure progress, speech quality is ultimately perceptual. You can hear the difference yourself on NVIDIA Build：https://build.nvidia.com/nvidia/magpie-tts-multilingual or the Hugging Face demo：https://huggingface.co/spaces/nvidia/magpie_tts_multilingual_demo.","Latency you can measure is useful. Latency you can control is even better.","Open weights give developers capabilities that come from owning the deployment. With Magpie you can:","For enterprises building production voice AI, this control over deployment, performance, and customization is often what matters most.","Voice AI in production is a system of models, not a single one. Magpie TTS is part of the NVIDIA Nemotron Voice Agent Developer Example：https://build.nvidia.com/nvidia/nemotron-voice-agent, a reference implementation showing how purpose-built speech, language, and reasoning models work together as a coordinated system — so you can build always-on voice agents, not just better-sounding speech.","The Nemotron Voice Agent developer example provides an end-to-end reference implementation that developers can clone, customize, and deploy in hours. It includes production patterns for:","Rather than assembling individual components from scratch, developers can start from a complete reference architecture and adapt it to their own applications.","Recommended inference configuration:","Nvidia Text-to-Speech with MagpieTTS."],"articleImages":[{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63d69da144f1d8fbe58a3a03/N3uTYOFMxI5CoSchv70x9.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsnh74w7078srohftrbxn3wn/e85a295987123e97.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b708fa53d91a364aa32bbf/f8gnDFShkI3VBhadko8qr.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsnh74w7078srohftrbxn3wn/fe5a3a6949d7a449.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b10be888060c5a7c5423b3/FIOvkNhjeFiksYG7nn3QJ.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsnh74w7078srohftrbxn3wn/73c74b935216e549.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68127936de624fb6f57d9989/_eCAtnPorf7jPrW1luUg4.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmsnh74w7078srohftrbxn3wn/f9c0e250dc4b5fe9.webp"}],"mediaStatus":"ok","articleBodyZh":["语音 AI 正在默认支持多语言：#voice-ai-is-becoming-multilingual-by-default 一个开放模型，支持十二种语言：#one-open-model-twelve-languages 用户实际感知的延迟：#the-latency-your-users-actually-notice 针对实时语音生成优化：#optimized-for-real-time-speech-generation 更快无意义，如果听起来不自然：#faster-doesnt-matter-if-it-doesnt-sound-natural 为什么开放权重很重要：#why-open-weights-matter 构建完整的语音代理——不仅仅是更好的语音：#build-complete-voice-agents--not-just-better-speech 开始使用：#get-started：https://cdn-uploads.huggingface.co/production/uploads/68127936de624fb6f57d9989/6ORCRF3ze_mI5BYb6hh9u.png","每一次语音交互都有一个延迟预算。","当用户听到你的应用响应时，你已经花费了宝贵的毫秒来捕捉音频、转录语音、运行大型语言模型 (LLM)、检索上下文以及生成响应。文本转语音 (TTS) 是最后一步——也是用户最注意到的一步。如果语音生成速度慢，整个体验就显得缓慢。","你自己能够运行和调优的管道越多，你就能从延迟预算中获得更多回报。","语音 AI 正在快速发展。集成语音模型提供了简便性——一次 API 调用，输入音频，输出音频——但它们牺牲了在你的领域微调每个组件的能力、在模型发布时替换更好的模型、强制数据驻留以及精确了解延迟来源。为了获得更多控制，级联架构——专用的 ASR、TTS 和 LLM 组件一起运行——可以保持每一层独立可调，并可部署在你自己的基础设施上。","NVIDIA Magpie 多语言 TTS：https://huggingface.co/nvidia/magpie_tts_multilingual_357m 就是为此而构建的。通过开放权重、生产就绪的 NVIDIA NIM：https://build.nvidia.com/nvidia/magpie-tts-multilingual，以及对 12 种语言的支持，你可以在自己的基础设施内部署多语言语音，为你的工作负载优化延迟，并为你的领域定制模型——从头到尾，完全在你自己的环境中。","最新版本通过加入现代标准阿拉伯语、韩语和巴西葡萄牙语来扩展多语言覆盖，同时通过更新的训练数据和模型改进提高了许多现有语言的质量。","无论您是在构建客户支持代理、医疗助理、企业辅助工具、翻译系统还是会话式 AI 应用，Magpie 都为生产级语音 AI 提供了开放的基础。","今天的语音应用不只服务单一语言。","全球客户支持、企业助手、医疗文档、零售自动化和翻译工作流越来越多地需要跨多语言的自然对话——同时保持低延迟。","支持更多语言只是挑战的一部分。开发者还需要具备以下能力：","开放模型改变了在每一种应用中可能实现的功能。","Magpie TTS 多语言模型是一个 3.64 亿参数的开放权重模型，支持：","英语 · 西班牙语 · 法语 · 德语 · 意大利语 · 越南语 · 普通话 · 印地语 · 日语 · 现代标准阿拉伯语（新增） · 韩语（新增） · 巴西葡萄牙语（新增）","每种语言都通过共享的多语言说话人表示提供男性和女性的语音。","此版本还通过支持印度语和日语的扩展代码切换，提升了多语言灵活性。这通过 IPA 字母-音素处理和自定义发音词典实现，使准确发音名称、技术术语和混合语言内容变得更容易。","开发者无需为不同地区维护独立的 TTS 模型，可以在单一开放基础上构建多语言应用。","在会话式 AI 中，文本转语音是用户听到响应前的最后阶段。这使得首次音频时间（TTFA）——从语音生成开始到第一段音频到达用户的延迟——成为语音管道中最重要的延迟指标之一。","由于 Magpie TTS 可以部署在您自己的环境中，您测量的延迟就是您实际可控的服务器端延迟，无需通过托管服务的往返。","来源：NVIDIA TTS NIM 性能文档：https://docs.nvidia.com/nim/speech/26.07.0/reference/performances/tts/performance.html（v26.07），三次试验平均结果，本地部署。TTFA = 首个音频延迟；RTFX = 实时倍数的吞吐量。","在 B200 上以 32ms 测得，Magpie 的 TTFA 将剩余的延迟预算留给 ASR 和 LLM 处理 —— 保持端到端总延迟在自然对话所需的 200ms 以下。跨 NVIDIA GPU，Magpie 在单条流上可在 32–79ms 内生成首个音频。在 64 条并发流下，B200 的 TTFA 达到 239ms，同时吞吐量为实时的 320 倍 —— 即使在并发负载下，也能生成音频的速度超过播放速度 300 倍以上。","上表显示 Magpie 作为 NVIDIA NIM 提供服务，并在本地测量 —— 即在您自己的 GPU 上运行的优化容器。开源的 Hugging Face 检查点是相同的模型，也是您进行研究和微调的路径；而 NIM 是调优的服务堆栈，可产生这些生产环境的延迟。两者都运行在您控制的硬件上。","因为模型运行在您自己的基础设施上，您可以直接进行性能基准测试，为部署进行调优，并根据工作负载进行扩展。对于实时语音代理，这决定了对话是感觉响应及时还是延迟感明显。","低延迟不是偶然的。Magpie 引入了两个互补的架构改进，以减少推理时间，同时保持语音质量。","帧堆叠。解码器在每个解码步骤中预测两个音频帧，而不是一个。这将解码器迭代次数减少一半，缩短生成时间并提高吞吐量。","局部 Transformer。单独的帧堆叠会通过在同时生成的码本 token 之间引入依赖关系而降低音频质量。局部 Transformer 对这些依赖关系进行建模并优化生成的音频，恢复帧堆叠可能牺牲的音质。","这些技术结合起来既能加快生成速度，又能实现自然语音合成。该架构在《Frame-Stacked Local Transformers for Efficient Multi-Codebook Speech Generation》中有详细描述：https://arxiv.org/abs/2509.19592（ICASSP 2026）。","此次发布不仅增加了语言支持，还提升了多种已有语言的合成质量。与上一个版本相比，Magpie 在多种语言上的字符错误率（CER）降低，讲话者相似度（SSIM）提高，其中法语和西班牙语的提升最为明显：","来源：Magpie TTS 多语言模型卡：https://huggingface.co/nvidia/magpie_tts_multilingual_357m。CER 越低越好；SSIM 越高越好。","新增加的阿拉伯语（1.62% CER）、韩语（2.69%）和巴西葡萄牙语（2.91%）模型为未来的改进建立了基线质量。","虽然客观指标有助于衡量进展，但语音质量最终取决于感知。你可以在 NVIDIA Build 上亲自感受差异：https://build.nvidia.com/nvidia/magpie-tts-multilingual 或 Hugging Face 演示：https://huggingface.co/spaces/nvidia/magpie_tts_multilingual_demo。","可测量的延迟很有用。可控制的延迟更佳。","开放权重赋予开发者来自部署拥有的能力。使用 Magpie，你可以：","对于建设生产语音 AI 的企业而言，能够控制部署、性能和定制化通常是最重要的。","生产中的语音 AI 是一个由多个模型组成的系统，而非单一模型。Magpie TTS 是 NVIDIA Nemotron Voice Agent 开发示例的一部分：https://build.nvidia.com/nvidia/nemotron-voice-agent，这是一个参考实现，展示了专用语音、语言和推理模型如何协调工作——因此你可以构建全天候语音代理，而不仅仅是更好听的语音。","Nemotron Voice Agent 开发示例提供了一个端到端的参考实现，开发者可以克隆、定制并在数小时内部署。它包含用于生产的模式：","开发者无需从零组装各个组件，可以从完整的参考架构出发，并根据自己的应用进行适配。","推荐的推理配置：","使用 MagpieTTS 的 Nvidia 文本转语音。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Hugging Face 发布 NVIDIA Magpie 多语言 TTS 介绍。该模型采用开放权重，支持 12 种语言，并提供 NVIDIA NIM 部署方式，面向需要自主部署、延迟优化和领域定制的实时语音应用。","background":"文章将语音交互延迟拆分为音频采集、语音识别、大语言模型处理、上下文检索和语音生成等环节，并指出 TTS 位于用户感知链路末端。其强调级联架构可分别调优和部署 ASR、TTS 与 LLM。","viewpoint":"Aioga 判断，Magpie 的重点不只是增加语言覆盖，而是把开放权重、可控部署和组件级调优结合起来。对于重视数据驻留或希望掌握延迟来源的团队，这种架构可能更具吸引力。","implications":"该方案可能降低企业构建多语言语音应用时对单一集成式接口的依赖，同时保留替换和调优各组件的空间。文章提到的客户支持、医疗助手、企业副驾驶、翻译和对话式应用，均可作为潜在使用方向，但实际效果仍取决于部署环境和业务适配。","nextStep":"值得关注的是，团队应依据自身语音场景评估 12 种语言的覆盖需求、实时生成表现、部署基础设施和领域定制空间，并进一步核验开放权重模型与 NVIDIA NIM 在目标环境中的适配情况。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-10T22:05:08.546Z","sourceHash":"8fef659e4f9d052d","review":{"approved":true,"groundedness":95,"clarity":91,"duplicationRisk":18,"blockingIssues":[],"notes":["“Hugging Face 发布 NVIDIA Magpie 多语言 TTS 介绍”可更精确地表述为“Hugging Face 博客刊发 NVIDIA Magpie 多语言 TTS 介绍”，以避免被理解为 Hugging Face 发布了该模型；此处不构成事实性阻断问题。","“可能降低对单一集成式接口的依赖”和“实际效果仍取决于部署环境和业务适配”属于基于来源内容作出的审慎推论，候选内容已使用“可能”“仍取决于”等限定语。","viewpoint 中的判断被明确标注为“Aioga 判断”，没有将观点冒充来源事实。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","Hugging Face：Blog（RSS）"],"translations":{"zh-CN":{"title":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS","summary":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"huggingface.co","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS - Aioga AI资讯","description":"🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/news/cmsnh74w7078srohftrbxn3wn/","articleBody":["语音 AI 正在默认支持多语言：#voice-ai-is-becoming-multilingual-by-default 一个开放模型，支持十二种语言：#one-open-model-twelve-languages 用户实际感知的延迟：#the-latency-your-users-actually-notice 针对实时语音生成优化：#optimized-for-real-time-speech-generation 更快无意义，如果听起来不自然：#faster-doesnt-matter-if-it-doesnt-sound-natural 为什么开放权重很重要：#why-open-weights-matter 构建完整的语音代理——不仅仅是更好的语音：#build-complete-voice-agents--not-just-better-speech 开始使用：#get-started：https://cdn-uploads.huggingface.co/production/uploads/68127936de624fb6f57d9989/6ORCRF3ze_mI5BYb6hh9u.png","每一次语音交互都有一个延迟预算。","当用户听到你的应用响应时，你已经花费了宝贵的毫秒来捕捉音频、转录语音、运行大型语言模型 (LLM)、检索上下文以及生成响应。文本转语音 (TTS) 是最后一步——也是用户最注意到的一步。如果语音生成速度慢，整个体验就显得缓慢。","你自己能够运行和调优的管道越多，你就能从延迟预算中获得更多回报。","语音 AI 正在快速发展。集成语音模型提供了简便性——一次 API 调用，输入音频，输出音频——但它们牺牲了在你的领域微调每个组件的能力、在模型发布时替换更好的模型、强制数据驻留以及精确了解延迟来源。为了获得更多控制，级联架构——专用的 ASR、TTS 和 LLM 组件一起运行——可以保持每一层独立可调，并可部署在你自己的基础设施上。","NVIDIA Magpie 多语言 TTS：https://huggingface.co/nvidia/magpie_tts_multilingual_357m 就是为此而构建的。通过开放权重、生产就绪的 NVIDIA NIM：https://build.nvidia.com/nvidia/magpie-tts-multilingual，以及对 12 种语言的支持，你可以在自己的基础设施内部署多语言语音，为你的工作负载优化延迟，并为你的领域定制模型——从头到尾，完全在你自己的环境中。","最新版本通过加入现代标准阿拉伯语、韩语和巴西葡萄牙语来扩展多语言覆盖，同时通过更新的训练数据和模型改进提高了许多现有语言的质量。","无论您是在构建客户支持代理、医疗助理、企业辅助工具、翻译系统还是会话式 AI 应用，Magpie 都为生产级语音 AI 提供了开放的基础。","今天的语音应用不只服务单一语言。","全球客户支持、企业助手、医疗文档、零售自动化和翻译工作流越来越多地需要跨多语言的自然对话——同时保持低延迟。","支持更多语言只是挑战的一部分。开发者还需要具备以下能力：","开放模型改变了在每一种应用中可能实现的功能。","Magpie TTS 多语言模型是一个 3.64 亿参数的开放权重模型，支持：","英语 · 西班牙语 · 法语 · 德语 · 意大利语 · 越南语 · 普通话 · 印地语 · 日语 · 现代标准阿拉伯语（新增） · 韩语（新增） · 巴西葡萄牙语（新增）","每种语言都通过共享的多语言说话人表示提供男性和女性的语音。","此版本还通过支持印度语和日语的扩展代码切换，提升了多语言灵活性。这通过 IPA 字母-音素处理和自定义发音词典实现，使准确发音名称、技术术语和混合语言内容变得更容易。","开发者无需为不同地区维护独立的 TTS 模型，可以在单一开放基础上构建多语言应用。","在会话式 AI 中，文本转语音是用户听到响应前的最后阶段。这使得首次音频时间（TTFA）——从语音生成开始到第一段音频到达用户的延迟——成为语音管道中最重要的延迟指标之一。","由于 Magpie TTS 可以部署在您自己的环境中，您测量的延迟就是您实际可控的服务器端延迟，无需通过托管服务的往返。","来源：NVIDIA TTS NIM 性能文档：https://docs.nvidia.com/nim/speech/26.07.0/reference/performances/tts/performance.html（v26.07），三次试验平均结果，本地部署。TTFA = 首个音频延迟；RTFX = 实时倍数的吞吐量。","在 B200 上以 32ms 测得，Magpie 的 TTFA 将剩余的延迟预算留给 ASR 和 LLM 处理 —— 保持端到端总延迟在自然对话所需的 200ms 以下。跨 NVIDIA GPU，Magpie 在单条流上可在 32–79ms 内生成首个音频。在 64 条并发流下，B200 的 TTFA 达到 239ms，同时吞吐量为实时的 320 倍 —— 即使在并发负载下，也能生成音频的速度超过播放速度 300 倍以上。","上表显示 Magpie 作为 NVIDIA NIM 提供服务，并在本地测量 —— 即在您自己的 GPU 上运行的优化容器。开源的 Hugging Face 检查点是相同的模型，也是您进行研究和微调的路径；而 NIM 是调优的服务堆栈，可产生这些生产环境的延迟。两者都运行在您控制的硬件上。","因为模型运行在您自己的基础设施上，您可以直接进行性能基准测试，为部署进行调优，并根据工作负载进行扩展。对于实时语音代理，这决定了对话是感觉响应及时还是延迟感明显。","低延迟不是偶然的。Magpie 引入了两个互补的架构改进，以减少推理时间，同时保持语音质量。","帧堆叠。解码器在每个解码步骤中预测两个音频帧，而不是一个。这将解码器迭代次数减少一半，缩短生成时间并提高吞吐量。","局部 Transformer。单独的帧堆叠会通过在同时生成的码本 token 之间引入依赖关系而降低音频质量。局部 Transformer 对这些依赖关系进行建模并优化生成的音频，恢复帧堆叠可能牺牲的音质。","这些技术结合起来既能加快生成速度，又能实现自然语音合成。该架构在《Frame-Stacked Local Transformers for Efficient Multi-Codebook Speech Generation》中有详细描述：https://arxiv.org/abs/2509.19592（ICASSP 2026）。","此次发布不仅增加了语言支持，还提升了多种已有语言的合成质量。与上一个版本相比，Magpie 在多种语言上的字符错误率（CER）降低，讲话者相似度（SSIM）提高，其中法语和西班牙语的提升最为明显：","来源：Magpie TTS 多语言模型卡：https://huggingface.co/nvidia/magpie_tts_multilingual_357m。CER 越低越好；SSIM 越高越好。","新增加的阿拉伯语（1.62% CER）、韩语（2.69%）和巴西葡萄牙语（2.91%）模型为未来的改进建立了基线质量。","虽然客观指标有助于衡量进展，但语音质量最终取决于感知。你可以在 NVIDIA Build 上亲自感受差异：https://build.nvidia.com/nvidia/magpie-tts-multilingual 或 Hugging Face 演示：https://huggingface.co/spaces/nvidia/magpie_tts_multilingual_demo。","可测量的延迟很有用。可控制的延迟更佳。","开放权重赋予开发者来自部署拥有的能力。使用 Magpie，你可以：","对于建设生产语音 AI 的企业而言，能够控制部署、性能和定制化通常是最重要的。","生产中的语音 AI 是一个由多个模型组成的系统，而非单一模型。Magpie TTS 是 NVIDIA Nemotron Voice Agent 开发示例的一部分：https://build.nvidia.com/nvidia/nemotron-voice-agent，这是一个参考实现，展示了专用语音、语言和推理模型如何协调工作——因此你可以构建全天候语音代理，而不仅仅是更好听的语音。","Nemotron Voice Agent 开发示例提供了一个端到端的参考实现，开发者可以克隆、定制并在数小时内部署。它包含用于生产的模式：","开发者无需从零组装各个组件，可以从完整的参考架构出发，并根据自己的应用进行适配。","推荐的推理配置：","使用 MagpieTTS 的 Nvidia 文本转语音。"]},"en":{"title":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS","summary":"🔗 Read the original article via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"Industry","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Build Low-Latency Multilingual Voice Agents： Open Weights & Full Deployment Control with NVIDIA Magpie TTS - Aioga AI News","description":"🔗 Read the original article via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/en/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:16.947Z"},"ja":{"title":"低遅延の多言語音声エージェントを構築する:NVIDIA Magpie TTSによるオープンウェイトとフルデプロイメントコントロール","summary":"🔗 原文記事はAIHOTより読むことができます。 https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"業界動向","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"低遅延の多言語音声エージェントを構築する:NVIDIA Magpie TTSによるオープンウェイトとフルデプロイメントコントロール - Aioga AIニュース","description":"🔗 原文記事はAIHOTより読むことができます。 https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/ja/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:17.226Z"},"ko":{"title":"NVIDIA Magpie TTS와 함께 저지연 다국어 음성 에이전트 구축: 오픈 가중치 및 완전한 배포 제어","summary":"🔗 원문 기사는 AIHOT를 통해 읽을 수 있습니다. https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"업계 동향","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"NVIDIA Magpie TTS와 함께 저지연 다국어 음성 에이전트 구축: 오픈 가중치 및 완전한 배포 제어 - Aioga AI 뉴스","description":"🔗 원문 기사는 AIHOT를 통해 읽을 수 있습니다. https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/ko/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:25.949Z"},"es":{"title":"Crea agentes de voz multilingües de baja latencia: Openweights y control total de despliegue con NVIDIA Magpie TTS","summary":"🔗 Lee el artículo original a través de AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"Industria","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Crea agentes de voz multilingües de baja latencia: Openweights y control total de despliegue con NVIDIA Magpie TTS - Aioga Noticias de IA","description":"🔗 Lee el artículo original a través de AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/es/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:25.812Z"},"fr":{"title":"Créer des agents vocaux multilingues à faible latence : poids ouverts et contrôle complet du déploiement avec NVIDIA Magpie TTS","summary":"🔗 Lisez l’article original via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"Industrie","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Créer des agents vocaux multilingues à faible latence : poids ouverts et contrôle complet du déploiement avec NVIDIA Magpie TTS - Aioga Actualités IA","description":"🔗 Lisez l’article original via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/fr/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:35.040Z"},"de":{"title":"Mehrsprachige Sprachagenten mit niedriger Latenz bauen: Offene Gewichte & vollständige Bereitstellungskontrolle mit NVIDIA Magpie TTS","summary":"🔗 Lesen Sie den Originalartikel über AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Mehrsprachige Sprachagenten mit niedriger Latenz bauen: Offene Gewichte & vollständige Bereitstellungskontrolle mit NVIDIA Magpie TTS - Aioga KI-News","description":"🔗 Lesen Sie den Originalartikel über AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/de/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:33.704Z"},"pt-BR":{"title":"Construa Agentes de Voz Multilíngues de Baixa Latência: Pesos Abertos e Controle Total de Implantação com NVIDIA Magpie TTS","summary":"🔗 Leia o artigo original via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Construa Agentes de Voz Multilíngues de Baixa Latência: Pesos Abertos e Controle Total de Implantação com NVIDIA Magpie TTS - Aioga Notícias de IA","description":"🔗 Leia o artigo original via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/pt-BR/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:43.277Z"},"ru":{"title":"Создайте многоязычные голосовые агенты с низкой задержкой: открытые веса и полное управление развертыванием с помощью NVIDIA Magpie TTS","summary":"🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Создайте многоязычные голосовые агенты с низкой задержкой: открытые веса и полное управление развертыванием с помощью NVIDIA Magpie TTS - Aioga Новости ИИ","description":"🔗 Прочитайте оригинальную статью на сайте AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/ru/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:43.794Z"},"ar":{"title":"بناء وكلاء صوتيين متعددي اللغات منخفض التأخير: أوزان مفتوحة وتحكم كامل في النشر باستخدام NVIDIA Magpie TTS","summary":"🔗 اقرأ المقال الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"بناء وكلاء صوتيين متعددي اللغات منخفض التأخير: أوزان مفتوحة وتحكم كامل في النشر باستخدام NVIDIA Magpie TTS - Aioga أخبار الذكاء الاصطناعي","description":"🔗 اقرأ المقال الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/ar/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:52.868Z"},"hi":{"title":"लो-लेटेंसी बहुभाषी वॉयस एजेंटों का निर्माण करें: एनवीडिया मैगपाई टीटीएस के साथ ओपन वेट और फुल डिप्लॉयमेंट कंट्रोल","summary":"🔗 AIHOT के माध्यम से मूल लेख पढ़ें · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"लो-लेटेंसी बहुभाषी वॉयस एजेंटों का निर्माण करें: एनवीडिया मैगपाई टीटीएस के साथ ओपन वेट और फुल डिप्लॉयमेंट कंट्रोल - Aioga AI समाचार","description":"🔗 AIHOT के माध्यम से मूल लेख पढ़ें · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/hi/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:02:53.076Z"},"it":{"title":"Crea agenti vocali multilingue a bassa latenza: pesi aperti e controllo completo del deployment con NVIDIA Magpie TTS","summary":"🔗 Leggi l'articolo originale su AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Crea agenti vocali multilingue a bassa latenza: pesi aperti e controllo completo del deployment con NVIDIA Magpie TTS - Aioga Notizie IA","description":"🔗 Leggi l'articolo originale su AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/it/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:01.688Z"},"nl":{"title":"Bouw laag-latentie meertalige spraakagenten: Open gewichten & volledige implementatiecontrole met NVIDIA Magpie TTS","summary":"🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Bouw laag-latentie meertalige spraakagenten: Open gewichten & volledige implementatiecontrole met NVIDIA Magpie TTS - Aioga AI-nieuws","description":"🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/nl/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:01.138Z"},"tr":{"title":"NVIDIA Magpie TTS ile Düşük Gecikmeli Çok Dilli Ses Ajanları Oluşturun: Açık Ağırlıklar ve Tam Dağıtım Kontrolü","summary":"🔗 Orijinal makaleyi AIHOT üzerinden okuyun · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"NVIDIA Magpie TTS ile Düşük Gecikmeli Çok Dilli Ses Ajanları Oluşturun: Açık Ağırlıklar ve Tam Dağıtım Kontrolü - Aioga AI Haberleri","description":"🔗 Orijinal makaleyi AIHOT üzerinden okuyun · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/tr/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:10.074Z"},"vi":{"title":"Xây dựng các tác nhân giọng nói đa ngôn ngữ với độ trễ thấp: Trọng số mở & Kiểm soát triển khai đầy đủ với NVIDIA Magpie TTS","summary":"🔗 Đọc bài viết gốc qua AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Xây dựng các tác nhân giọng nói đa ngôn ngữ với độ trễ thấp: Trọng số mở & Kiểm soát triển khai đầy đủ với NVIDIA Magpie TTS - Tin tức AI Aioga","description":"🔗 Đọc bài viết gốc qua AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/vi/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:10.302Z"},"id":{"title":"Bangun Agen Suara Multibahasa dengan Latensi Rendah: Bobot Terbuka & Kontrol Penyebaran Penuh dengan NVIDIA Magpie TTS","summary":"🔗 Baca artikel asli melalui AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Bangun Agen Suara Multibahasa dengan Latensi Rendah: Bobot Terbuka & Kontrol Penyebaran Penuh dengan NVIDIA Magpie TTS - Berita AI Aioga","description":"🔗 Baca artikel asli melalui AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/id/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:18.906Z"},"th":{"title":"สร้างเอเจนต์เสียงหลายภาษาที่มีความหน่วงต่ํา: น้ําหนักเปิดและการควบคุมการติดตั้งเต็มรูปแบบด้วย NVIDIA Magpie TTS","summary":"🔗 อ่านบทความต้นฉบับผ่าน AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"สร้างเอเจนต์เสียงหลายภาษาที่มีความหน่วงต่ํา: น้ําหนักเปิดและการควบคุมการติดตั้งเต็มรูปแบบด้วย NVIDIA Magpie TTS - ข่าว AI Aioga","description":"🔗 อ่านบทความต้นฉบับผ่าน AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/th/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:19.011Z"},"pl":{"title":"Buduj wielojęzyczne agenty głosowe o niskich opóźnieniach: otwarte wagi i pełną kontrolę wdrożenia z NVIDIA Magpie TTS","summary":"🔗 Przeczytaj oryginalny artykuł za pośrednictwem AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Buduj wielojęzyczne agenty głosowe o niskich opóźnieniach: otwarte wagi i pełną kontrolę wdrożenia z NVIDIA Magpie TTS - Aioga Wiadomości AI","description":"🔗 Przeczytaj oryginalny artykuł za pośrednictwem AIHOT · https://aihot.virxact.com/items/cmsnh74w7078srohftrbxn3wn","url":"https://www.aioga.com/pl/news/cmsnh74w7078srohftrbxn3wn/","contentTranslated":true,"sourceHash":"2749371b34b9d292","translatedAt":"2026-08-10T18:03:27.690Z"}}}}