{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-21T16:01:09.055Z","headline":"Meta Superintelligence Labs 发布 Muse Voice Transcribe：一个实时模型同时完成流式 ASR、说话人分离和端点检测","description":"Meta Superintelligence Labs 发布首个实时音频感知模型 Muse Voice Transcribe，将流式 ASR、20+ 说话人的说话人分离和端点检测合并为一个自回归模型，无需后处理。","url":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","mainEntityOfPage":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","datePublished":"2026-09-02T05:37:12.000Z","dateModified":"2026-09-02T05:37:12.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing","https://aihot.virxact.com/items/cmtjoufgu029rrotnspaovulr"],"canonicalUrl":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","directAnswer":{"@type":"Answer","text":"Meta Superintelligence Labs 发布 Muse Voice Transcribe，将流式自动语音识别、面向20多个说话人的说话人分离与端点检测合并到一个自回归模型中，并称无需后处理。","url":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","dateCreated":"2026-09-02T05:37:12.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing","datePublished":"2026-09-02T05:37:12.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmtjoufgu029rrotnspaovulr","datePublished":"2026-09-02T05:37:12.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmtjoufgu029rrotnspaovulr"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing"},"geoDeepAnswer":null,"article":{"id":"cmtjoufgu029rrotnspaovulr","slug":"cmtjoufgu029rrotnspaovulr","url":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","title":"Meta Superintelligence Labs 发布 Muse Voice Transcribe：一个实时模型同时完成流式 ASR、说话人分离和端点检测","title_en":"","summary":"Meta Superintelligence Labs 发布首个实时音频感知模型 Muse Voice Transcribe，将流式 ASR、20+ 说话人的说话人分离和端点检测合并为一个自回归模型，无需后处理。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing","aiHotUrl":"https://aihot.virxact.com/items/cmtjoufgu029rrotnspaovulr","publishedAt":"2026-09-02T05:37:12.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode.","Muse Voice Transcribe：https://research.meta.ai/blog/introducing-muse-voice-transcribe, announced by Meta Superintelligence Labs this week, collapses those three jobs into a single autoregressive model. Meta calls it its first real-time audio perception model. It performs streaming ASR, speaker diarization for 20+ speakers, and endpointing in one pass, with no required post-processing.","Is it deployable? Yes, but only as a hosted API. It is live on the Meta Model API：https://developer.meta.com/ai/models/muse-voice-transcribe as muse-voice-transcribe-1.0 at $3.00 per 1,000 audio minutes ($0.18 per hour), and it already powers dictation in Meta AI for Mac and Muse Code：https://developer.meta.com/ai/products/muse-code. No weights have been released, so there is no self-hosted path.","Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family. Audio arrives in 80ms chunks at 12.5 Hz. Each chunk is transformed into a single soft token.","After every chunk the model makes one binary choice. It either predicts a token and keeps listening, or it emits a text token. When the model predicts , that token is replaced by the actual next audio chunk in the input. When the stream ends, an token is inserted, and the model flushes all remaining text without requesting more audio.","Listening and writing share one decoder loop, so there is no separate alignment stage to drift.","Because the model controls when it listens, it also controls how much audio context sits behind each word. Meta calls that gap ‘delay.’ Longer delay means a more accurate transcript and higher latency.","Instead of fixing that trade-off, Meta trains it. Reinforcement learning combines a word error rate reward and a delay reward multiplicatively, producing a policy that varies delay per word by difficulty. Meta reports this puts the model on the Pareto front for speed against accuracy, measured by time to final transcription, ahead of the previous frontier formed by Soniox, Cartesia, and ElevenLabs systems.","Meta did not add a second model for speaker attribution. It added special tokens to the same stream.","For diarization, a token marks a potential speaker switch, and a tag identifies the speaker. The turn token fires as soon as a switch is possible, while the speaker tag is delayed to the end of the chunk. Audio from one speaker can be split across several segments that all resolve to the same tag.","For endpointing, marks the start of speech and marks the point where the user finished. Both tasks are trained jointly with streaming ASR, using extra rewards layered on top of the ASR reward.","The model was trained on 70+ languages, of which 25 are extensively verified and recommended at launch. Code-switching is native, both within a sentence and between sentences, which matters for bilingual speakers who mix languages mid-clause. Accuracy can be improved further with language, keyword, and context biasing.","Long-context handling is a practical differentiator. Meta states the model natively supports audio input exceeding one hour and 20+ speakers, with no required post-processing step.","Meta reports first place on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks, as of September 1, 2026.","On Artificial Analysis AA-WER Streaming：https://x.com/ArtificialAnlys/status/2094849283120128135, Muse Voice Transcribe records 3.1% final-transcript WER at 0.16s after end of speech. Cartesia Ink-2 with semantic endpoints is 3.4% at 0.43s. ElevenLabs Scribe v2 Realtime is 3.6% at 0.14s. Cartesia Ink-2 with external endpoints is fastest at 0.07s but least accurate at 4.0%. On first partial transcript, Muse Voice Transcribe records 3.6% WER at 0.13s.","On diarization, Meta reports a 17.5% average diarization error rate across AMI-IHM, AMI-SDM, and VoxConverse. Five other systems in the same chart range from 21.1% to 28.6%.","Price is the other axis. At $3.00 per 1,000 minutes, it undercuts Cartesia Ink-2 at $4.00 and is less than half the $6.50 for ElevenLabs Scribe v2 Realtime and Deepgram Flux.","Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us ：https://forms.gle/wbash1wF6efRj8G58","Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2025/07/a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg-300x300.png","alt":"","afterParagraph":17,"url":"/media/articles/cmtjoufgu029rrotnspaovulr/84e64b03066de40c.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/09/blog2222-100x70.png","alt":"Researchers from Princeton, Ant Group and Stanford Introduce AQuA","afterParagraph":18,"url":"/media/articles/cmtjoufgu029rrotnspaovulr/e783ad9836537901.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/08/blog2211-3-100x70.png","alt":"Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour","afterParagraph":18,"url":"/media/articles/cmtjoufgu029rrotnspaovulr/29f802a690594064.webp"}],"mediaStatus":"ok","articleBodyZh":["大多数生产语音堆栈是由三个系统拼接在一起的。一个模型进行转录，第二个模型区分说话者，检测器判断用户何时停止说话。每一次交接都会增加延迟并产生新的故障模式。","Muse Voice Transcribe：https://research.meta.ai/blog/introducing-muse-voice-transcribe，Meta Superintelligence Labs 本周宣布，将这三项工作合并为一个自回归模型。 Meta 称其为首个实时音频感知模型。它可以一次性执行流式 ASR、20 多个说话者的说话者分类和终结点，无需进行后处理。","它可以部署吗？可以，但仅限于托管API。它已经在Meta模型API上上线：https://developer.meta.com/ai/models/muse-voice-transcribe，型号为muse-voice-transcribe-1.0，价格为每1000分钟音频3美元（每小时0.18美元），并且已经为Mac上的Meta AI和Muse Code提供语音输入功能：https://developer.meta.com/ai/products/muse-code。模型权重尚未公开，因此没有自托管的途径。","Muse Voice Transcribe是Muse Spark系列中的一个自回归多模态模型。音频以12.5 Hz的频率以80ms的块到达。每个音频块被转换为一个软令牌。","每处理完一个音频块，模型会做出一个二元选择。它要么预测一个令牌并继续监听，要么输出一个文本令牌。当模型预测时，该令牌会被输入中的下一个实际音频块替换。当流结束时，会插入一个令牌，模型会在不请求更多音频的情况下刷新所有剩余文本。","听和写共享一个解码器循环，因此没有单独的对齐阶段出现漂移。","因为模型控制何时监听，它也控制每个单词后面有多少音频上下文。Meta称这一间隔为“延迟”。更长的延迟意味着转录更准确，但延迟更高。","Meta没有固定这种权衡，而是通过训练来调节。强化学习将词错误率奖励和延迟奖励相乘，生成一种策略，使每个单词的延迟根据难度而变化。Meta报告称，这使得模型在最终转录所需时间上，在速度与准确率的权衡上处于帕累托前沿，超过了由Soniox、Cartesia和ElevenLabs系统形成的先前前沿。","Meta 没有为说话人归属添加第二个模型，而是在同一流中添加了特殊标记。","对于分说话人标注，一个标记表示潜在的说话人切换，标签标识说话人。轮次标记在一旦可能发生切换时触发，而说话人标签延迟到该段落结束。一个说话人的音频可以被分割成多个片段，但都对应同一个标签。","对于端点检测，标记表示语音的开始，标记表示用户结束的点。两个任务与流式 ASR 联合训练，使用在 ASR 奖励基础上添加的额外奖励。","该模型在 70 多种语言上训练，其中 25 种在发布时经过广泛验证并推荐使用。代码切换是原生支持的，包括句子内和句子间，对于混合语言的双语说话人尤为重要。准确度可以通过语言、关键词和上下文偏置进一步提高。","长上下文处理是一个实际区分点。Meta 表示模型原生支持超过一小时的音频输入和 20 多个说话人，无需额外后处理步骤。","截至 2026 年 9 月 1 日，Meta 报告在流式语音转文本的人工分析和公开分说话人基准中获得第一名。","在人工分析 AA-WER 流式测评中：https://x.com/ArtificialAnlys/status/2094849283120128135，Muse Voice Transcribe 在语音结束 0.16 秒后最终抄本 WER 为 3.1%。使用语义端点的 Cartesia Ink-2 为 3.4%，延迟 0.43 秒。ElevenLabs Scribe v2 Realtime 为 3.6%，延迟 0.14 秒。使用外部端点的 Cartesia Ink-2 速度最快，仅需 0.07 秒，但准确度最低，为 4.0%。在首次部分抄本中，Muse Voice Transcribe 在 0.13 秒时记录 WER 为 3.6%。","在分说话人方面，Meta 报告在 AMI-IHM、AMI-SDM 和 VoxConverse 中的平均分说话人错误率为 17.5%。同一图表中的其他五个系统范围为 21.1% 至 28.6%。","价格是另一个维度。每 1,000 分钟 3.00 美元，比 Cartesia Ink-2 的 4.00 美元低，还不到 ElevenLabs Scribe v2 Realtime 和 Deepgram Flux 的 6.50 美元的一半。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等？请联系我们：https://forms.gle/wbash1wF6efRj8G58","米哈尔·萨特（Michal Sutter）是一名数据科学专业人士，拥有帕多瓦大学的数据科学硕士学位。凭借在统计分析、机器学习和数据工程方面的扎实基础，米哈尔擅长将复杂的数据集转化为可操作的洞察。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Meta Superintelligence Labs 发布 Muse Voice Transcribe，将流式自动语音识别、面向20多个说话人的说话人分离与端点检测合并到一个自回归模型中，并称无需后处理。","background":"该模型以80毫秒音频块接收输入，在同一解码循环中决定继续监听或输出文本。材料称其目前以托管API形式提供，型号为muse-voice-transcribe-1.0，未发布模型权重。","viewpoint":"Aioga 判断：Muse Voice Transcribe 的核心变化在于把语音转写、说话人区分和结束判断放进同一实时处理流程，减少多系统衔接这一产品架构特征值得关注；但材料未提供独立验证结果。","implications":"可能影响：统一模型可能减少传统语音栈中的组件交接，但不代表部署复杂度、转写错误或延迟问题已经消失。使用方需要评估托管API依赖、成本、数据处理要求及准确率与实时性的取舍。","nextStep":"后续观察：应关注Meta是否公布更完整的评测方法、不同场景下的错误与延迟表现，以及模型权重或更多部署方式是否开放；同时核实其报告的性能比较能否被独立复现。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-09-02T06:25:09.435Z","sourceHash":"301b6c7d6fa41829","review":{"approved":true,"groundedness":93,"clarity":92,"duplicationRisk":12,"blockingIssues":[],"notes":["“成本、数据处理要求”等属于使用方评估维度，来源未提供具体数据；当前表述为建议性关注点，并未冒充已证实事实。","可进一步明确“材料称”所指为Meta或原文，以区分来源陈述与编辑判断。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","editorial-labels","inference-boundary","low-source-overlap","no-html","independent-ai-review"]}},"tags":["行业动态","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Meta Superintelligence Labs 发布 Muse Voice Transcribe：一个实时模型同时完成流式 ASR、说话人分离和端点检测","summary":"Meta Superintelligence Labs 发布首个实时音频感知模型 Muse Voice Transcribe，将流式 ASR、20+ 说话人的说话人分离和端点检测合并为一个自回归模型，无需后处理。","category":"行业动态","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs 发布 Muse Voice Transcribe：一个实时模型同时完成流式 ASR、说话人分离和端点检测 - Aioga AI资讯","description":"Meta Superintelligence Labs 发布首个实时音频感知模型 Muse Voice Transcribe，将流式 ASR、20+ 说话人的说话人分离和端点检测合并为一个自回归模型，无需后处理。","url":"https://www.aioga.com/news/cmtjoufgu029rrotnspaovulr/","articleBody":["大多数生产语音堆栈是由三个系统拼接在一起的。一个模型进行转录，第二个模型区分说话者，检测器判断用户何时停止说话。每一次交接都会增加延迟并产生新的故障模式。","Muse Voice Transcribe：https://research.meta.ai/blog/introducing-muse-voice-transcribe，Meta Superintelligence Labs 本周宣布，将这三项工作合并为一个自回归模型。 Meta 称其为首个实时音频感知模型。它可以一次性执行流式 ASR、20 多个说话者的说话者分类和终结点，无需进行后处理。","它可以部署吗？可以，但仅限于托管API。它已经在Meta模型API上上线：https://developer.meta.com/ai/models/muse-voice-transcribe，型号为muse-voice-transcribe-1.0，价格为每1000分钟音频3美元（每小时0.18美元），并且已经为Mac上的Meta AI和Muse Code提供语音输入功能：https://developer.meta.com/ai/products/muse-code。模型权重尚未公开，因此没有自托管的途径。","Muse Voice Transcribe是Muse Spark系列中的一个自回归多模态模型。音频以12.5 Hz的频率以80ms的块到达。每个音频块被转换为一个软令牌。","每处理完一个音频块，模型会做出一个二元选择。它要么预测一个令牌并继续监听，要么输出一个文本令牌。当模型预测时，该令牌会被输入中的下一个实际音频块替换。当流结束时，会插入一个令牌，模型会在不请求更多音频的情况下刷新所有剩余文本。","听和写共享一个解码器循环，因此没有单独的对齐阶段出现漂移。","因为模型控制何时监听，它也控制每个单词后面有多少音频上下文。Meta称这一间隔为“延迟”。更长的延迟意味着转录更准确，但延迟更高。","Meta没有固定这种权衡，而是通过训练来调节。强化学习将词错误率奖励和延迟奖励相乘，生成一种策略，使每个单词的延迟根据难度而变化。Meta报告称，这使得模型在最终转录所需时间上，在速度与准确率的权衡上处于帕累托前沿，超过了由Soniox、Cartesia和ElevenLabs系统形成的先前前沿。","Meta 没有为说话人归属添加第二个模型，而是在同一流中添加了特殊标记。","对于分说话人标注，一个标记表示潜在的说话人切换，标签标识说话人。轮次标记在一旦可能发生切换时触发，而说话人标签延迟到该段落结束。一个说话人的音频可以被分割成多个片段，但都对应同一个标签。","对于端点检测，标记表示语音的开始，标记表示用户结束的点。两个任务与流式 ASR 联合训练，使用在 ASR 奖励基础上添加的额外奖励。","该模型在 70 多种语言上训练，其中 25 种在发布时经过广泛验证并推荐使用。代码切换是原生支持的，包括句子内和句子间，对于混合语言的双语说话人尤为重要。准确度可以通过语言、关键词和上下文偏置进一步提高。","长上下文处理是一个实际区分点。Meta 表示模型原生支持超过一小时的音频输入和 20 多个说话人，无需额外后处理步骤。","截至 2026 年 9 月 1 日，Meta 报告在流式语音转文本的人工分析和公开分说话人基准中获得第一名。","在人工分析 AA-WER 流式测评中：https://x.com/ArtificialAnlys/status/2094849283120128135，Muse Voice Transcribe 在语音结束 0.16 秒后最终抄本 WER 为 3.1%。使用语义端点的 Cartesia Ink-2 为 3.4%，延迟 0.43 秒。ElevenLabs Scribe v2 Realtime 为 3.6%，延迟 0.14 秒。使用外部端点的 Cartesia Ink-2 速度最快，仅需 0.07 秒，但准确度最低，为 4.0%。在首次部分抄本中，Muse Voice Transcribe 在 0.13 秒时记录 WER 为 3.6%。","在分说话人方面，Meta 报告在 AMI-IHM、AMI-SDM 和 VoxConverse 中的平均分说话人错误率为 17.5%。同一图表中的其他五个系统范围为 21.1% 至 28.6%。","价格是另一个维度。每 1,000 分钟 3.00 美元，比 Cartesia Ink-2 的 4.00 美元低，还不到 ElevenLabs Scribe v2 Realtime 和 Deepgram Flux 的 6.50 美元的一半。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等？请联系我们：https://forms.gle/wbash1wF6efRj8G58","米哈尔·萨特（Michal Sutter）是一名数据科学专业人士，拥有帕多瓦大学的数据科学硕士学位。凭借在统计分析、机器学习和数据工程方面的扎实基础，米哈尔擅长将复杂的数据集转化为可操作的洞察。"]},"en":{"title":"Meta Superintelligence Labs Releases Muse Voice Transcribe: A Real-Time Model That Performs Streaming ASR, Speaker Separation, and Endpoint Detection Simultaneously","summary":"Meta Superintelligence Labs released its first real-time audio-perception model, Muse Voice Transcribe, combining streaming ASR, speaker separation for more than 20 speakers, and endpoint detection into a single autoregressive model without the need for post-processing.","category":"Industry","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs Releases Muse Voice Transcribe: A Real-Time Model That Performs Streaming ASR, Speaker Separation, and Endpoint Detection Simultaneously - Aioga AI News","description":"Meta Superintelligence Labs released its first real-time audio-perception model, Muse Voice Transcribe, combining streaming ASR, speaker separation for more than 20 speakers, and e...","url":"https://www.aioga.com/en/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:06.158Z"},"ja":{"title":"Meta Superintelligence Labs、Muse Voice Transcribeを発表：リアルタイムモデルで同時にストリーミングASR、話者分離、エンドポイント検出を実現","summary":"Meta Superintelligence Labsは初のリアルタイム音声認識モデルMuse Voice Transcribeを発表し、ストリーミングASR、20人以上の話者の話者分離、エンドポイント検出を一つの自己回帰モデルに統合し、後処理を必要としない。","category":"業界動向","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs、Muse Voice Transcribeを発表：リアルタイムモデルで同時にストリーミングASR、話者分離、エンドポイント検出を実現 - Aioga AIニュース","description":"Meta Superintelligence Labsは初のリアルタイム音声認識モデルMuse Voice Transcribeを発表し、ストリーミングASR、20人以上の話者の話者分離、エンドポイント検出を一つの自己回帰モデルに統合し、後処理を必要としない。","url":"https://www.aioga.com/ja/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:07.295Z"},"ko":{"title":"Meta Superintelligence Labs, Muse Voice Transcribe 발표: 실시간 모델로 스트리밍 ASR, 화자 분리 및 엔드포인트 감지를 동시에 수행","summary":"Meta Superintelligence Labs는 최초의 실시간 오디오 인식 모델 Muse Voice Transcribe를 발표했으며, 스트리밍 ASR, 20명 이상 화자 분리, 엔드포인트 감지를 하나의 자동회귀 모델로 통합해 후처리 없이 수행할 수 있다.","category":"업계 동향","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs, Muse Voice Transcribe 발표: 실시간 모델로 스트리밍 ASR, 화자 분리 및 엔드포인트 감지를 동시에 수행 - Aioga AI 뉴스","description":"Meta Superintelligence Labs는 최초의 실시간 오디오 인식 모델 Muse Voice Transcribe를 발표했으며, 스트리밍 ASR, 20명 이상 화자 분리, 엔드포인트 감지를 하나의 자동회귀 모델로 통합해 후처리 없이 수행할 수 있다.","url":"https://www.aioga.com/ko/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:14.393Z"},"es":{"title":"Meta Superintelligence Labs lanza Muse Voice Transcribe: un modelo en tiempo real que realiza simultáneamente ASR en streaming, separación de hablantes y detección de puntos finales","summary":"Meta Superintelligence Labs lanzó su primer modelo de percepción de audio en tiempo real, Muse Voice Transcribe, que combina ASR en streaming, separación de hablantes de más de 20 voces y detección de puntos finales en un modelo autoregresivo, sin necesidad de posprocesamiento.","category":"Industria","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs lanza Muse Voice Transcribe: un modelo en tiempo real que realiza simultáneamente ASR en streaming, separación de hablantes y detección de puntos finales - Aioga Noticias de IA","description":"Meta Superintelligence Labs lanzó su primer modelo de percepción de audio en tiempo real, Muse Voice Transcribe, que combina ASR en streaming, separación de hablantes de más de 20...","url":"https://www.aioga.com/es/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:13.879Z"},"fr":{"title":"Meta Superintelligence Labs a publié Muse Voice Transcribe : un modèle temps réel qui effectue simultanément le streaming ASR, la séparation des haut-parleurs et la détection des points d’extrémité","summary":"Meta Superintelligence Labs a publié le premier modèle de perception audio en temps réel, Muse Voice Transcribe, qui combine le streaming ASR, la séparation de 20+ haut-parleurs et la détection des points d’extrémité en un seul modèle autorégressif sans besoin de post-traitement.","category":"Industrie","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs a publié Muse Voice Transcribe : un modèle temps réel qui effectue simultanément le streaming ASR, la séparation des haut-parleurs et la détection des points d’extrémité - Aioga Actualités IA","description":"Meta Superintelligence Labs a publié le premier modèle de perception audio en temps réel, Muse Voice Transcribe, qui combine le streaming ASR, la séparation de 20+ haut-parleurs et...","url":"https://www.aioga.com/fr/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:23.798Z"},"de":{"title":"Meta Superintelligence Labs veröffentlicht Muse Voice Transcribe: Ein Echtzeitmodell, das gleichzeitig Streaming-ASR, Sprechertrennung und Endpunkt-Erkennung durchführt","summary":"Meta Superintelligence Labs hat das erste Echtzeit-Audio-Wahrnehmungsmodell Muse Voice Transcribe veröffentlicht, das Streaming-ASR, Sprechertrennung für über 20 Sprecher und Endpunkt-Erkennung in einem autoregressiven Modell kombiniert, ohne Nachbearbeitung.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs veröffentlicht Muse Voice Transcribe: Ein Echtzeitmodell, das gleichzeitig Streaming-ASR, Sprechertrennung und Endpunkt-Erkennung durchführt - Aioga KI-News","description":"Meta Superintelligence Labs hat das erste Echtzeit-Audio-Wahrnehmungsmodell Muse Voice Transcribe veröffentlicht, das Streaming-ASR, Sprechertrennung für über 20 Sprecher und Endpu...","url":"https://www.aioga.com/de/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:21.532Z"},"pt-BR":{"title":"Meta Superintelligence Labs lança Muse Voice Transcribe: um modelo em tempo real que realiza ASR em streaming, separação de locutor e detecção de ponto final","summary":"Meta Superintelligence Labs lançou o primeiro modelo de percepção de áudio em tempo real Muse Voice Transcribe, que combina ASR em streaming, separação de locutores para mais de 20 falantes e detecção de ponto final em um único modelo autoregressivo, sem necessidade de pós-processamento.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs lança Muse Voice Transcribe: um modelo em tempo real que realiza ASR em streaming, separação de locutor e detecção de ponto final - Aioga Notícias de IA","description":"Meta Superintelligence Labs lançou o primeiro modelo de percepção de áudio em tempo real Muse Voice Transcribe, que combina ASR em streaming, separação de locutores para mais de 20...","url":"https://www.aioga.com/pt-BR/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:30.471Z"},"ru":{"title":"Meta Superintelligence Labs выпустила Muse Voice Transcribe: модель в реальном времени, которая одновременно выполняет потоковые ASR, разделение динамиков и обнаружение конечных точек","summary":"Meta Superintelligence Labs выпустила первую модель восприятия звука в реальном времени — Muse Voice Transcribe, которая объединяет потоковый ASR, разделение 20+ динамиков и обнаружение конечных точек в единую авторегрессивную модель без необходимости постобработки.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs выпустила Muse Voice Transcribe: модель в реальном времени, которая одновременно выполняет потоковые ASR, разделение динамиков и обнаружение конечных точек - Aioga Новости ИИ","description":"Meta Superintelligence Labs выпустила первую модель восприятия звука в реальном времени — Muse Voice Transcribe, которая объединяет потоковый ASR, разделение 20+ динамиков и обнару...","url":"https://www.aioga.com/ru/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:33.146Z"},"ar":{"title":"مختبرات Meta Superintelligence تصدر Muse Voice Transcribe: نموذج في الوقت الفعلي ينفذ التعرف الصوتي، فصل المتحدثين وكشف النقاط النهائية في وقت واحد","summary":"أصدرت مختبرات Meta Superintelligence أول نموذج إدراك صوتي في الوقت الفعلي Muse Voice Transcribe، الذي يجمع بين التعرف الصوتي المستمر، فصل أكثر من 20 متحدثًا وكشف النقاط النهائية في نموذج رجعي ذاتي واحد دون الحاجة لمعالجة لاحقة.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"مختبرات Meta Superintelligence تصدر Muse Voice Transcribe: نموذج في الوقت الفعلي ينفذ التعرف الصوتي، فصل المتحدثين وكشف النقاط النهائية في وقت واحد - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت مختبرات Meta Superintelligence أول نموذج إدراك صوتي في الوقت الفعلي Muse Voice Transcribe، الذي يجمع بين التعرف الصوتي المستمر، فصل أكثر من 20 متحدثًا وكشف النقاط النهائية في...","url":"https://www.aioga.com/ar/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:39.707Z"},"hi":{"title":"Meta Superintelligence Labs ने Muse Voice Transcribe जारी किया: एक रीयल-टाइम मॉडल जो एक साथ स्ट्रीमिंग ASR, स्पीकर सेपरेशन और एंडपॉइंट डिटेक्शन करता है","summary":"मेटा सुपरइंटेलिजेंस लैब्स ने पहला रीयल-टाइम ऑडियो परसेप्शन मॉडल, म्यूज़ वॉयस ट्रांसक्राइब जारी किया, जो स्ट्रीमिंग ASR, 20+ स्पीकर सेपरेशन और एंडपॉइंट डिटेक्शन को एक एकल ऑटोरेग्रेसिव मॉडल में जोड़ता है, जिसमें बिना किसी पोस्ट-प्रोसेसिंग की आवश्यकता होती है।","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs ने Muse Voice Transcribe जारी किया: एक रीयल-टाइम मॉडल जो एक साथ स्ट्रीमिंग ASR, स्पीकर सेपरेशन और एंडपॉइंट डिटेक्शन करता है - Aioga AI समाचार","description":"मेटा सुपरइंटेलिजेंस लैब्स ने पहला रीयल-टाइम ऑडियो परसेप्शन मॉडल, म्यूज़ वॉयस ट्रांसक्राइब जारी किया, जो स्ट्रीमिंग ASR, 20+ स्पीकर सेपरेशन और एंडपॉइंट डिटेक्शन को एक एकल ऑटोरेग्रेस...","url":"https://www.aioga.com/hi/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:42.326Z"},"it":{"title":"Meta Superintelligence Labs lancia Muse Voice Transcribe: un modello in tempo reale che realizza simultaneamente ASR a flusso continuo, separazione dei parlanti e rilevamento dei punti finali","summary":"Meta Superintelligence Labs ha rilasciato il primo modello di percezione audio in tempo reale Muse Voice Transcribe, che unisce ASR a flusso continuo, separazione dei parlanti per oltre 20 parlanti e rilevamento dei punti finali in un modello autoregressivo, senza necessità di post-elaborazione.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs lancia Muse Voice Transcribe: un modello in tempo reale che realizza simultaneamente ASR a flusso continuo, separazione dei parlanti e rilevamento dei punti finali - Aioga Notizie IA","description":"Meta Superintelligence Labs ha rilasciato il primo modello di percezione audio in tempo reale Muse Voice Transcribe, che unisce ASR a flusso continuo, separazione dei parlanti per...","url":"https://www.aioga.com/it/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:48.431Z"},"nl":{"title":"Meta Superintelligence Labs lanceert Muse Voice Transcribe: een realtime model dat tegelijk gestreamde ASR, sprekersscheiding en eindpuntsdetectie uitvoert","summary":"Meta Superintelligence Labs heeft het eerste realtime audio-perceptiemodel Muse Voice Transcribe uitgebracht, dat gestreamde ASR, sprekersscheiding voor 20+ sprekers en eindpuntsdetectie combineert in één autoregressief model, zonder nabewerking.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs lanceert Muse Voice Transcribe: een realtime model dat tegelijk gestreamde ASR, sprekersscheiding en eindpuntsdetectie uitvoert - Aioga AI-nieuws","description":"Meta Superintelligence Labs heeft het eerste realtime audio-perceptiemodel Muse Voice Transcribe uitgebracht, dat gestreamde ASR, sprekersscheiding voor 20+ sprekers en eindpuntsde...","url":"https://www.aioga.com/nl/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:49.842Z"},"tr":{"title":"Meta Superintelligence Labs, Muse Voice Transcribe’ı Yayınladı: Akış ASR, Konuşmacı Ayırma ve Uç Tespitini Aynı Anda Gerçekleştiren Bir Gerçek Zamanlı Model","summary":"Meta Superintelligence Labs, ilk gerçek zamanlı ses algılama modeli Muse Voice Transcribe’ı yayınladı; bu model, akış ASR, 20+ konuşmacının ayrımını ve uç tespitini tek bir oto-regresif modele entegre ediyor ve son işlem gerektirmiyor.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs, Muse Voice Transcribe’ı Yayınladı: Akış ASR, Konuşmacı Ayırma ve Uç Tespitini Aynı Anda Gerçekleştiren Bir Gerçek Zamanlı Model - Aioga AI Haberleri","description":"Meta Superintelligence Labs, ilk gerçek zamanlı ses algılama modeli Muse Voice Transcribe’ı yayınladı; bu model, akış ASR, 20+ konuşmacının ayrımını ve uç tespitini tek bir oto-reg...","url":"https://www.aioga.com/tr/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:57.312Z"},"vi":{"title":"Meta Superintelligence Labs đã phát hành Muse Voice Transcribe: một mô hình thời gian thực thực hiện đồng thời ASR phát trực tiếp, tách loa và phát hiện điểm cuối","summary":"Meta Superintelligence Labs đã phát hành mô hình nhận thức âm thanh thời gian thực đầu tiên, Muse Voice Transcribe, kết hợp ASR streaming, cách ly 20+ loa và phát hiện điểm cuối thành một mô hình tự hồi quy duy nhất mà không cần xử lý hậu kỳ.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs đã phát hành Muse Voice Transcribe: một mô hình thời gian thực thực hiện đồng thời ASR phát trực tiếp, tách loa và phát hiện điểm cuối - Tin tức AI Aioga","description":"Meta Superintelligence Labs đã phát hành mô hình nhận thức âm thanh thời gian thực đầu tiên, Muse Voice Transcribe, kết hợp ASR streaming, cách ly 20+ loa và phát hiện điểm cuối th...","url":"https://www.aioga.com/vi/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:21:59.432Z"},"id":{"title":"Meta Superintelligence Labs Merilis Muse Voice Transcribe: Model Real-Time yang Menyelesaikan ASR Streaming, Pemisahan Pembicara, dan Deteksi Endpoint","summary":"Meta Superintelligence Labs merilis model persepsi audio real-time pertama mereka, Muse Voice Transcribe, yang menggabungkan ASR streaming, pemisahan pembicara dari 20+ pembicara, dan deteksi endpoint menjadi satu model autoregresif tanpa perlu pasca-pemrosesan.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs Merilis Muse Voice Transcribe: Model Real-Time yang Menyelesaikan ASR Streaming, Pemisahan Pembicara, dan Deteksi Endpoint - Berita AI Aioga","description":"Meta Superintelligence Labs merilis model persepsi audio real-time pertama mereka, Muse Voice Transcribe, yang menggabungkan ASR streaming, pemisahan pembicara dari 20+ pembicara,...","url":"https://www.aioga.com/id/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:22:05.307Z"},"th":{"title":"Meta Superintelligence Labs ได้เปิดตัว Muse Voice Transcribe: โมเดลเรียลไทม์ที่ทํางานพร้อมกันระหว่างสตรีม ASR, แยกลําโพง และตรวจจับปลายทาง","summary":"Meta Superintelligence Labs ได้เปิดตัวโมเดลการรับรู้เสียงแบบเรียลไทม์ตัวแรก Muse Voice Transcribe ซึ่งรวมการสตรีม ASR การแยกลําโพง 20+ ตัว และการตรวจจับปลายทางเข้าด้วยกันเป็นโมเดลออโตเรเกรสซีฟเดียวโดยไม่ต้องประมวลผลหลัง","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs ได้เปิดตัว Muse Voice Transcribe: โมเดลเรียลไทม์ที่ทํางานพร้อมกันระหว่างสตรีม ASR, แยกลําโพง และตรวจจับปลายทาง - ข่าว AI Aioga","description":"Meta Superintelligence Labs ได้เปิดตัวโมเดลการรับรู้เสียงแบบเรียลไทม์ตัวแรก Muse Voice Transcribe ซึ่งรวมการสตรีม ASR การแยกลําโพง 20+ ตัว และการตรวจจับปลายทางเข้าด้วยกันเป็นโมเดลอ...","url":"https://www.aioga.com/th/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:22:08.784Z"},"pl":{"title":"Meta Superintelligence Labs wprowadza Muse Voice Transcribe: model działający w czasie rzeczywistym, wykonujący jednocześnie strumieniowy ASR, separację mówców i detekcję końca wypowiedzi","summary":"Meta Superintelligence Labs ogłosiła pierwszy model audio działający w czasie rzeczywistym Muse Voice Transcribe, łączący strumieniowy ASR, separację mówców dla ponad 20 osób i detekcję końca wypowiedzi w jeden model autoregresyjny, bez potrzeby dalszej obróbki.","category":"行业动态","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Meta Superintelligence Labs wprowadza Muse Voice Transcribe: model działający w czasie rzeczywistym, wykonujący jednocześnie strumieniowy ASR, separację mówców i detekcję końca wypowiedzi - Aioga Wiadomości AI","description":"Meta Superintelligence Labs ogłosiła pierwszy model audio działający w czasie rzeczywistym Muse Voice Transcribe, łączący strumieniowy ASR, separację mówców dla ponad 20 osób i det...","url":"https://www.aioga.com/pl/news/cmtjoufgu029rrotnspaovulr/","contentTranslated":true,"sourceHash":"a9c6a94d86e80ed0","translatedAt":"2026-09-02T06:22:16.035Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":""}}