{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T06:20:51.496Z","headline":"2026年最佳开源语音识别模型对比：WER、语言、延迟与许可证","description":"开源语音识别领域已不再是Whisper一家独大。Cohere Transcribe（2B，Apache 2.0）以5.42%平均词错误率登顶Hugging Face开源ASR排行榜，但IBM Granite Speech 4.1 2B（5.33%）和ARK-ASR-3B（5.04%）紧随其后，榜首差距不足1个WER点。","url":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/","mainEntityOfPage":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/","datePublished":"2026-07-23T09:26:40.000Z","dateModified":"2026-07-23T09:26:40.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared","https://aihot.virxact.com/items/cmrxc6h0d01a6rot3gv44kv9z"],"canonicalUrl":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：开源语音识别领域已不再是Whisper一家独大。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/","dateCreated":"2026-07-23T09:26:40.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared","datePublished":"2026-07-23T09:26:40.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrxc6h0d01a6rot3gv44kv9z","datePublished":"2026-07-23T09:26:40.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrxc6h0d01a6rot3gv44kv9z"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared"},"article":{"id":"cmrxc6h0d01a6rot3gv44kv9z","slug":"cmrxc6h0d01a6rot3gv44kv9z","url":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/","title":"2026年最佳开源语音识别模型对比：WER、语言、延迟与许可证","title_en":"Best Open Speech Recognition （ASR） Models in 2026： WER， Languages， Latency， and License Compared","summary":"开源语音识别领域已不再是Whisper一家独大。Cohere Transcribe（2B，Apache 2.0）以5.42%平均词错误率登顶Hugging Face开源ASR排行榜，但IBM Granite Speech 4.1 2B（5.33%）和ARK-ASR-3B（5.04%）紧随其后，榜首差距不足1个WER点。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared","aiHotUrl":"https://aihot.virxact.com/items/cmrxc6h0d01a6rot3gv44kv9z","publishedAt":"2026-07-23T09:26:40.000Z","category":"技巧观点","score":60,"selected":false,"articleBody":["Open speech recognition stopped being a Whisper monoculture some time in the last twelve months. In March 2026 Cohere released Transcribe：https://cohere.com/blog/transcribe, a 2B Apache 2.0 model that took the top of the Hugging Face Open ASR Leaderboard：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard at 5.42% average word error rate. Five weeks later IBM shipped Granite Speech 4.1 2B：https://huggingface.co/ibm-granite/granite-speech-4.1-2b at 5.33%. Since then ARK-ASR-3B：https://huggingface.co/AutoArk-AI/ARK-ASR-3B and MOSS-Transcribe-preview-2B：https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-preview-2B have posted lower numbers still.","The top of that leaderboard is now separated by less than one WER point. That has a specific consequence for anyone choosing a model: rank is no longer the deciding variable. License, language coverage, streaming support, and cost per audio-hour are. This roundup compares the field on all four.","The Open ASR Leaderboard average is not a single fixed quantity, and the models currently listed side by side were not all scored the same way.","Cohere’s 5.42% is an average across eight English test sets, including TED-LIUM. That checks out: AMI 8.13, Earnings-22 10.86, GigaSpeech 9.34, LibriSpeech clean 1.25, LibriSpeech other 2.37, SPGISpeech 3.08, TED-LIUM 2.49, VoxPopuli 5.87 averages to exactly 5.42.","ARK-ASR-3B’s 5.04% is an average across seven sets. TED-LIUM is absent. The MOSS-Transcribe-preview-2B card states this explicitly: TED-LIUM is not currently part of the leaderboard run and is therefore excluded.","TED-LIUM is one of the easier sets in the suite, so dropping it raises the average. Recompute Cohere’s published per-dataset scores over the same seven sets ARK reports and Cohere lands at 5.84, not 5.42. Do the same to Granite Speech 4.1 2B and it moves from 5.33 to 5.65. On a like-for-like basis ARK’s lead is larger than the headline numbers imply, not smaller — but the point is that you cannot subtract one published figure from another and get a meaningful answer.","Two further caveats belong on the same page:","Some scores are openly leaderboard-fitted : The MOSS-Transcribe-preview-2B card states the model was fine-tuned with reinforcement learning on the Open ASR Leaderboard training splits. That is disclosed, which is more than most, but it means the score measures the benchmark rather than the capability.","Private-track data reorders the board : Appen contributed held-back evaluation sets：https://www.appen.com/blog/hugging-face-open-llm-leaderboard covering Australian, Canadian, Indian, and American accents in scripted and conversational conditions. When those private sets are toggled on, zoom/scribe_v1 moves from #4 to #1 and the public-leaderboard leader drops a position. Models tuned for clean read speech degrade disproportionately on spontaneous conversational audio.","Use the leaderboard to build a shortlist. Do not use it to pick a winner.","Cohere Transcribe：https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 (2B, Apache 2.0, 14 languages) is the model that actually shipped into production. It has been downloaded over 620,000 times in the past month and has runtime support across transformers , vLLM, mlx-audio for Apple Silicon, a Rust port, and a WebGPU build. It is a Conformer encoder with a lightweight Transformer decoder, trained from scratch. Cohere also ran human preference evaluation, where trained annotators scored transcripts for meaning preservation, hallucination, and named entities: a 61% average win rate, 78% against IBM Granite 4.0 1B Speech and 64% against Whisper large-v3.","The limitations section of its model card is unusually honest and should be read before committing. There is no automatic language detection, no timestamps, and no diarization, and the model is eager to transcribe silence, so Cohere recommends prepending a VAD or noise gate. The repo is also gated behind a contact-information agreement despite the Apache 2.0 license.","Granite Speech 4.1 2B：https://huggingface.co/ibm-granite/granite-speech-4.1-2b (2B, Apache 2.0) is the better pick if you need capability rather than a lower number. Six languages for ASR plus bidirectional speech translation, keyword-list biasing for names and jargon, punctuation and truecasing including German noun capitalization. Trained on 174,000 hours. RTFx 231.29. IBM also ships two siblings: -plus adds speaker-attributed ASR and word-level timestamps, and -nar is discussed below.","Canary-Qwen-2.5B：https://huggingface.co/nvidia/canary-qwen-2.5b (2.5B, CC-BY-4.0, English) pairs a FastConformer encoder with a Qwen3-1.7B decoder and runs in two modes — pure transcription, or LLM mode where the decoder summarizes and answers questions about the transcript. 5.63% WER at RTFx 418. Note that AMI was oversampled to roughly 15% of training data, which biases output toward verbatim disfluency-preserving transcripts. That is a feature for legal work and a nuisance for meeting notes.","Qwen3-ASR-1.7B：https://huggingface.co/Qwen/Qwen3-ASR-1.7B (Apache 2.0) covers 52 languages and dialects — 30 languages plus 22 Chinese dialects — at 5.76%. It ships with a full inference toolkit and a separate forced-alignment model for timestamps in 11 languages. For anything touching Mandarin or Chinese regional speech this is the obvious starting point.","Accuracy across the top of the field now varies by about one WER point. Throughput varies by more than an order of magnitude, which means throughput usually decides the invoice.","Parakeet TDT 0.6B v3：https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 (0.6B, CC-BY-4.0) is the throughput leader among multilingual open models at RTFx 3332.74 across 25 European languages with automatic language ID, handling up to 24 minutes in a single pass on an A100 80GB. It costs 6.32% WER — roughly one point more than Granite 4.1 2B for roughly fourteen times the audio per GPU-second.","Granite Speech 4.1 2B-NAR：https://huggingface.co/ibm-granite/granite-speech-4.1-2b-nar is the more interesting engineering result. It is non-autoregressive: it edits a CTC hypothesis in a single forward pass using a bidirectional LLM, reaching RTFx ~1820 on one H100 at batch size 128. It gives up Japanese, speech translation, and keyword biasing to get there.","Qwen3-ASR-0.6B：https://huggingface.co/Qwen/Qwen3-ASR-0.6B keeps all 52 languages and reaches 2000× throughput at concurrency 128.","A batch WER seems to be the wrong test for a streaming model, and the leaderboard scores them anyway. Voxtral Realtime sits at 7.68% and Kyutai STT 2.6B at 6.40% — both below Whisper — and neither number tells you anything useful about their intended use.","Voxtral Mini 4B Realtime 2602：https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602 (Apache 2.0, 13 languages) is a 3.4B language model along with a 970M causal audio encoder trained from scratch, with sliding-window attention on both halves for effectively unbounded streaming. Transcription delay is configurable in 80ms steps from 80ms to 1200ms, plus a standalone 2400ms option; Mistral recommends 480ms as the sweet spot and reports that at that setting it matches leading offline open models. It runs on a single 16GB GPU and had day-0 vLLM Realtime API support：https://developers.redhat.com/articles/2026/02/06/run-voxtral-mini-4b-realtime-vllm-red-hat-ai.","Kyutai STT：https://huggingface.co/kyutai/stt-1b-en_fr (CC-BY-4.0) comes in two shapes: a ~1B English/French model with a 0.5s delay and a built-in semantic voice activity detector, and a 2.6B English-only model with a 2.5s delay. For voice agents, the semantic VAD matters more than the transcription delay — it predicts when the speaker has actually finished, which is what governs perceived turn-taking latency. An H100 serves 400 concurrent streams in real time.","Meta’s Omnilingual ASR：https://ai.meta.com/research/publications/omnilingual-asr-open-source-multilingual-speech-recognition-for-1600-languages/ (Apache 2.0, corpus CC-BY) is not competing on WER and should not be evaluated as if it were. It covers 1,600+ languages natively and extends to 5,400+ through zero-shot in-context learning, built on a wav2vec 2.0 encoder scaled to 7B and pre-trained on about 4.3M hours. The 7B LLM-ASR variant achieves character error rate below 10% on 78% of supported languages, including 500+ never previously served by any ASR system. Encoder sizes run from 300M to 7B. Meta also released the Omnilingual ASR Corpus covering 350+ underserved languages.","Whisper large-v3：https://huggingface.co/openai/whisper-large-v3 (1.55B, MIT, 99 languages) has been overtaken on accuracy by roughly ten open models, and it remains the correct default for a large class of projects. MIT is the least encumbered license in the field. The runtime ecosystem — whisper.cpp, faster-whisper, WhisperX — has no equivalent among the newer releases. If your requirement is “some language, some hardware, no license lawyer,” it is still the answer.","diffusion-gemma-asr-small：https://github.com/InterfazeAI/diffusion-gemma-asr from YC startup Interfaze is the most architecturally creative release of the year. It produces transcripts by parallel diffusion denoising over a 256-token canvas in 8 to 16 steps, so decoding cost does not grow with transcript length. Only ~42M parameters were trained — 0.16% of the weights — on top of a frozen 26B DiffusionGemma and a frozen whisper-small encoder. It reaches 6.6% WER on LibriSpeech test-clean at roughly 11 to 17× realtime. It also reaches 15.7% on FLEURS English and 29.6% CER on FLEURS Mandarin, so treat the LibriSpeech figure as the ceiling. Nobody should deploy this. Everybody working on ASR should read it.","MOSS-Transcribe-Diarize 0.9B：https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize (Apache 2.0, 50+ languages) solves the problem most roundups ignore: it emits speaker labels, word timestamps, and transcript in one generation instead of chaining ASR to a separate diarization stack. 128k context, roughly 90 minutes of audio in one pass, RTF ~0.017 on an RTX 4090, with hotword biasing.","This is the section that actually blocks shipping, and the field divides cleanly:","Apache 2.0 : Cohere Transcribe, Granite Speech 4.1 (all three variants), Qwen3-ASR (both sizes), Voxtral Mini Realtime, Omnilingual ASR, ARK-ASR, MOSS-Transcribe. No attribution obligation, commercial use unrestricted. Note that Cohere’s repo is gated behind a contact-information agreement even though the license itself is Apache 2.0.","MIT : Whisper large-v3. The most permissive option in the field.","CC-BY-4.0 : Canary-Qwen-2.5B, Parakeet TDT 0.6B v3, Kyutai STT. Commercially usable, but attribution is required. For an embedded product or a white-labeled API, that is a real compliance obligation, and it is the single most common reason teams end up shipping a model that is not the most accurate one they tested.","Meta’s Omnilingual ASR splits the two: Apache 2.0 for models, CC-BY for the corpus.","Run this order, not just the leaderboard order:","The useful summary of 2026 is not that any single model won. It is that a 2B open-weight model on a permissive license now beats what closed APIs were charging for eighteen months ago, and that the remaining decision is a procurement question rather than a research one.","Leaderboard positions are live and change frequently. All figures verified against primary sources on 23 July 2026.","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":33,"url":"/media/articles/cmrxc6h0d01a6rot3gv44kv9z/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog6171-1-100x70.png","alt":"Designing High-Performance GPU Kernels with TileLang: Tensor-Core GEMM, Fused Softmax, FlashAttention, and Autotuning","afterParagraph":34,"url":"/media/articles/cmrxc6h0d01a6rot3gv44kv9z/80925b48ffa65ad2.webp"}],"mediaStatus":"ok","articleBodyZh":["开放语音识别在过去十二个月中的某个时间，停止了仅由 Whisper 主导的状态。2026 年 3 月，Cohere 发布了 Transcribe：https://cohere.com/blog/transcribe，这是一款 2B 参数的 Apache 2.0 模型，以 5.42% 的平均词错误率登上了 Hugging Face 开放 ASR 排行榜的榜首：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard。五周后，IBM 推出了 Granite Speech 4.1 2B：https://huggingface.co/ibm-granite/granite-speech-4.1-2b，平均词错误率为 5.33%。从那时起，ARK-ASR-3B：https://huggingface.co/AutoArk-AI/ARK-ASR-3B 和 MOSS-Transcribe-preview-2B：https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-preview-2B 的数值更低。","现在该排行榜的榜首相差不到一个词错误率点。这对任何选择模型的人都有一个具体影响：排名不再是决定性因素。许可证、语言覆盖范围、流式支持和每小时音频成本才是。这篇综述将对四个方面进行比较。","开放 ASR 排行榜的平均值并不是一个固定数值，目前并列列出的模型并非都以相同方式评分。","Cohere 的 5.42% 是八个英语测试集的平均值，包括 TED-LIUM。这是正确的：AMI 8.13，Earnings-22 10.86，GigaSpeech 9.34，LibriSpeech clean 1.25，LibriSpeech other 2.37，SPGISpeech 3.08，TED-LIUM 2.49，VoxPopuli 5.87，平均值恰好为 5.42。","ARK-ASR-3B 的 5.04% 是七个测试集的平均值。TED-LIUM 未包含在内。MOSS-Transcribe-preview-2B 的说明卡明确指出：TED-LIUM 当前未被纳入排行榜运行，因此被排除在外。","TED-LIUM 是测试套件中较容易的一个集，因此去掉它会提高平均值。用 ARK 报告的同样七个测试集重新计算 Cohere 发布的每个数据集的成绩，Cohere 的平均值为 5.84，而不是 5.42。对 Granite Speech 4.1 2B 进行同样操作，它从 5.33 上升到 5.65。在同类比较基础上，ARK 的领先幅度比标题数字显示的更大，而不是更小——但关键是，你不能用一个公布数字减去另一个就得到有意义的结果。","还有两个进一步注意事项应列在同一页上：","一些分数明显是为了排行榜而调整的：MOSS-Transcribe-preview-2B 的模型卡说明该模型在 Open ASR 排行榜的训练数据上通过强化学习进行了微调。这个信息是公开的，比大多数模型都透明，但这意味着分数衡量的是基准测试而不是实际能力。","私有轨道数据会重新排列排行榜：Appen 提供了保留的评估集：https://www.appen.com/blog/hugging-face-open-llm-leaderboard，涵盖澳大利亚、加拿大、印度和美国口音的脚本文本和对话场景。当这些私有数据集被启用时，zoom/scribe_v1 从第 4 名升至第 1 名，而公共排行榜的领先模型下降一位。针对干净朗读语音微调的模型在自发性对话音频上的表现降幅较大。","使用排行榜来构建候选名单，不要用它来决定赢家。","Cohere Transcribe：https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 (2B，Apache 2.0，14 种语言) 是实际投入生产的模型。过去一个月内下载量超过 620,000 次，并且在 transformers、vLLM、mlx-audio (Apple Silicon)、Rust 移植版本以及 WebGPU 构建方面均有运行支持。它是一个 Conformer 编码器，配有轻量级 Transformer 解码器，从零开始训练。Cohere 还进行了人工偏好评估，由受过培训的标注者根据意义保留、幻觉现象和命名实体对转录结果打分：平均胜率 61%，对 IBM Granite 4.0 1B Speech 胜率 78%，对 Whisper large-v3 胜率 64%。","模型卡中对局限性的描述异常诚实，应在使用前阅读。模型不具备自动语言检测、时间戳和说话人分离功能，并且倾向于转录静音，因此 Cohere 建议在前置加一个 VAD 或噪声门。尽管采用 Apache 2.0 许可，代码仓库仍需提供联系信息才能访问。","花岗岩语音 4.1 2B：https://huggingface.co/ibm-granite/granite-speech-4.1-2b (2B, Apache 2.0) 是更好的选择，如果你需要能力而不仅仅是较低的参数量。支持六种语言的自动语音识别（ASR）以及双向语音翻译，支持名称和术语的关键词列表偏置，标点符号和大小写恢复，包括德语名词大写。训练数据量为 174,000 小时。RTFx 231.29。IBM 还提供两个兄弟版本：-plus 增加了说话人归属的 ASR 和词级时间戳，而 -nar 在下面讨论。","Canary-Qwen-2.5B：https://huggingface.co/nvidia/canary-qwen-2.5b (2.5B, CC-BY-4.0, 英语) 结合了 FastConformer 编码器和 Qwen3-1.7B 解码器，并可运行两种模式——纯转录模式，或者 LLM 模式，在此模式下解码器对转录内容进行总结并回答问题。WER 为 5.63%，RTFx 418。注意 AMI 数据被过采样约占训练数据的 15%，这会使输出偏向保留口误的逐字转录。这对于法律工作是一种功能，但对于会议记录来说是一种困扰。","Qwen3-ASR-1.7B：https://huggingface.co/Qwen/Qwen3-ASR-1.7B (Apache 2.0) 覆盖 52 种语言和方言——30 种语言加上 22 种中文方言——WER 为 5.76%。它提供完整的推理工具包，并提供一个单独的强制对齐模型，用于 11 种语言的时间戳。对于涉及普通话或中国地方语音的任何任务，这显然是起点。","当前顶级模型的准确率差约一 WER 点。吞吐量差异超过一个数量级，这意味着通常是吞吐量决定了发票金额。","Parakeet TDT 0.6B v3：https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 (0.6B, CC-BY-4.0) 是多语言开源模型中吞吐量的领先者，RTFx 为 3332.74，覆盖 25 种欧洲语言并支持自动语言识别，在单次处理时可在 A100 80GB 上处理长达 24 分钟的音频。其 WER 为 6.32%——比 Granite 4.1 2B 高约一百分点，但每个 GPU-秒可处理的音频量大约是其十四倍。","花岗岩语音 4.1 2B-NAR：https://huggingface.co/ibm-granite/granite-speech-4.1-2b-nar 是更有趣的工程成果。它是非自回归的：使用双向 LLM 在一次前向传播中编辑 CTC 假设，在单个 H100 上批量大小 128 时 RTFx 约为 1820。为了达到这一点，它舍弃了日语、语音翻译和关键词偏置功能。","Qwen3-ASR-0.6B：https://huggingface.co/Qwen/Qwen3-ASR-0.6B 保留所有52种语言，并在并发128时达到2000×吞吐量。","批量 WER 似乎是对流式模型来说不合适的测试，但排行榜仍然对它们进行评分。Voxtral Realtime 的 WER 为 7.68%，Kyutai STT 2.6B 为 6.40%——两者都低于 Whisper——且这些数字对其预期用途没有任何实际参考价值。","Voxtral Mini 4B Realtime 2602：https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602（Apache 2.0，13 种语言）是一个 3.4B 的语言模型，配有一个从零开始训练的 970M 因果音频编码器，对两半采用滑动窗口注意力，实现有效的无限流式处理。转录延迟可以以 80ms 为步长从 80ms 配置到 1200ms，并提供单独的 2400ms 选项；Mistral 推荐 480ms 作为最佳设置，并报告在该设置下与领先的离线开放模型相匹配。它可在单个 16GB GPU 上运行，并在第0天就支持 vLLM Realtime API：https://developers.redhat.com/articles/2026/02/06/run-voxtral-mini-4b-realtime-vllm-red-hat-ai。","Kyutai STT：https://huggingface.co/kyutai/stt-1b-en_fr（CC-BY-4.0）提供两种版本：一个约1B 参数的英语/法语模型，延迟0.5秒，内置语义语音活动检测器；另一个为 2.6B 英语专用模型，延迟 2.5 秒。对于语音代理来说，语义 VAD 比转录延迟更重要——它预测说话者何时真正说完，这决定了感知的轮次切换延迟。H100 可实时服务 400 个并发流。","Meta 的 Omnilingual ASR：https://ai.meta.com/research/publications/omnilingual-asr-open-source-multilingual-speech-recognition-for-1600-languages/（Apache 2.0，语料 CC-BY）并不是在 WER 上竞争，不应按此进行评估。它原生支持 1600+ 种语言，并通过零样本上下文学习扩展至 5400+ 种语言，基于扩展到 7B 的 wav2vec 2.0 编码器，并在约 430 万小时数据上预训练。7B LLM-ASR 版本在 78% 的支持语言上字符错误率低于 10%，包括 500+ 在任何 ASR 系统中未曾服务过的语言。编码器规模从 300M 到 7B 不等。Meta 还发布了覆盖 350+ 个服务不足语言的 Omnilingual ASR 语料库。","Whisper large-v3：https://huggingface.co/openai/whisper-large-v3 （1.55B，MIT，支持99种语言）在准确性上已被大约十个开放模型超越，但它仍然是大类项目的默认正确选择。MIT 是该领域最不受限制的许可证。运行时生态系统——whisper.cpp、faster-whisper、WhisperX——在新发布的模型中没有可比的替代。如果你的需求是“某种语言、某种硬件、不想找许可证律师”，它仍然是答案。","diffusion-gemma-asr-small：https://github.com/InterfazeAI/diffusion-gemma-asr 来自 YC 初创公司 Interfaze，是今年架构上最具创意的发布。它通过在 256 令牌画布上进行 8 到 16 步的并行扩散去噪来生成文本，因此解码成本不会随着文本长度增加而增长。仅训练了约 4200 万个参数——占权重的 0.16%——是在冻结的 26B DiffusionGemma 和冻结的 whisper-small 编码器之上训练的。在 LibriSpeech test-clean 上达到约实时 11 到 17 倍速度时 WER 为 6.6%。在 FLEURS 英语上达到 15.7%，在 FLEURS 中文上 CER 为 29.6%，所以应将 LibriSpeech 的数据视为上限。没人应该部署它，但所有从事 ASR 的人都应阅读它。","MOSS-Transcribe-Diarize 0.9B：https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize （Apache 2.0，支持50+种语言）解决了大多数综述忽略的问题：它在一次生成中输出说话者标签、词语时间戳和文本，而不是将 ASR 与独立的说话人分离堆栈链式连接。128k 上下文，一次可处理约 90 分钟音频，RTX 4090 上 RTF 约为 0.017，并支持热词偏置。","这是实际阻碍发布的部分，领域内意见分歧明显：","Apache 2.0：Cohere Transcribe，Granite Speech 4.1（所有三种变体），Qwen3-ASR（两种尺寸），Voxtral Mini Realtime，Omnilingual ASR，ARK-ASR，MOSS-Transcribe。无需署名义务，商业用途无限制。注意 Cohere 的仓库虽然许可证是 Apache 2.0，但需要通过联系方式认证。","MIT：Whisper large-v3。该领域最宽松的选择。","CC-BY-4.0：Canary-Qwen-2.5B, Parakeet TDT 0.6B v3, Kyutai STT。可商业使用，但需要署名。对于嵌入式产品或白标 API，这是一个真正的合规义务，也是团队最终发布并非他们测试过的最准确模型的最常见原因。","Meta 的 Omnilingual ASR 将两者分开：模型使用 Apache 2.0，语料库使用 CC-BY 许可。","按这个顺序运行，而不仅仅是按照排行榜顺序：","2026 年的有用总结不是某个单一模型获胜，而是一个许可宽松的 20 亿参数开放权重模型现在击败了闭源 API 十八个月前的收费水平，而剩下的决策更多是采购问题而非研究问题。","排行榜位置实时更新，变动频繁。所有数据已于 2026 年 7 月 23 日核对主要来源。","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为具有远见的企业家和工程师，Asif 致力于利用人工智能的潜力为社会公益服务。他最近的努力是推出人工智能媒体平台 Marktechpost，该平台因其对机器学习和深度学习新闻的深入报道而脱颖而出，报道内容既技术可靠，又易于广大受众理解。该平台每月浏览量超过 200 万，显示出其广受观众欢迎的程度。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：开源语音识别领域已不再是Whisper一家独大。 Aioga 将其归入「技巧观点」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-28T06:29:10.600Z","sourceHash":"45837c0722f87474","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["技巧观点","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"2026年最佳开源语音识别模型对比：WER、语言、延迟与许可证","summary":"开源语音识别领域已不再是Whisper一家独大。Cohere Transcribe（2B，Apache 2.0）以5.42%平均词错误率登顶Hugging Face开源ASR排行榜，但IBM Granite Speech 4.1 2B（5.33%）和ARK-ASR-3B（5.04%）紧随其后，榜首差距不足1个WER点。","category":"技巧观点","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"2026年最佳开源语音识别模型对比：WER、语言、延迟与许可证 - Aioga AI资讯","description":"开源语音识别领域已不再是Whisper一家独大。Cohere Transcribe（2B，Apache 2.0）以5.42%平均词错误率登顶Hugging Face开源ASR排行榜，但IBM Granite Speech 4.1 2B（5.33%）和ARK-ASR-3B（5.04%）紧随其后，榜首差距不足1个WER点。","url":"https://www.aioga.com/news/cmrxc6h0d01a6rot3gv44kv9z/"},"en":{"title":"Comparison of the Best Open Source Speech Recognition Models in 2026: WER, Language, Latency, and License","summary":"The field of open-source speech recognition is no longer dominated solely by Whisper. Cohere Transcribe (2B, Apache 2.0) tops the Hugging Face open-source ASR leaderboard with an average word error rate of 5.42%, but IBM Granite Speech 4.1 2B (5.33%) and ARK-ASR-3B (5.04%) closely follow, with less than 1 WER point separating them from the leader.","category":"Insights","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Comparison of the Best Open Source Speech Recognition Models in 2026: WER, Language, Latency, and License - Aioga AI News","description":"The field of open-source speech recognition is no longer dominated solely by Whisper. Cohere Transcribe (2B, Apache 2.0) tops the Hugging Face open-source ASR leaderboard with an a...","url":"https://www.aioga.com/en/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:42:21.389Z"},"ja":{"title":"2026年の最高のオープンソース音声認識モデル比較：WER、言語、遅延、ライセンス","summary":"オープンソース音声認識の分野はもはやWhisperだけが独占しているわけではありません。Cohere Transcribe（2B、Apache 2.0）は平均単語誤り率5.42%でHugging FaceのオープンソースASRランキングのトップに立ちましたが、IBM Granite Speech 4.1 2B（5.33%）とARK-ASR-3B（5.04%）が続き、トップとの差はWERで1ポイントにも満たない状況です。","category":"ヒントと視点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"2026年の最高のオープンソース音声認識モデル比較：WER、言語、遅延、ライセンス - Aioga AIニュース","description":"オープンソース音声認識の分野はもはやWhisperだけが独占しているわけではありません。Cohere Transcribe（2B、Apache 2.0）は平均単語誤り率5.42%でHugging FaceのオープンソースASRランキングのトップに立ちましたが、IBM Granite Speech 4.1 2B（5.33%）とARK-ASR-3B（5.04%）...","url":"https://www.aioga.com/ja/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:42:31.436Z"},"ko":{"title":"2026년 최고의 오픈소스 음성 인식 모델 비교: WER, 언어, 지연 및 라이선스","summary":"오픈소스 음성 인식 분야는 더 이상 Whisper가 독점하지 않습니다. Cohere Transcribe(2B, Apache 2.0)는 평균 단어 오류율 5.42%로 Hugging Face 오픈소스 ASR 순위에서 1위를 차지했지만, IBM Granite Speech 4.1 2B(5.33%)와 ARK-ASR-3B(5.04%)가 바로 뒤를 잇고 있어, 1위와의 차이는 1 WER 포인트도 되지 않습니다.","category":"인사이트","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"2026년 최고의 오픈소스 음성 인식 모델 비교: WER, 언어, 지연 및 라이선스 - Aioga AI 뉴스","description":"오픈소스 음성 인식 분야는 더 이상 Whisper가 독점하지 않습니다. Cohere Transcribe(2B, Apache 2.0)는 평균 단어 오류율 5.42%로 Hugging Face 오픈소스 ASR 순위에서 1위를 차지했지만, IBM Granite Speech 4.1 2B(5.33%)와 ARK-ASR-3B(5.04%...","url":"https://www.aioga.com/ko/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:43:11.825Z"},"es":{"title":"Comparación de los mejores modelos de reconocimiento de voz de código abierto en 2026: WER, idioma, latencia y licencia","summary":"El campo del reconocimiento de voz de código abierto ya no está dominado únicamente por Whisper. Cohere Transcribe (2B, Apache 2.0) encabeza la clasificación de ASR de código abierto de Hugging Face con una tasa promedio de error de palabras del 5,42%, pero IBM Granite Speech 4.1 2B (5,33%) y ARK-ASR-3B (5,04%) le siguen de cerca, con una diferencia en la cima de menos de un punto de WER.","category":"Ideas","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Comparación de los mejores modelos de reconocimiento de voz de código abierto en 2026: WER, idioma, latencia y licencia - Aioga Noticias de IA","description":"El campo del reconocimiento de voz de código abierto ya no está dominado únicamente por Whisper. Cohere Transcribe (2B, Apache 2.0) encabeza la clasificación de ASR de código abier...","url":"https://www.aioga.com/es/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:43:10.363Z"},"fr":{"title":"Comparaison des meilleurs modèles de reconnaissance vocale open source en 2026 : WER, langue, latence et licence","summary":"Le domaine de la reconnaissance vocale open source n'est plus dominé uniquement par Whisper. Cohere Transcribe (2B, Apache 2.0) atteint la première place du classement ASR open source de Hugging Face avec un taux d'erreur moyen des mots de 5,42 %, mais IBM Granite Speech 4.1 2B (5,33 %) et ARK-ASR-3B (5,04 %) ne sont pas loin derrière, l'écart avec le premier étant inférieur à 1 point de WER.","category":"Analyses","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Comparaison des meilleurs modèles de reconnaissance vocale open source en 2026 : WER, langue, latence et licence - Aioga Actualités IA","description":"Le domaine de la reconnaissance vocale open source n'est plus dominé uniquement par Whisper. Cohere Transcribe (2B, Apache 2.0) atteint la première place du classement ASR open sou...","url":"https://www.aioga.com/fr/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:43:49.933Z"},"de":{"title":"Vergleich der besten Open-Source-Spracherkennungsmodelle 2026: WER, Sprache, Latenz und Lizenz","summary":"Das Feld der Open-Source-Spracherkennung wird nicht mehr nur von Whisper dominiert. Cohere Transcribe (2B, Apache 2.0) führt die Hugging Face Open-Source-ASR-Rangliste mit einer durchschnittlichen Wortfehlerrate von 5,42 % an, doch IBM Granite Speech 4.1 2B (5,33 %) und ARK-ASR-3B (5,04 %) folgen dicht dahinter, der Abstand an der Spitze beträgt weniger als einen WER-Punkt.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Vergleich der besten Open-Source-Spracherkennungsmodelle 2026: WER, Sprache, Latenz und Lizenz - Aioga KI-News","description":"Das Feld der Open-Source-Spracherkennung wird nicht mehr nur von Whisper dominiert. Cohere Transcribe (2B, Apache 2.0) führt die Hugging Face Open-Source-ASR-Rangliste mit einer du...","url":"https://www.aioga.com/de/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:43:50.038Z"},"pt-BR":{"title":"Comparação dos melhores modelos de reconhecimento de voz de código aberto de 2026: WER, idioma, latência e licença","summary":"O campo de reconhecimento de fala de código aberto não é mais dominado apenas pelo Whisper. O Cohere Transcribe (2B, Apache 2.0) lidera o ranking de ASR de código aberto da Hugging Face com uma taxa média de erro de palavras de 5,42%, mas o IBM Granite Speech 4.1 2B (5,33%) e o ARK-ASR-3B (5,04%) vêm logo atrás, com a diferença para o primeiro colocado sendo inferior a 1 ponto de WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Comparação dos melhores modelos de reconhecimento de voz de código aberto de 2026: WER, idioma, latência e licença - Aioga Notícias de IA","description":"O campo de reconhecimento de fala de código aberto não é mais dominado apenas pelo Whisper. O Cohere Transcribe (2B, Apache 2.0) lidera o ranking de ASR de código aberto da Hugging...","url":"https://www.aioga.com/pt-BR/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:44:24.569Z"},"ru":{"title":"Сравнение лучших моделей распознавания речи с открытым исходным кодом в 2026 году: WER, язык, задержка и лицензия","summary":"Область открытого распознавания речи больше не принадлежит только Whisper. Cohere Transcribe (2B, Apache 2.0) возглавила рейтинг открытого ASR на Hugging Face с средней ошибкой слов 5,42%, но IBM Granite Speech 4.1 2B (5,33%) и ARK-ASR-3B (5,04%) следуют сразу за ней, при этом разрыв с лидером составляет менее одного пункта WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Сравнение лучших моделей распознавания речи с открытым исходным кодом в 2026 году: WER, язык, задержка и лицензия - Aioga Новости ИИ","description":"Область открытого распознавания речи больше не принадлежит только Whisper. Cohere Transcribe (2B, Apache 2.0) возглавила рейтинг открытого ASR на Hugging Face с средней ошибкой сло...","url":"https://www.aioga.com/ru/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:44:30.455Z"},"ar":{"title":"مقارنة أفضل نماذج التعرف على الصوت مفتوحة المصدر لعام 2026: معدل خطأ الكلمات، اللغة، التأخير والترخيص","summary":"لم يعد مجال التعرف على الكلام مفتوح المصدر محتكراً على Whisper وحدها. احتلت Cohere Transcribe (2B، Apache 2.0) المرتبة الأولى في قائمة Hugging Face للتعرف على الكلام مفتوح المصدر بمعدل خطأ كلمات متوسط قدره 5.42٪، لكن IBM Granite Speech 4.1 2B (5.33٪) و ARK-ASR-3B (5.04٪) تبعاهما مباشرة، والفارق مع القمة أقل من نقطة WER واحدة.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"مقارنة أفضل نماذج التعرف على الصوت مفتوحة المصدر لعام 2026: معدل خطأ الكلمات، اللغة، التأخير والترخيص - Aioga أخبار الذكاء الاصطناعي","description":"لم يعد مجال التعرف على الكلام مفتوح المصدر محتكراً على Whisper وحدها. احتلت Cohere Transcribe (2B، Apache 2.0) المرتبة الأولى في قائمة Hugging Face للتعرف على الكلام مفتوح المصدر ب...","url":"https://www.aioga.com/ar/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:45:08.820Z"},"hi":{"title":"2026 में सर्वश्रेष्ठ ओपन-सोर्स भाषण मान्यता मॉडल की तुलना: WER, भाषा, विलंब और लाइसेंस","summary":"ओपन-सोर्स भाषण पहचान के क्षेत्र में अब Whisper एकमात्र प्रमुख नहीं रह गया है। Cohere Transcribe (2B, Apache 2.0) ने औसत शब्द त्रुटि दर 5.42% के साथ Hugging Face के ओपन-सोर्स ASR रैंकिंग में शीर्ष स्थान हासिल किया, लेकिन IBM Granite Speech 4.1 2B (5.33%) और ARK-ASR-3B (5.04%) उसके तुरंत पीछे हैं, और शीर्ष स्थान का अंतर 1 WER अंक से कम है।","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"2026 में सर्वश्रेष्ठ ओपन-सोर्स भाषण मान्यता मॉडल की तुलना: WER, भाषा, विलंब और लाइसेंस - Aioga AI समाचार","description":"ओपन-सोर्स भाषण पहचान के क्षेत्र में अब Whisper एकमात्र प्रमुख नहीं रह गया है। Cohere Transcribe (2B, Apache 2.0) ने औसत शब्द त्रुटि दर 5.42% के साथ Hugging Face के ओपन-सोर्स ASR रै...","url":"https://www.aioga.com/hi/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:45:10.112Z"},"it":{"title":"Confronto dei migliori modelli di riconoscimento vocale open source del 2026: WER, lingua, latenza e licenza","summary":"Il campo del riconoscimento vocale open source non è più dominato esclusivamente da Whisper. Cohere Transcribe (2B, Apache 2.0) ha raggiunto la vetta della classifica ASR open source di Hugging Face con un tasso medio di errore delle parole del 5,42%, ma IBM Granite Speech 4.1 2B (5,33%) e ARK-ASR-3B (5,04%) lo seguono da vicino, con una differenza dalla prima posizione inferiore a un punto WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Confronto dei migliori modelli di riconoscimento vocale open source del 2026: WER, lingua, latenza e licenza - Aioga Notizie IA","description":"Il campo del riconoscimento vocale open source non è più dominato esclusivamente da Whisper. Cohere Transcribe (2B, Apache 2.0) ha raggiunto la vetta della classifica ASR open sour...","url":"https://www.aioga.com/it/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:45:51.634Z"},"nl":{"title":"Vergelijking van de beste open source spraakherkenningsmodellen van 2026: WER, taal, vertraging en licentie","summary":"Het open-source spraakherkenningsgebied wordt niet langer gedomineerd door alleen Whisper. Cohere Transcribe (2B, Apache 2.0) staat bovenaan de Hugging Face open-source ASR-ranglijst met een gemiddelde woordfoutpercentage van 5,42%, maar IBM Granite Speech 4.1 2B (5,33%) en ARK-ASR-3B (5,04%) volgen op de voet, met minder dan 1 WER-punt verschil met de koppositie.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Vergelijking van de beste open source spraakherkenningsmodellen van 2026: WER, taal, vertraging en licentie - Aioga AI-nieuws","description":"Het open-source spraakherkenningsgebied wordt niet langer gedomineerd door alleen Whisper. Cohere Transcribe (2B, Apache 2.0) staat bovenaan de Hugging Face open-source ASR-ranglij...","url":"https://www.aioga.com/nl/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:45:47.547Z"},"tr":{"title":"2026 Yılı En İyi Açık Kaynaklı Ses Tanıma Modelleri Karşılaştırması: WER, Dil, Gecikme ve Lisans","summary":"Açık kaynaklı konuşma tanıma alanı artık sadece Whisper'ın hakimiyetinde değil. Cohere Transcribe (2B, Apache 2.0) %5,42 ortalama kelime hatası oranı ile Hugging Face açık kaynak ASR sıralamasının zirvesine oturdu, ancak IBM Granite Speech 4.1 2B (%5,33) ve ARK-ASR-3B (%5,04) hemen ardından geldi ve ziraledeki fark 1 WER puanından azdı.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"2026 Yılı En İyi Açık Kaynaklı Ses Tanıma Modelleri Karşılaştırması: WER, Dil, Gecikme ve Lisans - Aioga AI Haberleri","description":"Açık kaynaklı konuşma tanıma alanı artık sadece Whisper'ın hakimiyetinde değil. Cohere Transcribe (2B, Apache 2.0) %5,42 ortalama kelime hatası oranı ile Hugging Face açık kaynak A...","url":"https://www.aioga.com/tr/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:46:33.269Z"},"vi":{"title":"So sánh các mô hình nhận dạng giọng nói mã nguồn mở tốt nhất năm 2026: WER, ngôn ngữ, độ trễ và giấy phép","summary":"Lĩnh vực nhận diện giọng nói mã nguồn mở không còn chỉ do Whisper thống trị. Cohere Transcribe (2B, Apache 2.0) dẫn đầu bảng xếp hạng ASR mã nguồn mở của Hugging Face với tỷ lệ lỗi từ trung bình 5,42%, nhưng IBM Granite Speech 4.1 2B (5,33%) và ARK-ASR-3B (5,04%) theo sát ngay sau, khoảng cách với vị trí dẫn đầu chưa đến 1 điểm WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"So sánh các mô hình nhận dạng giọng nói mã nguồn mở tốt nhất năm 2026: WER, ngôn ngữ, độ trễ và giấy phép - Tin tức AI Aioga","description":"Lĩnh vực nhận diện giọng nói mã nguồn mở không còn chỉ do Whisper thống trị. Cohere Transcribe (2B, Apache 2.0) dẫn đầu bảng xếp hạng ASR mã nguồn mở của Hugging Face với tỷ lệ lỗi...","url":"https://www.aioga.com/vi/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:46:28.648Z"},"id":{"title":"Perbandingan Model Pengenalan Suara Open Source Terbaik 2026: WER, Bahasa, Latensi, dan Lisensi","summary":"Bidang pengenalan suara sumber terbuka tidak lagi didominasi oleh Whisper saja. Cohere Transcribe (2B, Apache 2.0) menduduki puncak peringkat ASR sumber terbuka Hugging Face dengan tingkat kesalahan rata-rata kata 5,42%, tetapi IBM Granite Speech 4.1 2B (5,33%) dan ARK-ASR-3B (5,04%) mengikuti di belakangnya, selisih dari posisi teratas kurang dari 1 poin WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Perbandingan Model Pengenalan Suara Open Source Terbaik 2026: WER, Bahasa, Latensi, dan Lisensi - Berita AI Aioga","description":"Bidang pengenalan suara sumber terbuka tidak lagi didominasi oleh Whisper saja. Cohere Transcribe (2B, Apache 2.0) menduduki puncak peringkat ASR sumber terbuka Hugging Face dengan...","url":"https://www.aioga.com/id/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:47:09.558Z"},"th":{"title":"การเปรียบเทียบโมเดลรู้จำเสียงพูดโอเพนซอร์สที่ดีที่สุดในปี 2026: WER, ภาษา, ความล่าช้า และใบอนุญาต","summary":"ในด้านการรู้จำเสียงแบบโอเพ่นซอร์สไม่ใช่เรื่องที่ Whisper ครองตลาดเพียงรายเดียวอีกต่อไป Cohere Transcribe (2B, Apache 2.0) ขึ้นอันดับหนึ่งในตารางคะแนน ASR แบบโอเพ่นซอร์สของ Hugging Face ด้วยค่าอัตราความผิดพลาดคำเฉลี่ย 5.42% แต่ IBM Granite Speech 4.1 2B (5.33%) และ ARK-ASR-3B (5.04%) ตามมาอย่างใกล้ชิด โดยมีความต่างเพียงไม่ถึง 1 จุด WER จากอันดับหนึ่ง","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"การเปรียบเทียบโมเดลรู้จำเสียงพูดโอเพนซอร์สที่ดีที่สุดในปี 2026: WER, ภาษา, ความล่าช้า และใบอนุญาต - ข่าว AI Aioga","description":"ในด้านการรู้จำเสียงแบบโอเพ่นซอร์สไม่ใช่เรื่องที่ Whisper ครองตลาดเพียงรายเดียวอีกต่อไป Cohere Transcribe (2B, Apache 2.0) ขึ้นอันดับหนึ่งในตารางคะแนน ASR แบบโอเพ่นซอร์สของ Hugging...","url":"https://www.aioga.com/th/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:47:15.181Z"},"pl":{"title":"Porównanie najlepszych modeli rozpoznawania mowy open source w 2026 roku: WER, język, opóźnienie i licencja","summary":"Dziedzina otwartego oprogramowania do rozpoznawania mowy nie jest już zdominowana wyłącznie przez Whisper. Cohere Transcribe (2B, Apache 2.0) zajmuje pierwsze miejsce w rankingu otwartego ASR na Hugging Face z średnim wskaźnikiem błędów słów na poziomie 5,42%, jednak IBM Granite Speech 4.1 2B (5,33%) i ARK-ASR-3B (5,04%) są tuż za nim, a różnica na szczycie wynosi mniej niż 1 punkt WER.","category":"技巧观点","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Porównanie najlepszych modeli rozpoznawania mowy open source w 2026 roku: WER, język, opóźnienie i licencja - Aioga Wiadomości AI","description":"Dziedzina otwartego oprogramowania do rozpoznawania mowy nie jest już zdominowana wyłącznie przez Whisper. Cohere Transcribe (2B, Apache 2.0) zajmuje pierwsze miejsce w rankingu ot...","url":"https://www.aioga.com/pl/news/cmrxc6h0d01a6rot3gv44kv9z/","contentTranslated":true,"sourceHash":"56024f368d6b32c9","translatedAt":"2026-07-26T09:48:00.233Z"}}}}