{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-29T18:20:55.196Z","headline":"Open ASR 排行榜新增首个全球南方语言：印地语与印度英语评测集","description":"Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN 两个评测集，覆盖印地语与印度英语，其中印地语是该排行榜多语言板块首个非欧洲语言。数据集含公开与私有分割，共 4，888 位说话人，并记录 12 项说话人属性，旨在暴露按地区、年龄、性别等维度分布不均的语音识别误差。","url":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/","mainEntityOfPage":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/","datePublished":"2026-08-28T00:00:00.000Z","dateModified":"2026-08-28T00:00:00.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://huggingface.co/blog/open-asr-leaderboard-global-south","https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x"],"canonicalUrl":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN Aioga 将其归入「产品更新」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/","dateCreated":"2026-08-28T00:00:00.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"huggingface.co source article","url":"https://huggingface.co/blog/open-asr-leaderboard-global-south","datePublished":"2026-08-28T00:00:00.000Z","provider":{"@type":"Organization","name":"huggingface.co","url":"https://huggingface.co/blog/open-asr-leaderboard-global-south"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","datePublished":"2026-08-28T00:00:00.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x"}}],"aggregationSource":"Hugging Face：Blog（RSS）","originalPublisher":{"name":"huggingface.co","url":"https://huggingface.co/blog/open-asr-leaderboard-global-south"},"geoDeepAnswer":null,"article":{"id":"cmtdgo8nh04rurobxlzwr101x","slug":"cmtdgo8nh04rurobxlzwr101x","url":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/","title":"Open ASR 排行榜新增首个全球南方语言：印地语与印度英语评测集","title_en":"The Open ASR Leaderboard Adds Its First Global South Language","summary":"Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN 两个评测集，覆盖印地语与印度英语，其中印地语是该排行榜多语言板块首个非欧洲语言。数据集含公开与私有分割，共 4，888 位说话人，并记录 12 项说话人属性，旨在暴露按地区、年龄、性别等维度分布不均的语音识别误差。","source":"Hugging Face：Blog（RSS）","sourceUrl":"https://huggingface.co/blog/open-asr-leaderboard-global-south","aiHotUrl":"https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","publishedAt":"2026-08-28T00:00:00.000Z","category":"产品更新","score":66,"selected":true,"articleBody":["Design of the collection ：#design-of-the-collection Dataset composition ：#dataset-composition Speaker coverage ：#speaker-coverage Metadata fields ：#metadata-fields Collection and quality control ：#collection-and-quality-control Regional variation ：#regional-variation Orthographic variation in Hindi ：#orthographic-variation-in-hindi Getting evaluated ：#getting-evaluated What comes next ：#what-comes-next Voice Arena and Hugging Face partner to launch open ASR evaluation for Hindi and Indian English","Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:","All of that makes one number (WER) harder to game. It is still one number. A long line of work has shown that ASR error rates are not evenly distributed across the people using them. Racial disparities in automated speech recognition：https://www.pnas.org/doi/10.1073/pnas.1915768117 found commercial systems roughly twice as bad for Black speakers as for white speakers, and Quantifying Bias in Automatic Speech Recognition：https://huggingface.co/papers/2103.15122 found further differences by gender, age and accent. None of that is visible on a leaderboard, and not because the leaderboard is hiding it. The test sets it runs on record what was said and almost nothing about who said it.","To address this gap, we introduce two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN：https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN and Monsoon hi-IN：https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN. Hindi, spoken by more than half a billion people , is the first Indic language on a multilingual tab that currently covers only European languages. Each set is released as a public split, available for self-scoring, and a private split withheld to limit benchmark-specific optimisation. The four splits are speaker-disjoint, comprising 4,888 speakers, with 12 speaker attributes recorded for each.","A test set can only expose a failure mode it varies along. Most benchmarks are built from whatever audio was readily available. Monsoon was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio. Each is a way an aggregate WER can be right on average and wrong for a particular population.","：https://storage.googleapis.com/research_team_data/blog_figures/nine-axes-gray.png","The collection method follows from that.","Four splits, two languages, collected through one pipeline.","The data is sourced from unscripted dual-channel spontaneous conversations, with clips segmented from a single channel so that each clip carries one speaker. Along with the fields reported in the table, each clip also records occupation, education, marital status, income band, handset brand, current city and years in the current district.","Five clips from the public Indian English split, with the metadata each one carries:","29-year-old woman, West Tripura, Tripura. Student, samsung SM-G781B.","32-year-old woman, Satna, Madhya Pradesh. Unemployed, samsung SM-E146B.","22-year-old man, Rohtas, Bihar. Student, motorola moto g54 5G.","27-year-old woman, Warangal, Telangana. Unemployed, vivo V2247.","57-year-old man, Puducherry. Private job, Xiaomi M2006C3LI.","The English sets use standard string references, where the leaderboard's normaliser collapses most spelling variation. Hindi has far more of it, and no normaliser can resolve it, because the variants are not a fixed mapping between two conventions. The Hindi sets therefore ship a lattice: for each span of the transcript, a list of the spellings that are accepted as correct.","Monsoon is small measured in hours and large measured in speakers. That is the design, and it is where most of the value sits.","Speaker concentration and diversity beyond the fields above.","Three properties follow, and each is a claim about variance rather than volume.","No voice carries the score: The ten largest contributors account for between 2.8% and 6.8% of total duration, and more than half of all speakers appear exactly once. A result on Monsoon is an average over hundreds of distinct voices, not a small number of talkers recorded at length. Test sets of comparable duration are usually constructed the other way.","No region or handset carries it either: The Indian English public set draws on 428 native districts across 30 states and union territories; the Hindi sets, being a Hindi-belt language, concentrate more tightly but still span 202 and 295 districts. Recordings come from 315 to 582 distinct device models, with no single model exceeding 2.1% of segments in any subset. Corpora collected on standardised hardware overfit to one microphone response; this one cannot.","Indian English here is not one accent: This is English as it is spoken across the country, not the English of one region. All six zones are represented: in the public set, 35% of segments are contributed by southern speakers, 18% from the East, 18% from Central, 16% from the North and 11% from the West. The accent variation that follows from that spread is recorded in the metadata rather than asserted.","Monsoon ships 18 columns per segment, of which 12 are metadata, where most public ASR test sets ship an identifier, a transcript and a duration. Demographic fields are complete or near-complete; contributors consented to this use.","The two languages have different geographic shapes, and the shape is informative. The Hindi sets concentrate in the Hindi belt, with Uttar Pradesh accounting for roughly 40% of speakers, which is what a Hindi corpus sampled by population should look like. The Indian English sets are much flatter: no state exceeds 13%, and a third of speakers come from outside the eight largest. Public and private halves match closely on both.","Indian state boundaries were drawn along linguistic lines, so district and state carry real accent signal, which is why these fields are released rather than summarised away. Analysis of this kind has been reported at scale for Indian ASR: district-level error rates spanning roughly 4% to 44%, with underrepresented regions well behind the Hindi belt and the metros, plus disaggregation by audio quality, speaking rate, utterance duration, gender, age and device. Those runs were on a closed benchmark. Monsoon makes the same class of analysis possible on a public leaderboard test set.","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%201.png","Broad geographic coverage requires recruitment across hundreds of districts rather than longer sessions from fewer speakers, and distributed recruitment at this scale introduces failure modes that a smaller collection does not face: contributors gaming the task, played-back audio submitted as live speech, and inattentive annotation. Each is addressed by an explicit check.","Recruitment and recording: Contributors were recruited through the Voice Arena community, a global digital platform whose reach extends into the rural and semi-urban districts that speech corpora rarely cover. Pairs then recorded two-person conversations over a peer-to-peer interface, dual-channel, on assigned everyday topics. Contributors used their own handsets and their own connections. Many of those are low-end devices on unstable bandwidth, which is why that condition is present in the released audio rather than filtered out of it. Prospective contributors completed a language proficiency screening before being granted recording access, were compensated, and provided informed consent covering use in training and distribution. A per-speaker duration cap, calibrated per language to the population size and geographic distribution of its speakers, prevented a small number of prolific contributors from dominating a language or region; more than half of the speakers in these sets contribute exactly one segment.","Elicitation: Eliciting spontaneous speech at scale presents its own difficulty, as contributors tend to produce short and sparse responses without structured guidance. Each conversation was therefore seeded with an open-ended narrative cue and progressively revealed follow-up questions, spanning domains including travel, healthcare, agriculture, education and digital services, guiding the exchange toward extended description without scripting it. Candidate topics were generated with large language models, then reviewed and localised by native-speaker linguists.","Quality control: Every recording passed a set of gating checks prior to transcription. The spoken language was verified against the assigned language using language identification models trained on human-annotated data across more than 30 languages. Speaker gender was confirmed against the self-reported label using a dedicated classifier, applied as corroboration of the self-report rather than as a replacement for it. A further model distinguished genuine spontaneous conversation from pre-recorded or played-back audio. Signal-to-noise ratio estimation removed recordings degraded beyond intelligibility, while natural environmental background noise was deliberately preserved so that the acoustic realism of in-the-wild speech is retained. Recordings clearing these checks were segmented by voice activity detection, split at two seconds of continuous silence or at a fifteen-second soft cap closed at the next detected silence. Segmentation was applied independently per channel, so every segment is single-speaker and single-channel. Segments then passed a DNSMOS P.808 check.","What follows is one example, run on the public Indian English split, to show the kind of evaluation the metadata makes possible. It is not the finding the sets exist to deliver; it is an illustration of what becomes answerable once every clip carries a speaker.","Eight models on the leaderboard land between 4.81 and 4.99 WER on this set. That is 0.18 points from best to worst, inside what five hours can resolve. Ranked on the corpus, they are the same model.","Grouping speakers by region tells a different story. Each speaker's native district is rolled up to its zonal council, the Ministry of Home Affairs grouping of Indian states, giving five well-sampled zones. openai/whisper-large-v3-turbo：https://huggingface.co/openai/whisper-large-v3-turbo varies by 0.46 points across them. mistralai/Voxtral-Mini-3B-2507：https://huggingface.co/mistralai/Voxtral-Mini-3B-2507, fourteen hundredths of a point behind it on the corpus, varies by 1.68, running 4.38 in the Central zone against 6.06 in the East. Two systems that are indistinguishable on the leaderboard differ almost fourfold in how much their accuracy depends on where the speaker is from.","：https://storage.googleapis.com/research_team_data/csv_files/corpus_wer_versus_zone_range_coloured_by_zone.png","Which zone is hardest is not fixed either. ibm-granite/granite-speech-3.3-2b：https://huggingface.co/ibm-granite/granite-speech-3.3-2b is worst in the North, microsoft/VibeVoice-ASR-HF：https://huggingface.co/microsoft/VibeVoice-ASR-HF in the South, mistralai/Voxtral-Mini-3B-2507：https://huggingface.co/mistralai/Voxtral-Mini-3B-2507 in the East. If a single region were simply harder to transcribe, every model would rank the zones the same way. They do not, which points at the models rather than the audio.","Region is one of twelve recorded attributes, and the zones above are a coarse rollup of 428 districts. The same breakdown runs on age, education, occupation and handset, and the released files carry everything needed to reproduce it. None of it is available for a test set that records only what was said.","English orthographic variation is bounded. British against American spelling, punctuation, casing, digits against words: a normaliser can map most of it to a single form, and the leaderboard's does. Hindi is not bounded in the same way. Everyday speech is heavily code-mixed, English-origin words have no settled Devanagari spelling, and compound forms are written joined or separated according to preference. A single phrase can have ten or more valid written forms, and no fixed mapping collapses them, because there is no canonical side to map to.","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%202.png","Scored with a single reference, WER rewards a system for producing the spelling the annotator happened to choose. Two systems that recognised the audio equally well can differ by several points on orthography alone.","The Hindi sets therefore ship a lattice: for each span of the transcript, the set of written forms accepted as correct. Building it is manual work. Candidate variants are drawn from multiple ASR transcripts of the same audio and expanded with language models, then native-speaker linguists decide which are valid for that utterance and prune the rest, so only forms consistent with what was said are admitted.","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%203a.png","Thus, for Hindi we report the Orthographically-Informed Word Error Rate (OIWER)：https://huggingface.co/papers/2603.00941, introduced by AI4Bharat, instead of WER. A hypothesis is aligned against the accepted set at each span, so any admitted form counts as correct and only genuine recognition errors are charged.","To quantify the effect, the same hypotheses were scored twice. Flattening each lattice to its first variant per span yields a single string reference of the kind a conventional benchmark provides; any of the admitted variants would serve equally well, and a different choice would yield a different reference. Scored against the flattened reference, error rates rise for every system, and they do not rise uniformly. Rankings change as a consequence. The figure below shows two pairs of systems that reverse order between the two references: under a single reference a system is rewarded in part for reproducing the annotator's orthography, whereas the lattice scores only recognition.","：https://storage.googleapis.com/research_team_data/csv_files/hindi_oiwer_flips.png","We also open source our implementation, voi-oiwer：https://pypi.org/project/voi-oiwer, so every result on these sets can be reproduced directly.","For the private splits, get your model on the Open ASR Leaderboard and the Hugging Face team will run the evaluation. As before, the process for adding a model to the leaderboard takes place on the Open ASR Leaderboard GitHub：https://github.com/huggingface/open_asr_leaderboard:","Indian-English joins the main leaderboard：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard as Voice Arena Monsoon , in the default column set rather than as an opt-in toggle, so it contributes to the headline Average WER for every model. The private split feeds the aggregated Private (conversational) column alongside Appen and DataoceanAI data：https://huggingface.co/blog/open-asr-leaderboard-private-data. The public and private Hindi appears in the Multilingual tab：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard, where a model is ranked only if it supports every selected language, making that column a like-for-like comparison. Alternative, select \"Hindi\" from the \"Language dataset breakdown\" dropdown menu.","Hindi is a sharp case of a general problem. Any language written more than one way, spoken by people a benchmark has not sampled, carries both of the failures described here. These sets do not fix them. What they add is a way to see them: a test set on a leaderboard the field already watches, carrying enough about each speaker and each reference that a difference between two systems can be traced to who was talking and how they write it, instead of disappearing into one number.","These four sets are part of Monsoon, Voice Arena's broader dataset initiative for the Global South.","Explore and compare speech‑recognition models by WER and speed"],"articleImages":[{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62a0baa7c6754e9bfabf7e76/aogJMfrZ_Cvgx1ka_IDTV.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmtdgo8nh04rurobxlzwr101x/3b9098936998b8c0.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6384db7fb2906edaf835a91d/MOTXxaOmjlTZ8wONYifnD.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmtdgo8nh04rurobxlzwr101x/fd4789dda847dae4.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641ae854ec5b871c0bcb5f36/Nvb14xFDGQnExnOZMCasb.png","alt":"","afterParagraph":0,"url":"/media/articles/cmtdgo8nh04rurobxlzwr101x/43a2b8a03541fd77.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c8747255abbb02c8b285ed/XYs0wxVHu7BO20ZiQvugD.jpeg","alt":"","afterParagraph":0,"url":"/media/articles/cmtdgo8nh04rurobxlzwr101x/6df95acd71b2bdae.webp"},{"sourceUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67bc3e2f4b9d3615a6e2c982/8x33LQ0fwbnUtIn9CAwv2.png","alt":"","afterParagraph":0,"url":"/media/articles/cmtdgo8nh04rurobxlzwr101x/4f8cd9312c0f3c86.webp"}],"mediaStatus":"ok","articleBodyZh":["合集设计：#design-of-the-collection 数据集构成：#dataset-composition 说话人覆盖：#speaker-coverage 元数据字段：#metadata-fields 收集与质量控制：#collection-and-quality-control 区域差异：#regional-variation 印地语正字法差异：#orthographic-variation-in-hindi 获取评估：#getting-evaluated 接下来是什么：#what-comes-next Voice Arena 与 Hugging Face 合作推出印地语和印度英语的开放式语音识别评估","基准决定了什么会被开发。一个在开放 ASR 排行榜（https://huggingface.co/spaces/hf-audio/open_asr_leaderboard）上得分高的模型会被采用和迭代，而排行榜未衡量的能力往往不会得到提升。最近在排行榜上的许多工作都集中在使评估指标更可信：","所有这些使得一个指标（WER）更难被操控，但它仍然是一个数字。一长串研究显示，ASR 错误率在使用它的人群中分布并不均匀。《自动语音识别中的种族差异》（https://www.pnas.org/doi/10.1073/pnas.1915768117）发现商业系统对黑人说话人的表现大约是对白人的两倍差，而《量化自动语音识别中的偏差》（https://huggingface.co/papers/2103.15122）则发现按性别、年龄和口音还有其他差异。这些在排行榜上都看不到，并不是排行榜在隐藏它。这些测试集记录了说了什么，但几乎没有关于说话人是谁的信息。","为了解决这一空白，我们在开放 ASR 排行榜引入了两个评估集：Monsoon en-IN（https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN）和 Monsoon hi-IN（https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN）。印地语由超过五亿人使用，是多语言选项卡上的第一种印度语，目前多语言选项卡仅覆盖欧洲语言。每个数据集都作为公开分割发布，可用于自我评分，同时保留私有分割以限制基准特定的优化。这四个分割在说话人上不重叠，共包括 4,888 个说话人，每位说话人记录 12 个属性。","一个测试集只能揭示它变化的失败模式。大多数基准测试都是基于现成可用的音频构建的。Monsoon 测试集是沿九个轴构建的：地理位置、年龄、性别、词汇、设备、声学环境、语音类型、语速，以及同一音频存在多个有效转录的情况。每一项都是平均词错误率（WER）可能正确，但对特定人群可能错误的因素。","：https://storage.googleapis.com/research_team_data/blog_figures/nine-axes-gray.png","数据收集方法是由此而来的。","四个分集，两个语言，通过一个处理流程收集。","数据来源于未经脚本的双声道自发对话，剪辑从单声道中分割，以确保每个剪辑只有一个说话者。除了表中报告的字段外，每个剪辑还记录职业、教育、婚姻状况、收入等级、手机品牌、当前城市和在当前地区的居住年数。","来自公共印度英语分集的五个剪辑及其携带的元数据：","29岁女性，西特里普拉，特里普拉邦。学生，三星 SM-G781B。","32岁女性，萨特纳，中央邦。失业，三星 SM-E146B。","22岁男性，罗塔斯，比哈尔邦。学生，摩托罗拉 moto g54 5G。","27岁女性，瓦兰格尔，特伦甘纳邦。失业，vivo V2247。","57岁男性，印度本地直辖区。私企员工，小米 M2006C3LI。","英语集合使用标准字符串引用，排行榜的规范化器可以合并大部分拼写变体。印地语有更多的拼写变体，且没有任何规范化器可以解决，因为这些变体不是两种约定之间的固定映射。因此，印地语集合提供一个格点：对于转录的每个片段，列出被接受为正确的拼写列表。","Monsoon 的小时数上规模小，按说话者数量上规模大。这就是设计理念，也是其大部分价值所在。","说话者集中度和上述字段之外的多样性。","随之而来的是三个属性，每一个都是关于方差的声明，而非数量。","没有任何一个声音主导评分：前十名贡献者占总语音时长的比例在2.8%至6.8%之间，而超过一半的说话者只出现一次。Monsoon 上的结果是数百个不同声音的平均，而不是少数说话者的长时间录音。通常具有相当时长的测试集通常以相反的方式构建。","没有任何地区或设备主导评分：印度英语公共数据集涵盖了来自30个邦和联邦属地的428个本地区；而印地语数据集作为印地语带地区的语言，分布更加集中，但仍覆盖202到295个区。录音来自315到582种不同设备型号，在任何子集里没有单一型号占比超过2.1%。在标准化硬件上收集的语料会过拟合某一麦克风响应，而该语料不会。","这里的印度英语并非单一口音：这是全国各地讲的英语，而不是某一区的英语。六个区域都有代表：在公共数据集中，南部说话者贡献了35%的片段，东部贡献18%，中部贡献18%，北部贡献16%，西部贡献11%。由此分布产生的口音差异记录在元数据中，而不是被假定。","Monsoon 每个片段提供18列，其中12列为元数据，而大多数公共自动语音识别测试集只包含一个标识符、一个转录和一个时长。人口统计字段完整或接近完整；贡献者已同意使用这些数据。","两种语言的地理分布形状不同，而这种形状具有信息性。印地语数据集集中在印地语带，北方邦约占说话者总数的40%，这正是按人口抽样的印地语语料应有的样子。印度英语数据集则更加均衡：没有任何一个邦超过13%，八个最大邦之外的说话者占三分之一。公共和私有数据集在两者上匹配情况相近。","印度州界是按语言划定的，因此区和州带有真实的重音信号，这也是为什么这些字段被释放而非被总结。这类分析已在印度ASR的大规模报告中被报告：区级错误率约为4%至44%，代表性不足的地区远远落后于印地语带和大都市区，并按音频质量、发言速度、发言时长、性别、年龄和设备进行细分。这些测试基于封闭基准进行。Monsoon在公开排行榜测试集上实现了同样的分析。","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%201.png","广泛的地理覆盖需要跨数百个学区招募，而非由较少发言者进行更长时间的会议，而这种规模的分布式招聘引入了较小群体不会遇到的失败模式：贡献者钻空子任务、作为直播语音提交的音频播放，以及不专注的注释。每种情况都通过明确的检查来解决。","招募与录音：贡献者通过Voice Arena社区招募，这是一个全球数字平台，其影响力延伸至语音语料库很少覆盖的农村和半城市地区。随后，配对通过点对点接口（双频道）录制两人对话，内容涵盖日常指定话题。贡献者使用自己的手机和连接。其中许多设备是低端设备，带宽不稳定，因此这种状况存在于发布的音频中，而非被过滤掉。潜在贡献者在获得录音访问前完成语言能力筛查，获得报酬，并提供了关于培训和分发使用的知情同意。每个说话者的时长上限，根据语言的人口规模和使用者地理分布校准，防止少数高产贡献者主导某一语言或地区;这些集中超过一半的使用者只贡献一个片段。","引导：大规模引导自发语音本身具有难度，因为贡献者往往在没有结构性指导的情况下产生简短且零散的回应。因此，每次对话都以开放式叙事提示为起点，并逐步呈现后续问题，涵盖旅行、医疗、农业、教育和数字服务等领域，引导交流向详细描述发展而不进行台本化。候选主题由大型语言模型生成，然后由母语语言学家进行审查和本地化。","质量控制：每个录音在转录前都经过一系列门控检查。使用基于人工标注数据训练的语言识别模型，对所说语言与分配语言进行验证，覆盖30多种语言。使用专用分类器确认说话者的性别是否与自报标签一致，该分类器作为自报的辅助验证，而非替代。另有模型区分真实自发对话与预录或播放的音频。信噪比估计会删除可理解性受损的录音，同时有意保留自然环境背景噪声，以保持真实场景语音的声学真实感。通过这些检查的录音会通过语音活动检测进行分段，当出现连续两秒的静默或在15秒的软上限处并在下一次检测到静默时结束。每个通道独立分段，因此每个片段为单说话者、单通道。分段之后通过DNSMOS P.808检查。","下面是一个示例，运行在公开的印度英语数据分集上，以展示元数据可能进行的评估类型。这并不是数据集合的发现目的；它是一个说明，一旦每个片段都带有说话者信息，哪些问题可以得到解答。","排行榜上的八个模型在该数据集上的词错误率（WER）在4.81到4.99之间。最好到最差之间只有0.18个百分点，相当于五小时即可解决的范围。在语料库排名上，它们是同一型号。","按地区对说话者进行分组会讲述不同的故事。每位说话者的母语地区被汇总到其区域委员会，即印度内政部对各邦的分组，形成五个样本充分的区域。openai/whisper-large-v3-turbo：https://huggingface.co/openai/whisper-large-v3-turbo 在各区域之间的变化为0.46分。mistralai/Voxtral-Mini-3B-2507：https://huggingface.co/mistralai/Voxtral-Mini-3B-2507，在语料库上落后它十四百分点，变化为1.68，中部区域为4.38，而东部为6.06。排行榜上难以区分的两个系统，其准确率对说话者来源地的依赖程度几乎相差四倍。","：https://storage.googleapis.com/research_team_data/csv_files/corpus_wer_versus_zone_range_coloured_by_zone.png","哪个区域最难也不是固定的。ibm-granite/granite-speech-3.3-2b：https://huggingface.co/ibm-granite/granite-speech-3.3-2b 在北区表现最差，microsoft/VibeVoice-ASR-HF：https://huggingface.co/microsoft/VibeVoice-ASR-HF 在南区最差，mistralai/Voxtral-Mini-3B-2507：https://huggingface.co/mistralai/Voxtral-Mini-3B-2507 在东区最差。如果某个单独的区域只是更难转录，每个模型对区域的排名都会相同。但事实并非如此，这说明问题出在模型本身，而不是音频。","地区是记录的十二个属性之一，上述区域是428个地区的粗略汇总。相同的划分也适用于年龄、教育、职业和终端设备，发布的文件包含重现所需的一切内容。这些信息在仅记录所说内容的测试集中是不可用的。","英语正字法变异是有界的。英式与美式拼写、标点、字母大小写、数字与单词：标准化工具可以将大部分映射为单一形式，排行榜上的方法正是如此。印地语不具备相同的界限。日常语言中大量混合代码，来源于英语的词没有固定的天城文拼写，复合形式根据偏好可连写或分写。同一句话可能有十种或更多有效书写形式，没有固定映射可以将它们统一，因为不存在可映射的规范形式。","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%202.png","使用单一参考进行评分时，WER 会奖励生成了标注者碰巧选择的拼写的系统。两个在音频识别上同样出色的系统，仅在正字法上就可能相差几个百分点。","因此，印地语数据集提供了一个 lattice（词格）：对于转录的每个片段，列出被接受为正确的书写形式集合。构建它是人工工作。候选变体来源于同一音频的多个 ASR 转录，之后通过语言模型扩展，然后由母语语言学家决定哪些对该语句有效，并修剪掉其余形式，因此只允许与所说内容一致的形式。","：https://storage.googleapis.com/research_team_data/blog_figures/flowchart%203a.png","因此，对于印地语，我们报告正字法信息的词错误率（Orthographically-Informed Word Error Rate, OIWER）：https://huggingface.co/papers/2603.00941，由 AI4Bharat 提出，而不是 WER。将假设与每个片段中被接受的集合进行对齐，因此任何被允许的形式都计为正确，只有真正的识别错误才会被计入。","为了量化其影响，对相同的假设进行了两次评分。将每个 lattice 的每个片段扁平化为其第一个变体，可得到传统基准提供的单个字符串参考；任何被接受的变体都同样有效，而不同的选择会导致不同的参考。对扁平化参考进行评分时，每个系统的错误率都会上升，而且上升幅度不均匀。因此排名会发生变化。下图显示了两对系统在两个参考之间顺序逆转的情况：在单一参考下，系统在一定程度上因复制标注者的正字法而获得奖励，而 lattice 评分只考虑识别。","：https://storage.googleapis.com/research_team_data/csv_files/hindi_oiwer_flips.png","我们还将我们的实现开源了，voi-oiwer：https://pypi.org/project/voi-oiwer，因此这些数据集上的每个结果都可以直接复现。","对于私有划分，将你的模型提交到 Open ASR Leaderboard，Hugging Face 团队将进行评估。与之前一样，将模型添加到排行榜的流程在 Open ASR Leaderboard GitHub 上进行：https://github.com/huggingface/open_asr_leaderboard","印度英语加入了主排行榜：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard，以 Voice Arena Monsoon 的身份出现，在默认列设置中，而不是作为可选切换，因此它会对每个模型的总体平均 WER 做出贡献。私有分割会将聚合的私有（会话型）列与 Appen 和 DataoceanAI 的数据一起提供：https://huggingface.co/blog/open-asr-leaderboard-private-data。公共和私有的印地语出现在多语言标签页：https://huggingface.co/spaces/hf-audio/open_asr_leaderboard，在该标签页中，只有当模型支持所选的每种语言时才会被排名，使该列具有可比性。或者，从“语言数据集细分”下拉菜单中选择“印地语”。","印地语是一个典型的普遍性问题案例。任何一种有多种书写方式的语言，且其使用者未被基准测试抽样，都会存在这里描述的两类失败。现有的数据集并不能解决这些问题。它们所做的是提供一种观察这些问题的方法：在排行榜上的测试集，领域内的人员已经在关注，包含足够多关于每位说话者和每个参考的数据，使得两个系统之间的差异可以追溯到说话的人及其书写方式，而不是消失在一个数字中。","这四个数据集是 Monsoon 的一部分，Voice Arena 针对全球南方的更广泛数据集计划的一部分。","通过 WER 和速度探索并比较语音识别模型"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN Aioga 将其归入「产品更新」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-08-29T18:34:42.433Z","sourceHash":"da9f157eb7b227a6","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["产品更新","Hugging Face：Blog（RSS）"],"translations":{"zh-CN":{"title":"Open ASR 排行榜新增首个全球南方语言：印地语与印度英语评测集","summary":"Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN 两个评测集，覆盖印地语与印度英语，其中印地语是该排行榜多语言板块首个非欧洲语言。数据集含公开与私有分割，共 4，888 位说话人，并记录 12 项说话人属性，旨在暴露按地区、年龄、性别等维度分布不均的语音识别误差。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"huggingface.co","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR 排行榜新增首个全球南方语言：印地语与印度英语评测集 - Aioga AI资讯","description":"Voice Arena 与 Hugging Face 合作，为 Open ASR 排行榜引入 Monsoon en-IN 和 Monsoon hi-IN 两个评测集，覆盖印地语与印度英语，其中印地语是该排行榜多语言板块首个非欧洲语言。数据集含公开与私有分割，共 4，888 位说话人，并记录 12 项说话人属性，旨在暴露按地区、年龄、性别等维度分布不均的语音识...","url":"https://www.aioga.com/news/cmtdgo8nh04rurobxlzwr101x/"},"en":{"title":"Open ASR leaderboard adds the first global Southern language: Hindi and Indian English evaluation set","summary":"Voice Arena has partnered with Hugging Face to introduce two evaluation sets, Monsoon en-IN and Monsoon hi-IN, to the Open ASR leaderboard, covering Hindi and Indian English. Hindi is the first non-European language in the multilingual section of this leaderboard. The dataset contains both public and private splits, including a total of 4,888 speakers, and records 12 speaker attributes, aiming to expose speech recognition errors that are unevenly distributed across dimensions such as region, age, and gender. 🔗 Read the original via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"Industry","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR leaderboard adds the first global Southern language: Hindi and Indian English evaluation set - Aioga AI News","description":"Voice Arena has partnered with Hugging Face to introduce two evaluation sets, Monsoon en-IN and Monsoon hi-IN, to the Open ASR leaderboard, covering Hindi and Indian English. Hindi...","url":"https://www.aioga.com/en/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:22:35.545Z"},"ja":{"title":"Open ASRランキングに初のグローバルサザン言語が追加：ヒンディー語とインド英語の評価データセット","summary":"Voice ArenaはHugging Faceと協力して、Open ASRランキングにMonsoon en-INとMonsoon hi-INの2つの評価データセットを導入しました。これによりヒンディー語とインド英語をカバーしており、ヒンディー語はこのランキングの多言語セクションで初めての非ヨーロッパ言語となります。データセットには公開および非公開の分割が含まれており、合計4,888人の話者が収録され、12項目の話者属性が記録されています。これは地域、年齢、性別などの観点で不均衡な音声認識誤差を明らかにすることを目的としています。 🔗 原文を読む via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"業界動向","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASRランキングに初のグローバルサザン言語が追加：ヒンディー語とインド英語の評価データセット - Aioga AIニュース","description":"Voice ArenaはHugging Faceと協力して、Open ASRランキングにMonsoon en-INとMonsoon hi-INの2つの評価データセットを導入しました。これによりヒンディー語とインド英語をカバーしており、ヒンディー語はこのランキングの多言語セクションで初めての非ヨーロッパ言語となります。データセットには公開および非公開の分割が含...","url":"https://www.aioga.com/ja/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:22:48.278Z"},"ko":{"title":"Open ASR 차트에 최초의 글로벌 남방 언어 추가: 힌디어와 인도 영어 평가 집합","summary":"Voice Arena는 Hugging Face와 협력하여 Open ASR 순위를 위해 Monsoon en-IN 및 Monsoon hi-IN 두 개의 평가 세트를 도입했으며, 인도 영어와 힌디어를 포함하고 있습니다. 이 중 힌디어는 해당 순위의 다국어 섹션에서 최초의 비유럽 언어입니다. 데이터 세트는 공개 및 비공개 분할을 포함하며, 총 4,888명의 화자를 담고 있고, 12개의 화자 속성이 기록되어 있어 지역, 연령, 성별 등 차원에서 불균형하게 나타나는 음성 인식 오류를 드러내는 것을 목표로 합니다. 🔗 원문 읽기 via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"업계 동향","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR 차트에 최초의 글로벌 남방 언어 추가: 힌디어와 인도 영어 평가 집합 - Aioga AI 뉴스","description":"Voice Arena는 Hugging Face와 협력하여 Open ASR 순위를 위해 Monsoon en-IN 및 Monsoon hi-IN 두 개의 평가 세트를 도입했으며, 인도 영어와 힌디어를 포함하고 있습니다. 이 중 힌디어는 해당 순위의 다국어 섹션에서 최초의 비유럽 언어입니다. 데이터 세트는 공개 및 비공개 분할을...","url":"https://www.aioga.com/ko/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:23:49.100Z"},"es":{"title":"La lista de clasificación de Open ASR ha agregado el primer idioma del sur global: conjunto de evaluación de hindi e inglés indio","summary":"Voice Arena se asoció con Hugging Face para introducir en el ranking de Open ASR dos conjuntos de evaluación, Monsoon en-IN y Monsoon hi-IN, que cubren inglés de la India y hindi, siendo el hindi el primer idioma no europeo en la sección multilingüe de este ranking. Los conjuntos de datos contienen divisiones públicas y privadas, con un total de 4,888 hablantes, y registran 12 atributos de los hablantes, con el objetivo de destacar los errores de reconocimiento de voz distribuidos de manera desigual según región, edad, género, entre otros. 🔗 Leer el artículo original vía AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"Industria","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"La lista de clasificación de Open ASR ha agregado el primer idioma del sur global: conjunto de evaluación de hindi e inglés indio - Aioga Noticias de IA","description":"Voice Arena se asoció con Hugging Face para introducir en el ranking de Open ASR dos conjuntos de evaluación, Monsoon en-IN y Monsoon hi-IN, que cubren inglés de la India y hindi,...","url":"https://www.aioga.com/es/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:23:39.984Z"},"fr":{"title":"Le classement Open ASR ajoute la première langue du Sud global : le hindi et le corpus d'évaluation de l'anglais indien","summary":"Voice Arena collabore avec Hugging Face pour introduire les ensembles de test Monsoon en-IN et Monsoon hi-IN dans le classement Open ASR, couvrant l'anglais indien et l'hindi, ce dernier étant la première langue non européenne de la section multilingue de ce classement. L'ensemble de données comprend des segments publics et privés, totalisant 4 888 locuteurs, et enregistre 12 attributs des locuteurs, afin de mettre en évidence les erreurs de reconnaissance vocale réparties de manière inégale selon la région, l'âge, le sexe, etc. 🔗 Lire l'article original via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"Industrie","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Le classement Open ASR ajoute la première langue du Sud global : le hindi et le corpus d'évaluation de l'anglais indien - Aioga Actualités IA","description":"Voice Arena collabore avec Hugging Face pour introduire les ensembles de test Monsoon en-IN et Monsoon hi-IN dans le classement Open ASR, couvrant l'anglais indien et l'hindi, ce d...","url":"https://www.aioga.com/fr/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:24:44.507Z"},"de":{"title":"Open ASR-Rangliste fügt erstmals eine globale Sprache des Südens hinzu: Hindi- und indisches Englisch-Bewertungsset","summary":"Voice Arena arbeitet mit Hugging Face zusammen, um für das Open ASR-Ranking die beiden Evaluationsdatensätze Monsoon en-IN und Monsoon hi-IN einzuführen, die Hindi und indisches Englisch abdecken, wobei Hindi die erste nicht-europäische Sprache im mehrsprachigen Abschnitt dieses Rankings ist. Der Datensatz enthält öffentliche und private Segmente mit insgesamt 4.888 Sprechern und erfasst 12 Sprechereigenschaften, um ungleich verteilte Spracherkennungsfehler nach Regionen, Alter, Geschlecht und anderen Dimensionen aufzuzeigen. 🔗 Originaltext lesen via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR-Rangliste fügt erstmals eine globale Sprache des Südens hinzu: Hindi- und indisches Englisch-Bewertungsset - Aioga KI-News","description":"Voice Arena arbeitet mit Hugging Face zusammen, um für das Open ASR-Ranking die beiden Evaluationsdatensätze Monsoon en-IN und Monsoon hi-IN einzuführen, die Hindi und indisches En...","url":"https://www.aioga.com/de/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:24:44.313Z"},"pt-BR":{"title":"O ranking do Open ASR adiciona o primeiro idioma global do Sul: conjuntos de avaliação de hindi e inglês indiano","summary":"A Voice Arena colaborou com a Hugging Face para introduzir os conjuntos de avaliação Monsoon en-IN e Monsoon hi-IN na classificação Open ASR, cobrindo inglês indiano e hindi, sendo o hindi a primeira língua não europeia na seção multilíngue dessa classificação. O conjunto de dados contém divisões públicas e privadas, totalizando 4.888 falantes, e registra 12 atributos dos falantes, com o objetivo de expor erros de reconhecimento de voz com distribuição desigual por região, idade, gênero e outros fatores. 🔗 Leia o texto original via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"O ranking do Open ASR adiciona o primeiro idioma global do Sul: conjuntos de avaliação de hindi e inglês indiano - Aioga Notícias de IA","description":"A Voice Arena colaborou com a Hugging Face para introduzir os conjuntos de avaliação Monsoon en-IN e Monsoon hi-IN na classificação Open ASR, cobrindo inglês indiano e hindi, sendo...","url":"https://www.aioga.com/pt-BR/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:25:34.267Z"},"ru":{"title":"Открыт рейтинг ASR с добавлением первого глобального южного языка: набор тестов по хинди и индийскому английскому","summary":"Voice Arena в сотрудничестве с Hugging Face представила для рейтинга Open ASR два оценочных набора данных Monsoon en-IN и Monsoon hi-IN, охватывающих хинди и индийский английский, при этом хинди является первым неевропейским языком в многоязычном разделе данного рейтинга. Наборы данных содержат публичные и приватные сегменты, всего 4 888 говорящих, а также фиксируют 12 характеристик говорящего, что нацелено на выявление неравномерных ошибок распознавания речи по регионам, возрасту, полу и другим параметрам. 🔗 Читать оригинал via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Открыт рейтинг ASR с добавлением первого глобального южного языка: набор тестов по хинди и индийскому английскому - Aioga Новости ИИ","description":"Voice Arena в сотрудничестве с Hugging Face представила для рейтинга Open ASR два оценочных набора данных Monsoon en-IN и Monsoon hi-IN, охватывающих хинди и индийский английский,...","url":"https://www.aioga.com/ru/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:25:41.455Z"},"ar":{"title":"أُضيف أول لغة من دول الجنوب العالمي إلى قائمة ترتيب ASR: مجموعة تقييم الهندية والإنجليزية الهندية","summary":"تعاونت Voice Arena مع Hugging Face لإدراج مجموعتي التقييم Monsoon en-IN و Monsoon hi-IN في تصنيف Open ASR، لتغطية اللغة الهندية والإنجليزية الهندية، حيث تُعد الهندية أول لغة غير أوروبية في القسم متعدد اللغات لهذا التصنيف. تحتوي مجموعة البيانات على تقسيمات عامة وخاصة، مع 4,888 متحدث، وتسجيل 12 سمة للمتحدثين، بهدف الكشف عن أخطاء التعرف على الكلام غير المتوازنة من حيث المنطقة والعمر والجنس وغيرها من الأبعاد. 🔗 اقرأ النص الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"أُضيف أول لغة من دول الجنوب العالمي إلى قائمة ترتيب ASR: مجموعة تقييم الهندية والإنجليزية الهندية - Aioga أخبار الذكاء الاصطناعي","description":"تعاونت Voice Arena مع Hugging Face لإدراج مجموعتي التقييم Monsoon en-IN و Monsoon hi-IN في تصنيف Open ASR، لتغطية اللغة الهندية والإنجليزية الهندية، حيث تُعد الهندية أول لغة غير أو...","url":"https://www.aioga.com/ar/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:26:34.841Z"},"hi":{"title":"Open ASR रैंकिंग में पहली वैश्विक दक्षिणी भाषा शामिल: हिंदी और भारतीय अंग्रेज़ी मूल्यांकन सेट","summary":"Voice Arena ने Hugging Face के साथ सहयोग करके Open ASR रैंकिंग के लिए Monsoon en-IN और Monsoon hi-IN दो परीक्षण सेट पेश किए हैं, जो हिंदी और भारतीय अंग्रेज़ी को कवर करते हैं, जिसमें हिंदी इस रैंकिंग के बहुभाषी सेक्शन में पहली गैर-यूरोपीय भाषा है। डेटासेट में सार्वजनिक और निजी विभाजन शामिल हैं, कुल 4,888 बोलने वाले हैं, और इसमें 12 बोलने वाले की विशेषताएँ रिकॉर्ड की गई हैं, जिसका उद्देश्य क्षेत्र, उम्र, लिंग आदि के आधार पर असमान वितरण वाले आवाज़ पहचान त्रुटियों को उजागर करना है। 🔗 मूल लेख पढ़ें via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR रैंकिंग में पहली वैश्विक दक्षिणी भाषा शामिल: हिंदी और भारतीय अंग्रेज़ी मूल्यांकन सेट - Aioga AI समाचार","description":"Voice Arena ने Hugging Face के साथ सहयोग करके Open ASR रैंकिंग के लिए Monsoon en-IN और Monsoon hi-IN दो परीक्षण सेट पेश किए हैं, जो हिंदी और भारतीय अंग्रेज़ी को कवर करते हैं, जिसमे...","url":"https://www.aioga.com/hi/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:26:40.743Z"},"it":{"title":"La classifica Open ASR aggiunge la prima lingua globale del Sud: set di valutazione in hindi e inglese indiano","summary":"Voice Arena ha collaborato con Hugging Face per introdurre due set di valutazione, Monsoon en-IN e Monsoon hi-IN, nella classifica Open ASR, coprendo Hindi e inglese indiano, con l'hindi che rappresenta la prima lingua non europea nella sezione multilingue di questa classifica. I dataset contengono suddivisioni pubbliche e private, per un totale di 4.888 parlanti, e registrano 12 caratteristiche dei parlanti, con l'obiettivo di evidenziare errori di riconoscimento vocale distribuiti in modo non uniforme in base a regione, età, genere e altre dimensioni. 🔗 Leggi l'articolo originale via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"La classifica Open ASR aggiunge la prima lingua globale del Sud: set di valutazione in hindi e inglese indiano - Aioga Notizie IA","description":"Voice Arena ha collaborato con Hugging Face per introdurre due set di valutazione, Monsoon en-IN e Monsoon hi-IN, nella classifica Open ASR, coprendo Hindi e inglese indiano, con l...","url":"https://www.aioga.com/it/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:27:38.413Z"},"nl":{"title":"Open ASR-ranglijst voegt de eerste wereldwijde zuidelijke talen toe: Hindi en Indiaas Engels evaluatieset","summary":"Voice Arena werkt samen met Hugging Face om voor de Open ASR-ranglijst twee evaluatiedatasets, Monsoon en-IN en Monsoon hi-IN, te introduceren, die Hindi en Indiaas Engels bestrijken. Hindi is de eerste niet-Europese taal in het meertalige gedeelte van deze ranglijst. De datasets bevatten openbare en privé-secties, in totaal 4.888 sprekers, en registreren 12 sprekersattributen, met als doel spraakherkenningsfouten bloot te leggen die ongelijk verdeeld zijn over regio, leeftijd, geslacht en andere dimensies. 🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR-ranglijst voegt de eerste wereldwijde zuidelijke talen toe: Hindi en Indiaas Engels evaluatieset - Aioga AI-nieuws","description":"Voice Arena werkt samen met Hugging Face om voor de Open ASR-ranglijst twee evaluatiedatasets, Monsoon en-IN en Monsoon hi-IN, te introduceren, die Hindi en Indiaas Engels bestrijk...","url":"https://www.aioga.com/nl/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:27:32.687Z"},"tr":{"title":"Open ASR sıralamasına ilk küresel Güney dili eklendi: Hintçe ve Hindistan İngilizcesi değerlendirme koleksiyonu","summary":"Voice Arena, Hugging Face ile iş birliği yaparak Open ASR sıralaması için Hintçe ve Hindistan İngilizcesini kapsayan Monsoon en-IN ve Monsoon hi-IN olmak üzere iki değerlendirme seti sundu. Bu sırada, Hintçe bu sıralamanın çok dilli bölümündeki ilk Avrupa dışı dil oluyor. Veri seti, açık ve özel bölümlerden oluşmak üzere toplam 4.888 konuşmacıyı kapsıyor ve 12 konuşmacı özelliğini kayıt altına alıyor. Amaç, bölge, yaş, cinsiyet gibi boyutlara göre dengesiz dağılan konuşma tanıma hatalarını ortaya çıkarmaktır. 🔗 Orijinal metni okuyun via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR sıralamasına ilk küresel Güney dili eklendi: Hintçe ve Hindistan İngilizcesi değerlendirme koleksiyonu - Aioga AI Haberleri","description":"Voice Arena, Hugging Face ile iş birliği yaparak Open ASR sıralaması için Hintçe ve Hindistan İngilizcesini kapsayan Monsoon en-IN ve Monsoon hi-IN olmak üzere iki değerlendirme se...","url":"https://www.aioga.com/tr/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:28:40.376Z"},"vi":{"title":"Open ASR bảng xếp hạng thêm ngôn ngữ toàn cầu phía Nam đầu tiên: bộ đánh giá tiếng Hindi và tiếng Anh Ấn Độ","summary":"Voice Arena hợp tác với Hugging Face, giới thiệu hai bộ dữ liệu đánh giá Monsoon en-IN và Monsoon hi-IN cho bảng xếp hạng Open ASR, bao phủ tiếng Hindi và tiếng Anh Ấn Độ, trong đó tiếng Hindi là ngôn ngữ phi châu Âu đầu tiên trong mục đa ngôn ngữ của bảng xếp hạng này. Bộ dữ liệu bao gồm các phân đoạn công khai và riêng tư, tổng cộng 4.888 người nói, và ghi lại 12 thuộc tính người nói, nhằm chỉ ra các lỗi nhận dạng giọng nói phân bố không đều theo khu vực, độ tuổi, giới tính, v.v. 🔗 Đọc bài gốc qua AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR bảng xếp hạng thêm ngôn ngữ toàn cầu phía Nam đầu tiên: bộ đánh giá tiếng Hindi và tiếng Anh Ấn Độ - Tin tức AI Aioga","description":"Voice Arena hợp tác với Hugging Face, giới thiệu hai bộ dữ liệu đánh giá Monsoon en-IN và Monsoon hi-IN cho bảng xếp hạng Open ASR, bao phủ tiếng Hindi và tiếng Anh Ấn Độ, trong đó...","url":"https://www.aioga.com/vi/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:28:38.626Z"},"id":{"title":"Open ASR daftar peringkat menambahkan bahasa Global Selatan pertama: set evaluasi Hindi dan Inggris India","summary":"Voice Arena bekerja sama dengan Hugging Face untuk memperkenalkan dua set evaluasi, Monsoon en-IN dan Monsoon hi-IN, ke dalam peringkat Open ASR, mencakup bahasa Hindi dan Inggris India, di mana bahasa Hindi merupakan bahasa non-Eropa pertama dalam bagian multibahasa peringkat tersebut. Dataset ini mencakup segmen publik dan privat, dengan total 4.888 pembicara, serta mencatat 12 atribut pembicara, bertujuan untuk mengekspos kesalahan pengenalan suara yang tidak merata berdasarkan wilayah, usia, jenis kelamin, dan dimensi lainnya. 🔗 Baca artikel lengkap via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR daftar peringkat menambahkan bahasa Global Selatan pertama: set evaluasi Hindi dan Inggris India - Berita AI Aioga","description":"Voice Arena bekerja sama dengan Hugging Face untuk memperkenalkan dua set evaluasi, Monsoon en-IN dan Monsoon hi-IN, ke dalam peringkat Open ASR, mencakup bahasa Hindi dan Inggris...","url":"https://www.aioga.com/id/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:29:37.156Z"},"th":{"title":"Open ASR อันดับเพิ่มภาษาทางใต้ของโลกเป็นครั้งแรก: ชุดทดสอบภาษาฮินดีและอังกฤษแบบอินเดีย","summary":"Voice Arena ร่วมมือกับ Hugging Face เพื่อนำชุดทดสอบ Monsoon en-IN และ Monsoon hi-IN มาใช้ใน Open ASR Leaderboard ครอบคลุมภาษาอังกฤษของอินเดียและภาษาฮินดี ซึ่งภาษาฮินดีเป็นภาษานอกยุโรปชุดแรกของแผนกหลายภาษาในกระดานผู้นำ ชุดข้อมูลประกอบด้วยส่วนแบ่งแบบสาธารณะและส่วนตัว รวม 4,888 ผู้พูด และบันทึกคุณลักษณะผู้พูด 12 รายการ มีเป้าหมายเพื่อเปิดเผยความผิดพลาดในการรู้จำเสียงตามการกระจายที่ไม่เท่ากันตามภูมิภาค อายุ เพศ ฯลฯ 🔗 อ่านต้นฉบับ via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR อันดับเพิ่มภาษาทางใต้ของโลกเป็นครั้งแรก: ชุดทดสอบภาษาฮินดีและอังกฤษแบบอินเดีย - ข่าว AI Aioga","description":"Voice Arena ร่วมมือกับ Hugging Face เพื่อนำชุดทดสอบ Monsoon en-IN และ Monsoon hi-IN มาใช้ใน Open ASR Leaderboard ครอบคลุมภาษาอังกฤษของอินเดียและภาษาฮินดี ซึ่งภาษาฮินดีเป็นภาษานอกยุ...","url":"https://www.aioga.com/th/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:29:47.546Z"},"pl":{"title":"Open ASR ranking dodaje pierwszy globalny język południowy: zestaw testowy dla hindi i indyjskiego angielskiego","summary":"Voice Arena współpracuje z Hugging Face, wprowadzając do rankingu Open ASR dwa zestawy testowe Monsoon en-IN i Monsoon hi-IN, obejmujące język hindi i angielski indyjski, przy czym hindi jest pierwszym językiem spoza Europy w wielojęzycznej sekcji tego rankingu. Zestawy danych zawierają podziały publiczne i prywatne, obejmujące łącznie 4 888 mówców, oraz rejestrują 12 cech mówców, mające na celu ujawnienie nierównomiernych błędów rozpoznawania mowy według regionu, wieku, płci i innych wymiarów. 🔗 Przeczytaj oryginał via AIHOT · https://aihot.virxact.com/items/cmtdgo8nh04rurobxlzwr101x","category":"行业动态","source":"Hugging Face：Blog（RSS）","aggregationSource":"Hugging Face：Blog（RSS）","pageTitle":"Open ASR ranking dodaje pierwszy globalny język południowy: zestaw testowy dla hindi i indyjskiego angielskiego - Aioga Wiadomości AI","description":"Voice Arena współpracuje z Hugging Face, wprowadzając do rankingu Open ASR dwa zestawy testowe Monsoon en-IN i Monsoon hi-IN, obejmujące język hindi i angielski indyjski, przy czym...","url":"https://www.aioga.com/pl/news/cmtdgo8nh04rurobxlzwr101x/","contentTranslated":true,"sourceHash":"eb34139ad533b03a","translatedAt":"2026-08-29T10:30:58.751Z"}}}}