{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T07:21:26.498Z","headline":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","description":"Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/","mainEntityOfPage":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/","datePublished":"2026-07-15T21:02:48.000Z","dateModified":"2026-07-15T21:02:48.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english","https://aihot.virxact.com/items/cmrmksstz0419biul93dzfnfw"],"canonicalUrl":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/","dateCreated":"2026-07-15T21:02:48.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english","datePublished":"2026-07-15T21:02:48.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrmksstz0419biul93dzfnfw","datePublished":"2026-07-15T21:02:48.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrmksstz0419biul93dzfnfw"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english"},"article":{"id":"cmrmksstz0419biul93dzfnfw","slug":"cmrmksstz0419biul93dzfnfw","url":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/","title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","title_en":"Soofi Consortium Releases Soofi S 30B-A3B： An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English","summary":"Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english","aiHotUrl":"https://aihot.virxact.com/items/cmrmksstz0419biul93dzfnfw","publishedAt":"2026-07-15T21:02:48.000Z","category":"模型更新","score":63,"selected":false,"articleBody":["A German research consortium has published the pretraining report for Soofi S 30B-A3B：https://www.soofi.info/soofi-s/ . It is an open base model for German and English. Training ran end to end on Deutsche Telekom’s Industrial AI Cloud in Munich. Preview weights are on Hugging Face. It is worth noting that among some of the fully open base models tested, Soofi S records the highest English and German aggregate scores.","Soofi S is a Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model. It totals ~31.6B parameters and activates ~3.2B per token. As a base model, it has no instruction tuning, alignment, or safety tuning. The KI Bundesverband coordinates the consortium, funded by the German Federal Ministry for Economic Affairs and Energy. Participants include Fraunhofer IAIS, DFKI, TU Darmstadt, ellamind, and Merantix Momentum.","The efficiency claim starts with the layer stack. The network holds 52 layers. That is 23 Mamba-2 sequence-mixing layers, 23 granular MoE layers, and 6 Grouped-Query Attention (GQA) layers. Only those 6 GQA layers maintain a KV cache. Each MoE layer holds 128 routed experts, activates 6 per token, and adds 2 shared experts. Other details: model dimension 2688, squared ReLU, RMSNorm, and no positional embeddings.","Soofi S adopts the Nemotron 3 Nano reference design without modification. The research team gives three reasons for that choice. Those are deployability on stacks such as vLLM, serving efficiency, and scientific control. Because the backbone is fixed, Nemotron 3 Nano becomes an architecture-identical baseline. The data recipe is the only moving part.","That recipe follows a Warmup–Stable–Decay (WSD) schedule with a minus_sqrt decay segment. Phase 1 consumed ~20T tokens on a diverse, quality-tiered mixture at a 1e-3 plateau. Phase 2 consumed ~6.58T tokens of high-quality annealing data. It decays 1e-3 to 1e-5, then continues at a constant 1e-5. Phase 3 consumed ~0.10T tokens at a 1,048,576-token sequence length. It extends the usable context window up to 1M tokens.","German is the deliberate variable. It rises from 7.2% of Phase 1 effective tokens to 15.32% in Phase 2. The reference Nemotron 3 Nano mixture allocates about 5% to all non-English languages combined. German sources include HPLT v3 and v4, German Commons, German FinePDFs, and FineWiki. Genios adds 193M articles from 916 newspaper and trade-press archives, commercially licensed.","Infrastructure follows the same sovereignty logic. The run used up to 512 NVIDIA B200 GPUs, from 24 March to 13 May 2026. It consumed ~253,000 B200 GPU-hours.","Those choices show up in the evaluation. Soofi S ran against 16 other open base models. All used the same lm-evaluation-harness pipeline, prompts, and few-shot settings.","Against its architecture-identical reference, Soofi S gains 1.8 points on the English aggregate. German gains 4.2, and held-out English 6.7. That isolates the data recipe from the backbone.","The picture changes against larger open-weight models. Qwen3.5 35B-A3B holds the highest English, German, and held-out means. Soofi S scores 70.1 English against 70.3 for Gemma 3 27B and Ministral 3 14B. On German it leads both, 79.1 to 78.4 and 78.3.","Reproducing any of this starts with the weights. The base repo is a gated preview, and it ships custom modeling code.","Together, the numbers suggest three deployment shapes. First, German document work: GLP-DE 88.8 and INCLUDE-DE 61.2 suit an insurer fine-tuning on policy PDFs. Second, bilingual code assistance: MBPP-DE 84.2 suits teams prompting in German against Python tasks. Third, high-concurrency long-context serving: a support-ticket RAG system at batch 32 and 40K context matches the measured regime. For that case, test retrieval against the RULER and NaturalQuestions gaps.","Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us ：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences."],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2019/06/Screen-Shot-2021-09-14-at-9.02.24-AM-300x300.png","alt":"","afterParagraph":12,"url":"/media/articles/cmrmksstz0419biul93dzfnfw/787a6d54564e8e19.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/high-level-description-a-developer-focus_feCO3rqGV4ig6q7LaBcN2w_K5t5TwjZTPWy666HxX0epA-100x70.png","alt":"Patter SDK Guide to Building a Restaurant Booking Phone Agent with Dynamic Variables, Guardrails, Latency Dashboards, and Eval Checks","afterParagraph":13,"url":"/media/articles/cmrmksstz0419biul93dzfnfw/c998049732b3b311.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/high-level-description-a-technical-blog-_jc7iI1-vWAa65GFgmnBhwQ_U7n28gZXRDeiQEimQU9cAw_cover-100x70.png","alt":"Building a Gin Config Controlled PyTorch Pipeline with Configurable MLP Variants, Cosine Scheduling, and Runtime Parameter Overrides","afterParagraph":13,"url":"/media/articles/cmrmksstz0419biul93dzfnfw/7ac38241502a486f.webp"}],"mediaStatus":"ok","articleBodyZh":["一个德国研究联盟发布了Soofi S 30B-A3B的预训练报告：https：//www.soofi.info/soofi-s/。它是德语和英语的开放式基础模型。培训在慕尼黑的德国电信工业人工智能云端到头进行。预览权重在Hugging Face上。值得注意的是，在测试过的一些完全开放基模型中，Soofi S 在英语和德语中取得了最高的综合分数。","Soofi S 是一种专家混合（MoE）混合型 Mamba 变压器基础模型。它总计有 ~31.6B 参数，每个令牌激活 ~3.2B。作为基础型号，它没有指令调校、校准或安全调校功能。KI联邦联盟协调该联盟，资金来源于德国联邦经济与能源部。参与者包括弗劳恩霍夫IAIS、DFKI、达姆施塔特工业大学、ellamind和Merantix Momentum。","效率的说法从层叠加开始。该网络包含52个层。即23个Mamba-2序列混合层、23个粒度MoE层和6个分组查询注意力（GQA）层。只有这6个GQA层维护KV缓存。每层MoE层可容纳128名被淘汰的专家，每个标记激活6个，并新增2个共享专家。其他细节：模型尺寸2688，ReLU平方，RMSNorm，且无位置嵌入。","Soofi S 采用了 Nemotron 3 Nano 参考设计，未作修改。研究团队给出了三个理由。这些包括可部署性，比如vLLM、服务效率和科学控制。由于骨干是固定的，Nemotron 3 Nano 成为一个架构相同的基线。数据配方是唯一可移动的部分。","该配方遵循预热-稳定-衰变（WSD）程序，并包含minus_sqrt衰变段。第一阶段消耗了约20吨代币，采用多样化且按质量等级分级的混合，处于1e-3平台期。第二阶段消耗了约6.58吨高质量退火数据。它从1e-3衰减到1e-5，然后以恒定的1e-5继续。第三阶段消耗了 ~0.10T 代币，序列长度为 1,048,576 个。它将可用的上下文窗口扩展到最多100万个令牌。","德语是刻意增加的变量。从第一阶段有效令牌的 7.2% 提升到第二阶段的 15.32%。参考 Nemotron 3 Nano 混合模型将约 5% 分配给所有非英语语言。德语来源包括 HPLT v3 和 v4、German Commons、German FinePDFs 以及 FineWiki。Genios 提供 193M 篇文章，来自 916 个报纸和行业报刊档案，拥有商业授权。","基础设施遵循相同的主权逻辑。运行使用了最多 512 个 NVIDIA B200 GPU，从 2026 年 3 月 24 日到 5 月 13 日。消耗约 253,000 B200 GPU 小时。","这些选择在评估中有所体现。Soofi S 与其他 16 个开放基础模型进行了对比。所有模型使用相同的 lm-evaluation-harness 流程、提示和少量示例设置。","与架构完全相同的参考模型相比，Soofi S 在英语综合得分上提升 1.8 分。德语提升 4.2 分，保留英语提升 6.7 分。这将数据配方与主干模型隔离开来。","在与更大的开放权重模型比较时，情况有所不同。Qwen3.5 35B-A3B 拥有最高的英语、德语和保留均值。Soofi S 英语得分 70.1，相比 Gemma 3 27B 和 Ministral 3 14B 分别为 70.3。在德语方面，Soofi S 领先两者，79.1 对 78.4 和 78.3。","复制这些工作首先要从权重开始。基础仓库为受控预览版，并且提供自定义建模代码。","综合来看，这些数字显示了三种部署形态。第一，德语文档处理：GLP-DE 88.8 和 INCLUDE-DE 61.2 适合针对保单 PDF 进行微调的保险公司。第二，双语代码辅助：MBPP-DE 84.2 适合以德语提示处理 Python 任务的团队。第三，高并发长上下文服务：在批量 32 和 40K 上下文下的支持工单 RAG 系统符合测得的配置。针对该场景，需测试 RULER 和 NaturalQuestions 的检索差距。","需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等吗？请联系我们：https://forms.gle/wbash1wF6efRj8G58","Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位有远见的企业家和工程师，Asif 致力于利用人工智能的潜力造福社会。他最近的努力是推出人工智能媒体平台 Marktechpost，该平台因其对机器学习和深度学习新闻的深入报道而脱颖而出，既技术上可靠，又易于广大受众理解。该平台每月浏览量超过 200 万次，显示了其在观众中的受欢迎程度。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE Aioga 将其归入「模型更新」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：模型与研究类动态需要结合能力边界、开放方式、成本、可用性和真实任务表现判断，单项指标领先不等于已经形成稳定采用。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文档、实际可用性、价格变化、开发者反馈和竞品回应。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T07:30:40.678Z","sourceHash":"179a186f569b9e98","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["模型更新","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AI资讯","description":"Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/news/cmrmksstz0419biul93dzfnfw/"},"en":{"title":"Soofi Consortium Releases Soofi S 30B-A3B： An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English","summary":"Aioga tracks this update from MarkTechPost（RSS） under Models. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"Models","source":"MarkTechPost（RSS）","pageTitle":"Soofi Consortium Releases Soofi S 30B-A3B： An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English - Aioga AI News","description":"Aioga tracks this update from MarkTechPost（RSS） under Models. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分","url":"https://www.aioga.com/en/news/cmrmksstz0419biul93dzfnfw/"},"ja":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aiogaは「モデル更新」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"モデル更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AIニュース","description":"Aiogaは「モデル更新」の動きとして、MarkTechPost（RSS） からの更新を追跡しています。Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/ja/news/cmrmksstz0419biul93dzfnfw/"},"ko":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga는 MarkTechPost（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"모델 업데이트","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AI 뉴스","description":"Aioga는 MarkTechPost（RSS）의 업데이트를 모델 업데이트 흐름으로 추적합니다. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/ko/news/cmrmksstz0419biul93dzfnfw/"},"es":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Modelos. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"Modelos","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Noticias de IA","description":"Aioga sigue esta actualización de MarkTechPost（RSS） dentro de Modelos. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（","url":"https://www.aioga.com/es/news/cmrmksstz0419biul93dzfnfw/"},"fr":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Modèles. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"Modèles","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Actualités IA","description":"Aioga suit cette mise à jour de MarkTechPost（RSS） dans la catégorie Modèles. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1","url":"https://www.aioga.com/fr/news/cmrmksstz0419biul93dzfnfw/"},"de":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga KI-News","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/de/news/cmrmksstz0419biul93dzfnfw/"},"pt-BR":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Notícias de IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/pt-BR/news/cmrmksstz0419biul93dzfnfw/"},"ru":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Новости ИИ","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/ru/news/cmrmksstz0419biul93dzfnfw/"},"ar":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/ar/news/cmrmksstz0419biul93dzfnfw/"},"hi":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AI समाचार","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/hi/news/cmrmksstz0419biul93dzfnfw/"},"it":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Notizie IA","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/it/news/cmrmksstz0419biul93dzfnfw/"},"nl":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AI-nieuws","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/nl/news/cmrmksstz0419biul93dzfnfw/"},"tr":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga AI Haberleri","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/tr/news/cmrmksstz0419biul93dzfnfw/"},"vi":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Tin tức AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/vi/news/cmrmksstz0419biul93dzfnfw/"},"id":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Berita AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/id/news/cmrmksstz0419biul93dzfnfw/"},"th":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - ข่าว AI Aioga","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/th/news/cmrmksstz0419biul93dzfnfw/"},"pl":{"title":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型","summary":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","category":"模型更新","source":"MarkTechPost（RSS）","pageTitle":"Soofi 联盟发布 Soofi S 30B-A3B：面向德语和英语的开源混合 Mamba-Transformer MoE 基础模型 - Aioga Wiadomości AI","description":"Aioga tracks this update from MarkTechPost（RSS） under 模型更新. Soofi S 30B-A3B 是一个总参数约 31.6B、每 token 激活约 3.2B 参数的混合 Mamba-Transformer MoE 基础模型，在完全开源的基础模型中取得最高的英语（70.1%）和德语（79.1%）聚合分数。","url":"https://www.aioga.com/pl/news/cmrmksstz0419biul93dzfnfw/"}}}}