一个德国研究联盟发布了Soofi S 30B-A3B的预训练报告:https://www.soofi.info/soofi-s/。它是德语和英语的开放式基础模型。培训在慕尼黑的德国电信工业人工智能云端到头进行。预览权重在Hugging Face上。值得注意的是,在测试过的一些完全开放基模型中,Soofi S 在英语和德语中取得了最高的综合分数。
Soofi S 是一种专家混合(MoE)混合型 Mamba 变压器基础模型。它总计有 ~31.6B 参数,每个令牌激活 ~3.2B。作为基础型号,它没有指令调校、校准或安全调校功能。KI联邦联盟协调该联盟,资金来源于德国联邦经济与能源部。参与者包括弗劳恩霍夫IAIS、DFKI、达姆施塔特工业大学、ellamind和Merantix Momentum。
需要与我们合作推广您的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会等吗?请联系我们:https://forms.gle/wbash1wF6efRj8G58
Asif Razzaq 是 Marktechpost Media Inc. 的首席执行官。作为一位有远见的企业家和工程师,Asif 致力于利用人工智能的潜力造福社会。他最近的努力是推出人工智能媒体平台 Marktechpost,该平台因其对机器学习和深度学习新闻的深入报道而脱颖而出,既技术上可靠,又易于广大受众理解。该平台每月浏览量超过 200 万次,显示了其在观众中的受欢迎程度。
A German research consortium has published the pretraining report for Soofi S 30B-A3B:https://www.soofi.info/soofi-s/ . It is an open base model for German and English. Training ran end to end on Deutsche Telekom’s Industrial AI Cloud in Munich. Preview weights are on Hugging Face. It is worth noting that among some of the fully open base models tested, Soofi S records the highest English and German aggregate scores.
Soofi S is a Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model. It totals ~31.6B parameters and activates ~3.2B per token. As a base model, it has no instruction tuning, alignment, or safety tuning. The KI Bundesverband coordinates the consortium, funded by the German Federal Ministry for Economic Affairs and Energy. Participants include Fraunhofer IAIS, DFKI, TU Darmstadt, ellamind, and Merantix Momentum.
The efficiency claim starts with the layer stack. The network holds 52 layers. That is 23 Mamba-2 sequence-mixing layers, 23 granular MoE layers, and 6 Grouped-Query Attention (GQA) layers. Only those 6 GQA layers maintain a KV cache. Each MoE layer holds 128 routed experts, activates 6 per token, and adds 2 shared experts. Other details: model dimension 2688, squared ReLU, RMSNorm, and no positional embeddings.
Soofi S adopts the Nemotron 3 Nano reference design without modification. The research team gives three reasons for that choice. Those are deployability on stacks such as vLLM, serving efficiency, and scientific control. Because the backbone is fixed, Nemotron 3 Nano becomes an architecture-identical baseline. The data recipe is the only moving part.
That recipe follows a Warmup–Stable–Decay (WSD) schedule with a minus_sqrt decay segment. Phase 1 consumed ~20T tokens on a diverse, quality-tiered mixture at a 1e-3 plateau. Phase 2 consumed ~6.58T tokens of high-quality annealing data. It decays 1e-3 to 1e-5, then continues at a constant 1e-5. Phase 3 consumed ~0.10T tokens at a 1,048,576-token sequence length. It extends the usable context window up to 1M tokens.
German is the deliberate variable. It rises from 7.2% of Phase 1 effective tokens to 15.32% in Phase 2. The reference Nemotron 3 Nano mixture allocates about 5% to all non-English languages combined. German sources include HPLT v3 and v4, German Commons, German FinePDFs, and FineWiki. Genios adds 193M articles from 916 newspaper and trade-press archives, commercially licensed.
Infrastructure follows the same sovereignty logic. The run used up to 512 NVIDIA B200 GPUs, from 24 March to 13 May 2026. It consumed ~253,000 B200 GPU-hours.
Those choices show up in the evaluation. Soofi S ran against 16 other open base models. All used the same lm-evaluation-harness pipeline, prompts, and few-shot settings.
Against its architecture-identical reference, Soofi S gains 1.8 points on the English aggregate. German gains 4.2, and held-out English 6.7. That isolates the data recipe from the backbone.
The picture changes against larger open-weight models. Qwen3.5 35B-A3B holds the highest English, German, and held-out means. Soofi S scores 70.1 English against 70.3 for Gemma 3 27B and Ministral 3 14B. On German it leads both, 79.1 to 78.4 and 78.3.
Reproducing any of this starts with the weights. The base repo is a gated preview, and it ships custom modeling code.
Together, the numbers suggest three deployment shapes. First, German document work: GLP-DE 88.8 and INCLUDE-DE 61.2 suit an insurer fine-tuning on policy PDFs. Second, bilingual code assistance: MBPP-DE 84.2 suits teams prompting in German against Python tasks. Third, high-concurrency long-context serving: a support-ticket RAG system at batch 32 and 40K context matches the measured regime. For that case, test retrieval against the RULER and NaturalQuestions gaps.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us :https://forms.gle/wbash1wF6efRj8G58
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.