蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-flash,总参数量124B,激活参数量仅5.1B,在传统推理、指令遵循与长文本等指标上对标甚至超越上一代旗舰Ring-2.6-1T。
模型采用原生混合线性注意力架构与1/64稀疏MoE,并扩展至10,000+可交互训练环境,长输入下TTFT降低60%至80%以上。
Ant Group's BaLing has released the new generation native hybrid reasoning model Ling-3.0-flash, with a total parameter count of 124B and only 5.1B active...
Ant Group's BaLing has released the new generation native hybrid reasoning model Ling-3.0-flash,
with a total parameter count of 124B and only 5.1B active parameters. In traditional reasoning, instruction following, and long-text benchmarks, it matches or even surpasses the previous flagship Ring-2.6-1T. The model adopts a native hybrid linear attention architecture with 1/64 sparse MoE and has been expanded to over 10,000 interactive training environments, reducing TTFT under long inputs by more than 60% to 80%.
蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-flash,总参数量124B,激活参数量仅5.1B,在传统推理、指令遵循与长文本等指标上对标甚至超越上一代旗舰Ring-2.6-1T。
模型采用原生混合线性注意力架构与1/64稀疏MoE,并扩展至10,000+可交互训练环境,长输入下TTFT降低60%至80%以上。
蚂蚁百灵发布原生混合推理模型Ling-3.0-flash,总参数量124B、激活参数量5.1B,并称其在传统推理、指令遵循和长文本等指标上对标甚至超越Ring-2.6-1T。
公开材料显示,该模型采用原生混合线性注意力架构与1/64稀疏MoE,训练范围扩展至超过10,000个可交互环境;在长输入场景下,TTFT降低60%至80%以上。
Aioga 判断,Ling-3.0-flash的重点是以较低激活参数量兼顾推理、指令遵循和长文本表现,但现有材料未提供具体评测集、测试条件及第三方验证结果。
Aioga 判断,混合线性注意力与稀疏MoE可能有助于改善长输入处理效率;不过,官方披露的指标能否稳定复现,以及不同任务下的实际表现,仍值得关注。 后续可重点关注完整技术报告、评测方法、基准成绩与测试环境,并核验TTFT降幅的对照模型、输入长度及部署条件,以判断相关性能表述的适用范围。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: mp.weixin.qq.com
Source: 公众号:蚂蚁百灵(Ling)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: summary-fallback · Updated: 2026-07-24T13:40:30.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.