腾讯混元开源端到端投机解码框架 AngelSpec,支持训练与部署。
在 Hy3-A21B 模型上,其 DFly 方案相比自回归解码实现 1.98-2.40 倍端到端加速,吞吐量比 DFlash 高 10.5-11.8%。
训练代码及 Hy3-A21B MTP/DFly 草稿模型权重已开源。
Tencent's Hunyuan open-source end-to-end speculative decoding framework AngelSpec supports both training and deployment. On the Hy3-A21B model, its DFly so...
Tencent's Hunyuan open-source end-to-end speculative decoding framework AngelSpec supports both
training and deployment. On the Hy3-A21B model, its DFly solution achieves 1.98-2.40 times end-to-end acceleration compared to autoregressive decoding, and the throughput is 10.5-11.8% higher than DFlash. The training code and Hy3-A21B MTP/DFly draft model weights have been open-sourced.
腾讯混元开源端到端投机解码框架 AngelSpec,支持训练与部署。
在 Hy3-A21B 模型上,其 DFly 方案相比自回归解码实现 1.98-2.40 倍端到端加速,吞吐量比 DFlash 高 10.5-11.8%。
训练代码及 Hy3-A21B MTP/DFly 草稿模型权重已开源。
腾讯混元宣布开源端到端投机解码框架 AngelSpec,覆盖训练与部署。摘要称,DFly 在 Hy3-A21B 模型上较自回归解码实现 1.98—2.40 倍端到端加速,吞吐量较 DFlash 高 10.5%—11.8%。
公开材料来自腾讯混元官方 X 账号。除 AngelSpec 框架外,腾讯混元还表示已开源训练代码,以及面向 Hy3-A21B 的 MTP、DFly 草稿模型权重;材料未披露具体测试环境与配置。
Aioga 判断,此次发布的重点是将投机解码的训练、部署和配套草稿模型权重一并开放,而非仅公布性能结果。现有加速与吞吐数据仅对应材料所述 Hy3-A21B 场景,不宜直接外推。
该框架可能为研究者和开发者提供复现、训练及部署投机解码方案的公开基础。值得关注的是,材料未说明其他模型、硬件或负载下的表现,因此其通用性能收益仍需更多公开测试验证。 建议后续核验开源训练代码与权重的许可、运行依赖和复现说明,并关注不同硬件、序列长度及并发条件下的测试结果。同时应比较自回归解码与 DFlash 的统一测试配置,以判断数据可比性。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: x.com
Source: X:腾讯混元 (@TencentHunyuan)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: social-summary · Updated: 2026-07-29T12:43:55.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.