SANA-Video 2.0 是一个混合视频扩散 Transformer,提供 5B 和 14B 两种规模,可在单 GPU 上生成最高 720p 视频。
其 Hybrid Linear-Softmax Attention 以 3:1 比例混合线性与 softmax 注意力,配合 Block Attention Residuals 将深层有效秩提升约 12%。
SANA-Video 2.0 is a hybrid video diffusion Transformer, available in 5B and 14B sizes, capable of generating videos up to 720p on a single GPU. Its Hybrid...
SANA-Video 2.0 is a hybrid video diffusion Transformer, available in 5B and 14B sizes, capable of
generating videos up to 720p on a single GPU. Its Hybrid Linear-Softmax Attention mixes linear and softmax attention at a 3:1 ratio, and combined with Block Attention Residuals, it increases the effective rank of deep layers by about 12%.
SANA-Video 2.0 是一个混合视频扩散 Transformer,提供 5B 和 14B 两种规模,可在单 GPU 上生成最高 720p 视频。
其 Hybrid Linear-Softmax Attention 以 3:1 比例混合线性与 softmax 注意力,配合 Block Attention Residuals 将深层有效秩提升约 12%。
SANA-Video 2.0 是混合视频扩散 Transformer,包含 5B 与 14B 两种规模,公开材料称其可在单 GPU 上生成最高 720p 视频。
该模型采用 Hybrid Linear-Softmax Attention,按 3:1 比例混合线性与 softmax 注意力,并引入 Block Attention Residuals。
Aioga 判断,方案重点是在视频生成中组合不同注意力机制,并通过残差设计改善深层表示;其实际效率与生成质量仍需更多公开结果支撑。
材料称 Block Attention Residuals 可将深层有效秩提升约 12%。这可能为控制视频扩散 Transformer 的计算负担与深层表示能力提供研究参考。 值得关注后续是否披露 5B 与 14B 模型的生成速度、显存占用、质量评测和测试条件,以验证单 GPU 生成最高 720p 视频的适用范围。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: arxiv.org
Source: HuggingFace Daily Papers(社区热门论文)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: summary-fallback · Updated: 2026-07-23T00:00:00.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.