Black Forest Labs has launched the FLUX 3 multimodal foundation model in Early Access, using a unified architecture to jointly learn images, videos, and au...
模型更新IT之家(RSS)
Today AI Intelligence Brief
Black Forest Labs has launched the FLUX 3 multimodal foundation model in Early Access, using a
unified architecture to jointly learn images, videos, and audio. This model is extended based on the Self-Flow learning framework and can output videos up to 20 seconds long with native audio in a single generation, supporting tasks such as text-to-video, image-to-video, and multi-shot concatenation.
Aioga 判断,统一处理图像、视频与音频是此次更新的核心看点,20 秒单次生成和原生音频有助于减少分段制作环节。不过模型仍处于 Early Access 阶段,其稳定性与实际工作流表现仍值得关注。
影响与后续
该模型可能推动视频生成从单一画面输出转向音画联合生成,并为多镜头内容制作提供新的实现方式。官方人工评测给出了与三款模型的对比胜率,但这些结果仍需结合评测条件和独立测试理解。 后续应关注 Early Access 的开放范围、实际生成质量与任务稳定性,并核验其在不同语言、分辨率、多镜头衔接和音视频续写场景中的表现。机器人行为预测合作的研究进展也值得持续跟踪。
Source and Copyright
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Black Forest Labs has launched the FLUX 3 multimodal foundation model in Early Access, using a unified architecture to jointly learn images, videos, and au...
IT之家(RSS)2026-07-24T06:48:04.000Z
Scan to open this article
Aioga aggregates global AI updates and preserves source information for verification and citation.