{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-28T06:20:51.496Z","headline":"Black Forest Labs 发布 FLUX 3：统一图像、视频、音频与机器人动作预测的多模态基础模型","description":"Black Forest Labs 发布 FLUX 3，这是首个在单一架构中联合学习图像、视频和音频，并从同一组权重输出视频、音频和动作预测的多模态基础模型。FLUX 3 Video 可单次生成长达 20 秒、带原生音频的视频，在 10 秒 720p 文本生成视频的人类偏好测试中，以 93% 的偏好率击败 Luma Ray 3.2。","url":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/","mainEntityOfPage":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/","datePublished":"2026-07-26T17:50:23.000Z","dateModified":"2026-07-26T17:50:23.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.marktechpost.com/2026/07/26/black-forest-labs-releases-flux-3-a-multimodal-flow-model-for-image-video-audio-and-robot-action-prediction","https://aihot.virxact.com/items/cms23ou2z022fro9fuxwjdsli"],"canonicalUrl":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/","directAnswer":{"@type":"Answer","text":"Black Forest Labs 发布 FLUX 3，称其以单一架构联合学习图像、视频与音频，并由同一组权重进行视频、音频和动作预测。FLUX 3 Video 支持单次生成最长20秒且带原生音频的视频。","url":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/","dateCreated":"2026-07-26T17:50:23.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"marktechpost.com source article","url":"https://www.marktechpost.com/2026/07/26/black-forest-labs-releases-flux-3-a-multimodal-flow-model-for-image-video-audio-and-robot-action-prediction","datePublished":"2026-07-26T17:50:23.000Z","provider":{"@type":"Organization","name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/26/black-forest-labs-releases-flux-3-a-multimodal-flow-model-for-image-video-audio-and-robot-action-prediction"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cms23ou2z022fro9fuxwjdsli","datePublished":"2026-07-26T17:50:23.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cms23ou2z022fro9fuxwjdsli"}}],"aggregationSource":"MarkTechPost（RSS）","originalPublisher":{"name":"marktechpost.com","url":"https://www.marktechpost.com/2026/07/26/black-forest-labs-releases-flux-3-a-multimodal-flow-model-for-image-video-audio-and-robot-action-prediction"},"article":{"id":"cms23ou2z022fro9fuxwjdsli","slug":"cms23ou2z022fro9fuxwjdsli","url":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/","title":"Black Forest Labs 发布 FLUX 3：统一图像、视频、音频与机器人动作预测的多模态基础模型","title_en":"Black Forest Labs Releases FLUX 3： A Multimodal Flow Model for Image， Video， Audio and Robot Action Prediction","summary":"Black Forest Labs 发布 FLUX 3，这是首个在单一架构中联合学习图像、视频和音频，并从同一组权重输出视频、音频和动作预测的多模态基础模型。FLUX 3 Video 可单次生成长达 20 秒、带原生音频的视频，在 10 秒 720p 文本生成视频的人类偏好测试中，以 93% 的偏好率击败 Luma Ray 3.2。","source":"MarkTechPost（RSS）","sourceUrl":"https://www.marktechpost.com/2026/07/26/black-forest-labs-releases-flux-3-a-multimodal-flow-model-for-image-video-audio-and-robot-action-prediction","aiHotUrl":"https://aihot.virxact.com/items/cms23ou2z022fro9fuxwjdsli","publishedAt":"2026-07-26T17:50:23.000Z","category":"模型更新","score":70,"selected":false,"articleBody":["Black Forest Labs (BFL) has released FLUX 3：https://bfl.ai/blog/flux-3, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model to ship video, audio and action prediction from one set of weights.","The Black Forest Labs (BFL) research team argues that no single modality gives a complete description of the world. Images capture spatial structure at one instant. Video restores time and exposes physical dynamics. Audio reveals causal relationships between mechanical events and sound. Each is treated as a lossy projection of the same underlying reality.","Training on all of them at once means the modalities constrain each other. The sound has to match the impact. The motion has to obey the mass. The research team calls FLUX 3 its first model built entirely on that principle.","FLUX 3 builds on Self-Flow：https://bfl.ai/research/self-flow, BFL’s method for aligning multimodal generation and understanding in one architecture. Self-Flow combines the flow matching objective with a self-supervised feature reconstruction objective. The reference implementation on GitHub：https://github.com/black-forest-labs/Self-Flow is Apache-2.0 and uses SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token mask ratio and self-distillation from an EMA teacher at layer 20 to a student at layer 8.","That released checkpoint is an ImageNet 256×256 research model, not FLUX 3. BFL states that it ‘significantly scaled up compute and data resources’ on the same approach to train FLUX 3 across video, images and audio simultaneously. Self-Flow itself was introduced in March 2026, so it is not new to this launch. What is new is the scale.","FLUX 3 Video generates clips up to 20 seconds long in a single generation, with native audio. The supported modes cover text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for controlled transitions, and generative video-audio continuation from input video and audio.","BFL also lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and strong typography generation with animated designs. The BFL team reports particular strength in human facial expressions and in associating sounds with physical events.","BFL team published preliminary human preference results. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. Against Grok Imagine Video the figure is up to 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result is 52%, close to a coin flip.","Check out the FLUX 3 announcement：https://bfl.ai/blog/flux-3, the FLUX 3 x mimic technical post：https://bfl.ai/blog/flux-3-mimic and the Self-Flow paper：https://arxiv.org/abs/2603.06507. All credit for this research goes to the researchers of this project.","Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.","Build an Agentic Event Venue Operator [Full Codes]：https://pxllnk.co/twdn5","Thanks! Our team will contact you soon 🙌"],"articleImages":[{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2025/07/a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg-300x300.png","alt":"","afterParagraph":8,"url":"/media/articles/cms23ou2z022fro9fuxwjdsli/84e64b03066de40c.webp"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog6171-6-100x70.png","alt":"KwaiKAT Team Releases KAT-Coder-V2.5","afterParagraph":9,"url":"/media/articles/cms23ou2z022fro9fuxwjdsli/2225a1a0a768ea5a.png"},{"sourceUrl":"https://www.marktechpost.com/wp-content/uploads/2026/07/blog6171-4-100x70.png","alt":"FAIRChem v2 UMA for Multidomain Atomistic Simulation across Molecules, Catalysts, Materials, Vibrations, and Molecular Dynamics","afterParagraph":9,"url":"/media/articles/cms23ou2z022fro9fuxwjdsli/a8de8cc0d3522061.png"}],"mediaStatus":"ok","articleBodyZh":["黑森林实验室（BFL）发布了 FLUX 3：https://bfl.ai/blog/flux-3，这是一种多模态基础模型，在单一架构中从图像、视频和音频中学习。它也是首个可通过一组权重实现视频、音频和动作预测的 FLUX 模型。","黑森林实验室（BFL）的研究团队认为，没有任何单一模态能完整描述世界。图像在瞬间捕捉空间结构。视频还原时间并揭示物理动态。音频揭示机械事件与声音之间的因果关系。每种模态都被视为对同一底层现实的有损投影。","同时在所有模态上进行训练意味着模态之间相互约束。声音必须与撞击匹配。运动必须遵循质量规律。研究团队称 FLUX 3 是其完全基于这一原则构建的首个模型。","FLUX 3 建立在 Self-Flow 基础上：https://bfl.ai/research/self-flow，这是 BFL 在单一架构中对齐多模态生成和理解的方法。Self-Flow 将流匹配目标与自监督特征重建目标结合起来。GitHub 上的参考实现：https://github.com/black-forest-labs/Self-Flow 使用 Apache-2.0 许可证，并采用 SiT-XL/2，每个 token 配置时间步条件。训练采用每 token 25% 的掩码比率，并从第 20 层 EMA 教师到第 8 层学生进行自蒸馏。","已发布的检查点是 ImageNet 256×256 的研究模型，并不是 FLUX 3。BFL 表示，他们在同一方法上‘显著扩大了计算和数据资源’，以同时在视频、图像和音频上训练 FLUX 3。Self-Flow 本身是在 2026 年 3 月推出的，因此对于此次发布来说并不新颖。新颖之处在于规模。","FLUX 3 视频可在一次生成中制作最长 20 秒的片段，并附带原生音频。支持的模式包括文本到视频、图像到视频、参考片段的视频到视频、关键帧到视频以实现受控过渡，以及基于输入视频和音频的生成型视频-音频延续。","BFL 还列出了多语言对话、将片段串联成多镜头序列的代理性链式操作，以及带有动画设计的强大排版生成。BFL 团队报告在人物面部表情以及将声音与物理事件关联方面表现尤为突出。","BFL 团队发布了初步的人类偏好结果。实验设置为 10 秒的文字转视频片段，分辨率为 720p 并带音频。在比较中，FLUX 3 在 93% 的情况下被偏好于 Luma Ray 3.2，在 77% 的情况下被偏好于 Runway Gen-4.5。与 Grok Imagine Video 对比时，这一比例高达 69%，然后是 Kling v3 Pro 为 60%，Happy Horse v1 为 59%，Happy Horse 1.1 为 57%。与 Seedance 2.0 和 Gemini Omni Flash 比较时，结果为 52%，接近抛硬币般的随机。","查看 FLUX 3 公告：https://bfl.ai/blog/flux-3，FLUX 3 x mimic 技术文章：https://bfl.ai/blog/flux-3-mimic 以及 Self-Flow 论文：https://arxiv.org/abs/2603.06507。所有研究成果均归功于该项目的研究人员。","Michal Sutter 是一名数据科学专业人士，拥有帕多瓦大学的数据科学理学硕士学位。凭借在统计分析、机器学习和数据工程方面的坚实基础，Michal 擅长将复杂数据集转化为可操作的洞见。","构建一个自主事件场地运营者 [完整代码]：https://pxllnk.co/twdn5","谢谢！我们的团队会尽快联系您 🙌"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Black Forest Labs 发布 FLUX 3，称其以单一架构联合学习图像、视频与音频，并由同一组权重进行视频、音频和动作预测。FLUX 3 Video 支持单次生成最长20秒且带原生音频的视频。","background":"FLUX 3 基于 BFL 的 Self-Flow 方法，将流匹配目标与自监督特征重建目标结合。公开的 Self-Flow 检查点只是 ImageNet 256×256研究模型，并非 FLUX 3；BFL 称此次主要变化是扩大计算与数据资源。","viewpoint":"Aioga 判断，FLUX 3 的核心看点不是单项生成能力，而是以统一权重覆盖多种模态及动作预测。其让声音、运动与视觉信息相互约束的设计思路值得关注，但实际一致性仍需更多公开测试验证。","implications":"该模型覆盖文本生成视频、图生视频、视频转换、关键帧过渡及视频音频续写等模式。Aioga 判断，这可能为多镜头内容制作和物理事件声画匹配提供更统一的工作路径，但材料未说明实际部署成本。","nextStep":"后续应关注 BFL 是否披露完整评测方法、样本规模及更多第三方测试。现有结果由团队发布，显示其在10秒、720p、带音频测试中对 Luma Ray 3.2 获得93%偏好率，不宜直接外推至全部场景。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-07-27T04:10:53.969Z","sourceHash":"b4c99c3111d1b0cc","review":{"approved":true,"groundedness":96,"clarity":92,"duplicationRisk":18,"blockingIssues":[],"notes":["候选内容准确区分了已公开的 Self-Flow 研究检查点与 FLUX 3，避免将两者混同。","关于统一工作路径、实际一致性和后续评测需求的内容均明确标注为判断或审慎结论，没有冒充来源事实。","对 93% 偏好率注明了测试条件、结果发布方及不可外推性，表述审慎。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["模型更新","MarkTechPost（RSS）"],"translations":{"zh-CN":{"title":"Black Forest Labs 发布 FLUX 3：统一图像、视频、音频与机器人动作预测的多模态基础模型","summary":"Black Forest Labs 发布 FLUX 3，这是首个在单一架构中联合学习图像、视频和音频，并从同一组权重输出视频、音频和动作预测的多模态基础模型。FLUX 3 Video 可单次生成长达 20 秒、带原生音频的视频，在 10 秒 720p 文本生成视频的人类偏好测试中，以 93% 的偏好率击败 Luma Ray 3.2。","category":"模型更新","source":"marktechpost.com","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs 发布 FLUX 3：统一图像、视频、音频与机器人动作预测的多模态基础模型 - Aioga AI资讯","description":"Black Forest Labs 发布 FLUX 3，这是首个在单一架构中联合学习图像、视频和音频，并从同一组权重输出视频、音频和动作预测的多模态基础模型。FLUX 3 Video 可单次生成长达 20 秒、带原生音频的视频，在 10 秒 720p 文本生成视频的人类偏好测试中，以 93% 的偏好率击败 Luma Ray 3.2。","url":"https://www.aioga.com/news/cms23ou2z022fro9fuxwjdsli/"},"en":{"title":"Black Forest Labs releases FLUX 3: a multimodal foundation model unifying image, video, audio, and robotic motion prediction","summary":"Black Forest Labs has released FLUX 3, the first multimodal foundation model that jointly learns images, video, and audio within a single architecture and outputs video, audio, and action predictions from the same set of weights. FLUX 3 Video can generate videos up to 20 seconds long with native audio in a single shot, and in a human preference test for 10-second 720p text-to-video generation, it outperformed Luma Ray 3.2 with a 93% preference rate.","category":"Models","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs releases FLUX 3: a multimodal foundation model unifying image, video, audio, and robotic motion prediction - Aioga AI News","description":"Black Forest Labs has released FLUX 3, the first multimodal foundation model that jointly learns images, video, and audio within a single architecture and outputs video, audio, and...","url":"https://www.aioga.com/en/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:42:10.635Z"},"ja":{"title":"Black Forest Labs が FLUX 3 を発表：画像、動画、音声およびロボット動作予測を統合したマルチモーダル基盤モデル","summary":"Black Forest Labs は FLUX 3 を発表しました。これは、単一のアーキテクチャで画像、ビデオ、オーディオを共同学習し、同じ重みセットからビデオ、オーディオ、動作予測を出力する初のマルチモーダル基盤モデルです。FLUX 3 Video は、最長 20 秒のネイティブオーディオ付きビデオを単発で生成でき、10 秒 720p のテキスト生成ビデオにおける人間の好みテストでは、93% の支持率で Luma Ray 3.2 を打ち負かしました。","category":"モデル更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs が FLUX 3 を発表：画像、動画、音声およびロボット動作予測を統合したマルチモーダル基盤モデル - Aioga AIニュース","description":"Black Forest Labs は FLUX 3 を発表しました。これは、単一のアーキテクチャで画像、ビデオ、オーディオを共同学習し、同じ重みセットからビデオ、オーディオ、動作予測を出力する初のマルチモーダル基盤モデルです。FLUX 3 Video は、最長 20 秒のネイティブオーディオ付きビデオを単発で生成でき、10 秒 720p のテキスト生成ビデ...","url":"https://www.aioga.com/ja/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:42:21.358Z"},"ko":{"title":"Black Forest Labs, FLUX 3 출시: 이미지, 비디오, 오디오 및 로봇 동작 예측을 통합한 멀티모달 기초 모델","summary":"Black Forest Labs는 FLUX 3를 발표했습니다. 이는 단일 아키텍처에서 이미지, 비디오 및 오디오를 공동 학습하고 동일한 가중치 세트에서 비디오, 오디오 및 동작 예측을 출력하는 최초의 다중 모달 기반 모델입니다. FLUX 3 Video는 최대 20초 길이의 원본 오디오가 포함된 비디오를 단번에 생성할 수 있으며, 10초 720p 텍스트 생성 비디오 인간 선호도 테스트에서 93%의 선호도로 Luma Ray 3.2를 능가했습니다.","category":"모델 업데이트","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs, FLUX 3 출시: 이미지, 비디오, 오디오 및 로봇 동작 예측을 통합한 멀티모달 기초 모델 - Aioga AI 뉴스","description":"Black Forest Labs는 FLUX 3를 발표했습니다. 이는 단일 아키텍처에서 이미지, 비디오 및 오디오를 공동 학습하고 동일한 가중치 세트에서 비디오, 오디오 및 동작 예측을 출력하는 최초의 다중 모달 기반 모델입니다. FLUX 3 Video는 최대 20초 길이의 원본 오디오가 포함된 비디오를 단번에 생성할 수...","url":"https://www.aioga.com/ko/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:43:06.370Z"},"es":{"title":"Black Forest Labs lanza FLUX 3: un modelo base multimodal unificado para la predicción de imágenes, videos, audio y acciones de robots","summary":"Black Forest Labs lanzó FLUX 3, el primer modelo base multimodal que aprende imágenes, videos y audio dentro de una única arquitectura, y genera predicciones de video, audio y acciones desde el mismo conjunto de pesos. FLUX 3 Video puede generar videos de hasta 20 segundos con audio nativo en una sola pasada, y en pruebas de preferencia humana para videos generados a partir de texto de 10 segundos a 720p, superó a Luma Ray 3.2 con una tasa de preferencia del 93%.","category":"Modelos","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs lanza FLUX 3: un modelo base multimodal unificado para la predicción de imágenes, videos, audio y acciones de robots - Aioga Noticias de IA","description":"Black Forest Labs lanzó FLUX 3, el primer modelo base multimodal que aprende imágenes, videos y audio dentro de una única arquitectura, y genera predicciones de video, audio y acci...","url":"https://www.aioga.com/es/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:43:00.147Z"},"fr":{"title":"Black Forest Labs publie FLUX 3 : un modèle fondamental multimodal unifiant la prédiction d'images, de vidéos, d'audio et de mouvements robotiques","summary":"Black Forest Labs a publié FLUX 3, le premier modèle de base multimodal qui apprend conjointement l'image, la vidéo et l'audio dans une seule architecture, et produit à partir du même ensemble de poids des prédictions vidéo, audio et d'action. FLUX 3 Video peut générer en une seule fois des vidéos jusqu'à 20 secondes avec audio natif, et dans un test de préférence humaine de vidéos générées par texte de 10 secondes en 720p, il a dépassé Luma Ray 3.2 avec un taux de préférence de 93%.","category":"Modèles","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs publie FLUX 3 : un modèle fondamental multimodal unifiant la prédiction d'images, de vidéos, d'audio et de mouvements robotiques - Aioga Actualités IA","description":"Black Forest Labs a publié FLUX 3, le premier modèle de base multimodal qui apprend conjointement l'image, la vidéo et l'audio dans une seule architecture, et produit à partir du m...","url":"https://www.aioga.com/fr/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:43:48.290Z"},"de":{"title":"Black Forest Labs veröffentlicht FLUX 3: Ein multimodales Grundmodell zur einheitlichen Vorhersage von Bildern, Videos, Audio und Roboterbewegungen","summary":"Black Forest Labs hat FLUX 3 veröffentlicht, das erste multimodale Basis-Modell, das in einer einzigen Architektur Bild, Video und Audio gemeinsam lernt und aus demselben Satz von Gewichtungen Video-, Audio- und Bewegungsprognosen ausgibt. FLUX 3 Video kann auf einmal Videos von bis zu 20 Sekunden mit nativer Audioausgabe erzeugen und hat in einem menschlichen Vorzugstest für 10-Sekunden-720p-Textgenerierungsvideos mit einer Vorzugsrate von 93 % Luma Ray 3.2 übertroffen.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs veröffentlicht FLUX 3: Ein multimodales Grundmodell zur einheitlichen Vorhersage von Bildern, Videos, Audio und Roboterbewegungen - Aioga KI-News","description":"Black Forest Labs hat FLUX 3 veröffentlicht, das erste multimodale Basis-Modell, das in einer einzigen Architektur Bild, Video und Audio gemeinsam lernt und aus demselben Satz von...","url":"https://www.aioga.com/de/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:43:51.163Z"},"pt-BR":{"title":"Black Forest Labs lança FLUX 3: modelo fundamental multimodal unificado para previsão de imagens, vídeos, áudios e ações de robôs","summary":"A Black Forest Labs lançou o FLUX 3, o primeiro modelo fundamental multimodal que aprende imagens, vídeos e áudios em uma única arquitetura, e gera previsões de vídeo, áudio e ações a partir do mesmo conjunto de pesos. O FLUX 3 Video pode gerar vídeos de até 20 segundos com áudio nativo em uma única vez, e em testes de preferência humana de vídeos gerados por texto de 10 segundos em 720p, superou o Luma Ray 3.2 com uma taxa de preferência de 93%.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs lança FLUX 3: modelo fundamental multimodal unificado para previsão de imagens, vídeos, áudios e ações de robôs - Aioga Notícias de IA","description":"A Black Forest Labs lançou o FLUX 3, o primeiro modelo fundamental multimodal que aprende imagens, vídeos e áudios em uma única arquitetura, e gera previsões de vídeo, áudio e açõe...","url":"https://www.aioga.com/pt-BR/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:44:36.626Z"},"ru":{"title":"Black Forest Labs выпустила FLUX 3: унифицированная мультимодальная базовая модель для предсказания изображений, видео, аудио и действий роботов","summary":"Black Forest Labs выпустила FLUX 3, первую мультимодальную базовую модель, которая в единой архитектуре обучает изображения, видео и аудио, и с одного и того же набора весов выводит прогнозы видео, аудио и действий. FLUX 3 Video может за один раз создавать видео длительностью до 20 секунд с оригинальным аудио, и в тесте человеческих предпочтений для 10-секундного видео 720p, сгенерированного по тексту, одержала победу над Luma Ray 3.2 с предпочтением в 93% случаев.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs выпустила FLUX 3: унифицированная мультимодальная базовая модель для предсказания изображений, видео, аудио и действий роботов - Aioga Новости ИИ","description":"Black Forest Labs выпустила FLUX 3, первую мультимодальную базовую модель, которая в единой архитектуре обучает изображения, видео и аудио, и с одного и того же набора весов выводи...","url":"https://www.aioga.com/ru/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:44:34.334Z"},"ar":{"title":"أصدرت مختبرات بلاك فورست FLUX 3: نموذج أساسي متعدد الوسائط يوحد التنبؤ بالصور والفيديو والصوت وحركة الروبوتات","summary":"أصدرت Black Forest Labs نموذج FLUX 3، وهو أول نموذج أساسي متعدد الوسائط يتعلم الصور ومقاطع الفيديو والصوت ضمن هيكل واحد، ويخرج توقعات الفيديو والصوت والحركة من نفس مجموعة الأوزان. يمكن لـ FLUX 3 Video إنتاج مقاطع فيديو تصل مدتها إلى 20 ثانية بصوت أصلي في مرة واحدة، وفي اختبار تفضيل الإنسان لمقاطع الفيديو التي تم إنشاؤها نصيًا بدقة 720 بكسل لمدة 10 ثوانٍ، تفوق على Luma Ray 3.2 بنسبة تفضيل بلغت 93%.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"أصدرت مختبرات بلاك فورست FLUX 3: نموذج أساسي متعدد الوسائط يوحد التنبؤ بالصور والفيديو والصوت وحركة الروبوتات - Aioga أخبار الذكاء الاصطناعي","description":"أصدرت Black Forest Labs نموذج FLUX 3، وهو أول نموذج أساسي متعدد الوسائط يتعلم الصور ومقاطع الفيديو والصوت ضمن هيكل واحد، ويخرج توقعات الفيديو والصوت والحركة من نفس مجموعة الأوزان....","url":"https://www.aioga.com/ar/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:45:24.872Z"},"hi":{"title":"ब्लैक फॉरेस्ट लैब्स ने FLUX 3 जारी किया: इमेज, वीडियो, ऑडियो और रोबोटिक मूवमेंट प्रिडिक्शन को एकीकृत करने वाला मल्टीमॉडल बेस मॉडल","summary":"ब्लैक फॉरेस्ट लैब्स ने FLUX 3 जारी किया, जो एक एकल आर्किटेक्चर में छवि, वीडियो और ऑडियो को संयुक्त रूप से सीखने और एक ही सेट के वजनों से वीडियो, ऑडियो और क्रिया पूर्वानुमान देने वाला पहला बहु-मोडल बेस मॉडल है। FLUX 3 वीडियो एक बार में 20 सेकंड तक की वीडियो जनरेट कर सकता है, जिसमें मूल ऑडियो शामिल है। 10 सेकंड 720p टेक्स्ट-जनरेट वीडियो के मानवीय प्राथमिकता परीक्षण में, इसने Luma Ray 3.2 को 93% प्राथमिकता दर के साथ हराया।","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"ब्लैक फॉरेस्ट लैब्स ने FLUX 3 जारी किया: इमेज, वीडियो, ऑडियो और रोबोटिक मूवमेंट प्रिडिक्शन को एकीकृत करने वाला मल्टीमॉडल बेस मॉडल - Aioga AI समाचार","description":"ब्लैक फॉरेस्ट लैब्स ने FLUX 3 जारी किया, जो एक एकल आर्किटेक्चर में छवि, वीडियो और ऑडियो को संयुक्त रूप से सीखने और एक ही सेट के वजनों से वीडियो, ऑडियो और क्रिया पूर्वानुमान देने वा...","url":"https://www.aioga.com/hi/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:45:28.056Z"},"it":{"title":"Black Forest Labs ha lanciato FLUX 3: un modello di base multimodale unificato per la previsione di immagini, video, audio e azioni robotiche","summary":"Black Forest Labs ha rilasciato FLUX 3, il primo modello di base multimodale che apprende immagini, video e audio in un'unica architettura e produce previsioni di video, audio e azioni dallo stesso insieme di pesi. FLUX 3 Video può generare in una sola volta video di fino a 20 secondi con audio nativo e, in un test di preferenza umano per video generati da testo a 720p di 10 secondi, ha superato Luma Ray 3.2 con un tasso di preferenza del 93%.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs ha lanciato FLUX 3: un modello di base multimodale unificato per la previsione di immagini, video, audio e azioni robotiche - Aioga Notizie IA","description":"Black Forest Labs ha rilasciato FLUX 3, il primo modello di base multimodale che apprende immagini, video e audio in un'unica architettura e produce previsioni di video, audio e az...","url":"https://www.aioga.com/it/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:46:19.586Z"},"nl":{"title":"Black Forest Labs lanceert FLUX 3: een multimodaal basismodel dat beeld-, video-, audio- en robotbewegingsvoorspellingen verenigt","summary":"Black Forest Labs heeft FLUX 3 uitgebracht, het eerste multimodale basismodel dat binnen één enkele architectuur beeld, video en audio leert en video-, audio- en bewegingsvoorspellingen uit dezelfde set gewichten kan genereren. FLUX 3 Video kan in één keer video’s van maximaal 20 seconden met originele audio genereren en versloeg Luma Ray 3.2 met een voorkeur van 93% in menselijke voorkeurstests voor 10 seconden 720p tekst-gegenereerde video’s.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs lanceert FLUX 3: een multimodaal basismodel dat beeld-, video-, audio- en robotbewegingsvoorspellingen verenigt - Aioga AI-nieuws","description":"Black Forest Labs heeft FLUX 3 uitgebracht, het eerste multimodale basismodel dat binnen één enkele architectuur beeld, video en audio leert en video-, audio- en bewegingsvoorspell...","url":"https://www.aioga.com/nl/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:46:11.658Z"},"tr":{"title":"Black Forest Labs, FLUX 3’ü yayınladı: Görüntü, video, ses ve robot hareket tahminini birleştiren çok modlu temel model","summary":"Black Forest Labs, FLUX 3'ü yayınladı; bu, tek bir mimaride görüntü, video ve sesi birlikte öğrenen ve aynı ağırlık setinden video, ses ve hareket tahmini çıktısı üreten ilk çok modlu temel modeldir. FLUX 3 Video, tek seferde 20 saniyeye kadar, orijinal sesli video üretebilir ve 10 saniyelik 720p metin tabanlı video insan tercih testinde, %93 tercih oranıyla Luma Ray 3.2'yi geride bırakmıştır.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs, FLUX 3’ü yayınladı: Görüntü, video, ses ve robot hareket tahminini birleştiren çok modlu temel model - Aioga AI Haberleri","description":"Black Forest Labs, FLUX 3'ü yayınladı; bu, tek bir mimaride görüntü, video ve sesi birlikte öğrenen ve aynı ağırlık setinden video, ses ve hareket tahmini çıktısı üreten ilk çok mo...","url":"https://www.aioga.com/tr/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:47:05.896Z"},"vi":{"title":"Black Forest Labs phát hành FLUX 3: Mô hình cơ bản đa phương thức thống nhất dự đoán hình ảnh, video, âm thanh và chuyển động robot","summary":"Black Forest Labs phát hành FLUX 3, đây là mô hình nền tảng đa phương tiện đầu tiên học đồng thời hình ảnh, video và âm thanh trong một kiến trúc duy nhất, và xuất ra dự đoán video, âm thanh và hành động từ cùng một tập trọng số. FLUX 3 Video có thể tạo ra video dài tới 20 giây với âm thanh gốc chỉ trong một lần, và trong thử nghiệm ưu tiên của con người đối với video tạo từ văn bản 720p dài 10 giây, mô hình này đạt tỷ lệ ưa chuộng 93%, đánh bại Luma Ray 3.2.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs phát hành FLUX 3: Mô hình cơ bản đa phương thức thống nhất dự đoán hình ảnh, video, âm thanh và chuyển động robot - Tin tức AI Aioga","description":"Black Forest Labs phát hành FLUX 3, đây là mô hình nền tảng đa phương tiện đầu tiên học đồng thời hình ảnh, video và âm thanh trong một kiến trúc duy nhất, và xuất ra dự đoán video...","url":"https://www.aioga.com/vi/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:47:06.327Z"},"id":{"title":"Black Forest Labs merilis FLUX 3: model dasar multimodal yang menyatukan prediksi gambar, video, audio, dan gerakan robot","summary":"Black Forest Labs merilis FLUX 3, ini adalah model dasar multimodal pertama yang mempelajari gambar, video, dan audio secara bersamaan dalam satu arsitektur, dan mengeluarkan prediksi video, audio, dan aksi dari satu set bobot yang sama. FLUX 3 Video dapat menghasilkan video hingga 20 detik dengan audio asli sekaligus, dan dalam uji preferensi manusia untuk video teks yang dihasilkan 10 detik 720p, mengalahkan Luma Ray 3.2 dengan tingkat preferensi 93%.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs merilis FLUX 3: model dasar multimodal yang menyatukan prediksi gambar, video, audio, dan gerakan robot - Berita AI Aioga","description":"Black Forest Labs merilis FLUX 3, ini adalah model dasar multimodal pertama yang mempelajari gambar, video, dan audio secara bersamaan dalam satu arsitektur, dan mengeluarkan predi...","url":"https://www.aioga.com/id/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:47:48.365Z"},"th":{"title":"Black Forest Labs เปิดตัว FLUX 3: โมเดลพื้นฐานหลายโหมดที่รวมการทำนายภาพ วิดีโอ เสียง และการเคลื่อนไหวของหุ่นยนต์","summary":"Black Forest Labs เปิดตัว FLUX 3 ซึ่งเป็นโมเดลพื้นฐานมัลติโมดอลรุ่นแรกที่เรียนรู้ภาพ วิดีโอ และเสียงร่วมกันบนสถาปัตยกรรมเดียว และสามารถส่งออกการทำนายวิดีโอ เสียง และการเคลื่อนไหวจากชุดน้ำหนักเดียวกัน FLUX 3 Video สามารถสร้างวิดีโอความยาวสูงสุด 20 วินาทีพร้อมเสียงต้นฉบับในการสั่งสร้างเพียงครั้งเดียว ในการทดสอบความชอบของมนุษย์ในการสร้างวิดีโอ 720p ความยาว 10 วินาทีด้วยข้อความ มีอัตราความชอบ 93% เอาชนะ Luma Ray 3.2","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs เปิดตัว FLUX 3: โมเดลพื้นฐานหลายโหมดที่รวมการทำนายภาพ วิดีโอ เสียง และการเคลื่อนไหวของหุ่นยนต์ - ข่าว AI Aioga","description":"Black Forest Labs เปิดตัว FLUX 3 ซึ่งเป็นโมเดลพื้นฐานมัลติโมดอลรุ่นแรกที่เรียนรู้ภาพ วิดีโอ และเสียงร่วมกันบนสถาปัตยกรรมเดียว และสามารถส่งออกการทำนายวิดีโอ เสียง และการเคลื่อนไหวจา...","url":"https://www.aioga.com/th/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:48:10.630Z"},"pl":{"title":"Black Forest Labs wydaje FLUX 3: ujednolicony multimodalny model bazowy do przewidywania obrazów, wideo, dźwięku i ruchów robotów","summary":"Black Forest Labs wydało FLUX 3, pierwszy multimodalny model podstawowy, który w pojedynczej architekturze uczy się obrazów, wideo i dźwięku, a z tego samego zestawu wag generuje prognozy wideo, dźwięku i ruchu. FLUX 3 Video może jednorazowo wygenerować wideo o długości do 20 sekund z natywnym dźwiękiem, a w teście preferencji człowieka dla 10-sekundowych wideo w jakości 720p generowanych z tekstu pokonało Luma Ray 3.2 z 93% preferencją.","category":"模型更新","source":"MarkTechPost（RSS）","aggregationSource":"MarkTechPost（RSS）","pageTitle":"Black Forest Labs wydaje FLUX 3: ujednolicony multimodalny model bazowy do przewidywania obrazów, wideo, dźwięku i ruchów robotów - Aioga Wiadomości AI","description":"Black Forest Labs wydało FLUX 3, pierwszy multimodalny model podstawowy, który w pojedynczej architekturze uczy się obrazów, wideo i dźwięku, a z tego samego zestawu wag generuje p...","url":"https://www.aioga.com/pl/news/cms23ou2z022fro9fuxwjdsli/","contentTranslated":true,"sourceHash":"5e4b2fd5e0a23f3d","translatedAt":"2026-07-26T18:48:53.297Z"}}}}