为了实现这一点,世界模型:https://www.nvidia.com/en-us/glossary/world-models/ 学习物理环境的行为、接下来可能发生的情况以及哪些后续动作是合理的。它们可以生成物理基础的世界和动作数据,模拟未来状态,并为团队提供一个基础,使其可以针对机器人、自动驾驶车辆或视觉 AI 系统进行专业化。
开放世界模型已经被用于生成训练数据、测试策略并专业化物理 AI 系统。NVIDIA Cosmos 3:https://www.nvidia.com/en-us/ai/cosmos/ 将这些能力整合到一个开放模型家族中,在基准测试中取得领先成绩,并在机器人、自动驾驶车辆和视觉 AI 领域得到广泛采用。
NVIDIA Omniverse:https://www.nvidia.com/en-us/omniverse/ 库是 NVIDIA Agent Toolkit 的一部分,提供构建可模拟世界的预构建功能,物理 AI 团队可以利用这些功能在真实部署之前对系统进行训练、测试和验证。
物理 AI 背后的数据难以收集,并且成本高昂,尤其是在所需规模下。罕见事件和长尾情景尤其难以安全且重复地再现。
除了 Cosmos 外,NVIDIA 的物理 AI 堆栈还包括用于机器人技术的 Isaac GR00T:https://developer.nvidia.com/isaac/gr00t,用于自动驾驶车辆的 Alpamayo:https://www.nvidia.com/en-us/solutions/autonomous-vehicles/alpamayo/ 以及用于视觉 AI 的 Metropolis:https://www.nvidia.com/en-us/autonomous-machines/intelligent-video-analytics-platform/。
在各行各业,开发者正在基于 NVIDIA Cosmos 构建物理 AI 应用:Doosan Robotics、LG Electronics、Samsung Electronics 和 Skild AI 在机器人领域;理想汽车、小米和 Afari 在自动驾驶汽车领域;以及 Centific:https://www.centific.com/blog/centific-brings-last-mile-physical-ai-to-production-with-nvidia-cosmos-3、Fogsphere:https://fogsphere.com/fogsphere-announces-cosmos-3-support/、Linker Vision:https://www.linkervision.com/post/linker-vision-unveils-application-driven-ai-grid-for-agentic-video-reasoning-at-scale、Milestone Systems:https://www.milestonesys.com/resources/content/articles/milestone-hafnia-nvidia-cosmos-3/ 和 Yuan:https://www.yuan.com.tw/news/preview-news?id=336&t=d74eb353274d4fd78e460600ae11a561 在视觉 AI 代理方面:https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/,推动工业 AI 和智能空间应用的发展。
NVIDIA Cosmos 联盟:https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai 通过汇聚世界模型构建者、AI 开发者和物理 AI 领导者,贡献模型、研究和评估方法,从而推进这一工作。NVIDIA 最近将该联盟扩展到日本:https://nvidianews.nvidia.com/news/japans-robotics-and-manufacturing-leaders-build-on-nvidia-cosmos-to-advance-physical-ai-frontier,在那里,机器人和制造业领导者计划加入并开发用于工厂、物流、农业、建筑、医疗和交通的开放世界模型。
这些实现和合作正共同将开放世界模型确立为适用于机器人、自动驾驶汽车和视觉 AI 系统的物理 AI 可适应基础。
通过探索以下资源,了解更多关于世界模型、OpenUSD 和物理 AI 开发的信息:
探索开放的 Cosmos 3 模型集合:https://huggingface.co/collections/nvidia/cosmos3 以及 Hugging Face 数据集:https://huggingface.co/nvidia 和 GitHub:https://github.com/nvidia-cosmos。
Editor’s note: This post is part of Into the Omniverse :https://www.nvidia.com/en-us/omniverse/news/ , a series focused on how developers, 3D practitioners and enterprises can transform their workflows using the latest advancements in OpenUSD :https://www.nvidia.com/en-us/omniverse/usd/ and NVIDIA Omniverse :https://www.nvidia.com/en-us/omniverse/ .
In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership:https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf,” an open letter arguing that AI leadership will be measured not by any single frontier model but by whether an open ecosystem reaches every sector.
Open models:https://www.nvidia.com/en-us/glossary/open-models/, which anyone can download, inspect, modify and run on their own infrastructure, are what make that possible. Nowhere is that more crucial than in physical AI:https://www.nvidia.com/en-us/glossary/generative-physical-ai/, where every deployment is a specialization problem.
Physical AI has to understand and predict consequences, not just appearances.
To make this possible, world models:https://www.nvidia.com/en-us/glossary/world-models/ learn how physical environments behave, what may happen next and which following actions make sense. They can generate physically grounded world and action data, simulate future states and provide a foundation that teams can specialize for a robot, autonomous vehicle or vision AI system.
Open world models are already being used to generate training data, test policies and specialize physical AI systems. NVIDIA Cosmos 3:https://www.nvidia.com/en-us/ai/cosmos/ brings these capabilities together in an open model family, with leading benchmark results and adoption across robotics, autonomous vehicles and vision AI.
And NVIDIA Omniverse:https://www.nvidia.com/en-us/omniverse/ libraries, part of NVIDIA Agent Toolkit, provides prebuilt capabilities for building simulation-ready worlds that physical AI teams can use to train, test and validate systems before real-world deployment.
The data behind physical AI is difficult and expensive to collect at the scale required. Rare events and long-tail scenarios can be especially difficult to reproduce safely and repeatedly.
More useful data by learning physical relationships from large-scale multimodal scenarios.
More diverse environments that vary in weather, lighting, objects and trajectories.
A better foundation to build on and adapt to a particular robot, vehicle, sensor configuration, task or operating environment.
A general model hasn’t seen a team’s particular robot, sensors or operating environment. Closing that gap requires access to model weights, a license that permits adaptation and the tools needed for post-training.
NVIDIA Cosmos world foundation models are available under the Linux Foundation’s OpenMDW 1.1 license, enabling teams to post-train models on their own data and hardware. Specialization is where openness becomes a practical technical requirement.
Specializing a model is only part of the workflow. Teams also need environments to generate data, run simulations and test behavior.
Omniverse libraries help developers build simulation-ready environments, while OpenUSD:https://www.nvidia.com/en-us/glossary/openusd/ provides the open framework for composing, reusing and exchanging complex 3D data across digital twins:https://www.nvidia.com/en-us/glossary/digital-twin/, simulations and synthetic data generation:https://www.nvidia.com/en-us/glossary/synthetic-data-generation/ workflows. Together, Omniverse and OpenUSD cut the duplicated work that can otherwise pile up every time assets, sensor configurations or environmental conditions change.
NVIDIA Cosmos 3 — a frontier open physical AI foundation omni-model:https://www.nvidia.com/en-us/glossary/omni-model/ built on a mixture-of-transformers:https://www.nvidia.com/en-us/glossary/mixture-of-transformers/ architecture — combines vision reasoning, world generation and action prediction, letting developers use one model family to understand scenes, generate synthetic data, simulate future states and build specialized world action models:https://www.nvidia.com/en-us/glossary/world-action-model/.
Developers can use Cosmos 3 as a vision language model:https://www.nvidia.com/en-us/glossary/vision-language-models/, as a physics-grounded world simulator that predicts future world states and generates large-scale synthetic data, or as the backbone for world action models, instead of assembling and maintaining a separate model for each capability.
The family includes Cosmos 3 Super (64B) for high-fidelity world modeling, Cosmos 3 Nano (16B) for efficient reasoning and post-training, and Cosmos 3 Edge:https://huggingface.co/nvidia/Cosmos3-Edge (4B) for on-device vision reasoning and robot policy deployment. Lightweight enough to run on edge GPUs, Cosmos 3 Edge can be deployed across NVIDIA RTX GPUs, NVIDIA DGX systems and NVIDIA Jetson, including Jetson Thor platforms.
Across benchmark evaluations, Cosmos 3 ranks No. 1 on Artificial Analysis:https://artificialanalysis.ai/image/leaderboard/text-to-image/open-weights for open weights text-to-image and image-to-video generation, on PAI-Bench:https://huggingface.co/spaces/shi-labs/physical-ai-bench-leaderboard for world generation and in the image-to-video category of Physics-IQ:https://physics-iq.github.io/. For robot policy, it ranks No. 1 on RoboLab:https://research.nvidia.com/labs/srl/projects/robolab/leaderboard.html. Cosmos 3 Super is also the highest-ranked open model on VANTAGE-Bench:https://huggingface.co/spaces/clemson-computing/VANTAGE-Bench-Leaderboard for vision understanding.
In addition to Cosmos, NVIDIA’s physical AI stack includes Isaac GR00T:https://developer.nvidia.com/isaac/gr00t for robotics, Alpamayo:https://www.nvidia.com/en-us/solutions/autonomous-vehicles/alpamayo/ for autonomous vehicles and Metropolis:https://www.nvidia.com/en-us/autonomous-machines/intelligent-video-analytics-platform/ for vision AI.
Across industries, developers are building on NVIDIA Cosmos for physical AI applications: Doosan Robotics, LG Electronics, Samsung Electronics and Skild AI in robotics; Li Auto, Xiaomi and Afari in autonomous vehicles; and Centific:https://www.centific.com/blog/centific-brings-last-mile-physical-ai-to-production-with-nvidia-cosmos-3, Fogsphere:https://fogsphere.com/fogsphere-announces-cosmos-3-support/, Linker Vision:https://www.linkervision.com/post/linker-vision-unveils-application-driven-ai-grid-for-agentic-video-reasoning-at-scale, Milestone Systems:https://www.milestonesys.com/resources/content/articles/milestone-hafnia-nvidia-cosmos-3/ and Yuan:https://www.yuan.com.tw/news/preview-news?id=336&t=d74eb353274d4fd78e460600ae11a561 for vision AI agents:https://www.nvidia.com/en-us/use-cases/video-analytics-ai-agents/ powering industrial AI and smart spaces applications.
The NVIDIA Cosmos Coalition:https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai extends this work by bringing together world model builders, AI developers and physical AI leaders to contribute models, research and evaluation methods. NVIDIA recently expanded the coalition to Japan:https://nvidianews.nvidia.com/news/japans-robotics-and-manufacturing-leaders-build-on-nvidia-cosmos-to-advance-physical-ai-frontier, where robotics and manufacturing leaders intend to join and develop open world models for factories, logistics, agriculture, construction, healthcare and transportation.
Together, these implementations and collaborations are establishing open world models as an adaptable foundation for physical AI across robots, autonomous vehicles and vision AI systems.
Learn more about world models, OpenUSD and physical AI development by exploring these resources:
Explore the open Cosmos 3 model collection:https://huggingface.co/collections/nvidia/cosmos3 and datasets on Hugging Face:https://huggingface.co/nvidia and GitHub:https://github.com/nvidia-cosmos.
Read the Cosmos 3 technical report:https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf for full architecture details and evaluations.
Read the Cosmos 3 technical blog:https://developer.nvidia.com/blog/.
Tune in to the Cosmos Labs livestreams:https://www.addevent.com/calendar/ss55fmjpm04t.
Learn about the NVIDIA Cosmos Coalition:https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai.
情报判断
Aioga 编辑摘要
NVIDIA 发布 Cosmos 3,将视觉推理、世界生成和动作预测整合进开放物理 AI 模型家族。材料称,该系列可生成训练数据、模拟未来状态,并支持机器人、自动驾驶与视觉 AI 系统的专门化。
背景分析
物理 AI 不仅要识别外观,还要理解和预测行动后果。世界模型通过学习物理环境的运行方式、后续可能状态及合理动作,为训练、策略测试和部署前验证提供基础;相关数据通常难以大规模采集。