World Labs 发布新一代世界模型 Atlas,从零预训练,原生处理文本、图像、视频和 3D,架构为多模态自回归扩散 Transformer,输入以 3D 姿态形成空间上下文。
模型支持相机控制生成(最长 1 分钟 1440p 视频)、稀疏视图 3D 重建、时空仿真和图像生成,在相机控制生成与 3D 重建任务上优于更专用的模型; 现已开放早期访问申请,未来将驱动 Marble 等产品。
世界模型可以生成、重建和模拟任何可能的世界。它们理解世界的外观、行为和演变方式,使我们能够为创意用户呈现想象中的世界,高保真地模拟现实世界,并帮助机器人规划行动。在World Labs,我们致力于构建这些通用的世界模型,以追求空间智能。
今天我们推出了Atlas,我们的下一代世界模型。Atlas是一种全能模型,我们从零开始进行预训练,使其能够原生处理文本、图像、视频和3D数据。它是一种多模态自回归扩散变换器:所有输入都被整合到一个共享的空间上下文中。Atlas利用该上下文生成下一步内容,在三维空间中保持与已看到内容的一致性,并想象其之外的部分。Atlas设计为可扩展:随着训练计算的增加,其性能会提高,我们预计在持续扩展的过程中这一趋势将保持。
Atlas可以执行广泛的任务,包括世界生成、重建和模拟:
Atlas将为未来版本的Marble(https://marble.worldlabs.ai/)及World Labs的其他产品提供动力。
Atlas可以接受一个或多个参考图像,并在您指定的任意相机位置和角度生成新视图。生成的视图与输入图像的内容和几何形状一致,平滑地推展超出输入图像范围,从而想象输入中未显示的场景部分。
Atlas能够处理各种类型的场景、视觉风格和相机运动。
视频可以由一到六张输入图像生成,并配合手动设计的相机路径。
Atlas将精确的相机几何作为原生输入类型,超越了基于文本的粗略相机控制指令。这使您能够构图每一个镜头,并控制每一个动作。
在这里的示例中,Atlas从单张输入图像生成完整场景。它利用输入图像的内容及其广泛的世界知识,从新的角度想象场景应有的样子。例如,它生成机器人的背面,并推测泳池旁边应该有一片草坪。
从单张图像,Atlas可以生成任意角度的视图。拖动以改变视角。
类似于大型语言模型(LLM),Atlas首先将其输入编码为上下文,然后基于上下文生成输出。然而,Atlas的独特之处在于每张图像都在空间中的3D位置上有对应的基点;这形成了空间上下文。
管理这种空间上下文可以解锁全新类型的创意控制。例如,你可以在上下文中放置两张不相关的参考图像,将它们定位在3D空间中,Atlas就能生成一个在它们之间平滑插值的世界。
这些示例展示了模型的世界知识和创造力;它能够构想出门口、走廊、角落以及其他在原本无关的输入图像之间的过渡场景。
选择左右两帧以填充空间上下文,Atlas会将它们拼接在一起。
Atlas让你通过结合摄像机运动和空间上下文管理生成长视频,并实现精确控制。你可以设计每一个场景和每一个镜头角度。这让你成为导演:你在布置场景,而不是像拉老虎机一样等待结果。
在下面的示例中,我们使用少量参考图像生成一段1分钟、1440p分辨率的视频。我们手工设计了摄像机在场景中的路径,Atlas生成了一个连贯的世界。本页面上的其余视频已经被压缩以优化页面性能。
Atlas可以根据一张或多张输入图像重建真实世界的空间。它不需要特殊的拍摄设备,也不需要数百张密集视图就能忠实重建物体和场景。我们认为Atlas是在从稀疏输入图像生成新视图这一长期存在的3D计算机视觉基础问题上迈出的重要一步。
Atlas可以接受可变数量的场景输入视图。当世界的某些部分在输入视图中不可见时,Atlas会通过利用其丰富的世界知识,想象出一种填补空白的合理方式。
但有时你不需要想象力;你可能需要对现实世界地点的精确重建。提供更多输入图像可以给 Atlas 更多的上下文:它看到的越多,想象就越少。Atlas 通常只需两到三张图片就能提供逼真的重建效果,表现超过专门仅为 3D 重建训练的模型的最先进结果。然而,Atlas 也可以利用一百多张输入图片的空间上下文,从而忠实地再现现实世界环境。
在下面的第一个例子中,Atlas 仅用一张地面照片就生成了场景的俯视图。单张输入照片中可见的花园在模型输出中得到了准确再现,但场景的其他部分是想象的。在加入第二张花园旁小屋的真实世界输入图片后,模型输出显示了花园和小屋,但它仍然想象左边的房子。在加入第三张主屋的输入图片后,整个场景被准确描绘。
在第二个例子中,我们逐步建立斯坦福大学的主广场,从长满草的主入口开始,到纪念教堂外立面装饰的彩色马赛克结束。尽管 Atlas 只接收两到二十五张地面输入图像,但它仍然可以生成从高空俯瞰校园的路径。
Atlas 可以在同一场景中生成许多不同的轨迹,从而对相同的输入图像提供新的视角。无论你多少次更换摄像机路径,场景都保持一致。
在下面的例子中,我们展示了在给定少量输入图像的情况下,Atlas 可以生成通过同一场景的多条摄像机路径。不同的摄像机路径可以强调场景的不同部分,或通过改变速度、长度或复杂性来改变气氛。
Atlas 可以使用少量输入图像生成许多不同的路径通过同一场景。
在上面的结果中,你已看到 Atlas 输出了二维图像和视频,这对于某些应用来说已足够。但机器人、游戏、设计、视觉特效等领域的工作流程通常需要明确的三维输出。Atlas 原生同时操作二维图像帧和三维深度图,使其能够将世界输出为点云或三维高斯斑点。
从一张图像,Atlas生成新的视图和3D几何,然后转换为3D高斯点云
从单张输入图像,Atlas通过联合生成新视图和估计其几何形状来生成完整的3D世界。对于真实空间的视频,它预测每一帧的深度,并将它们组合成3D重建。在任何情况下,Atlas都会填补摄像机从未拍摄过的区域。
Atlas可以从输入视频重建三维点云
点云估计场景的几何形状,但3D高斯点云使其可用。Atlas填补剩余的空隙,将点云变成完整的点渲染场景,并可以在设备上高分辨率、高帧率渲染。这与Marble使用的表示相同:https://marble.worldlabs.ai/,使Atlas能够自然地与我们其他产品整合。
Atlas作为世界模拟器。它理解世界的空间结构以及世界随时间的演变。结合其空间和时间能力,带来了VFX、机器人及其他领域的新应用。
Atlas将少量普通摄像机变成“子弹时间”多视角拍摄工作室。只需三台摄像机的拍摄,Atlas就能冻结时间并重新构图,让你从不可能的角度观看事件。
无需昂贵的拍摄工作室,就能从新的摄像机角度重新构图真实世界视频
值得注意的是,这些镜头不需要专业摄影师或专业设备。每个镜头都是由少数工程师和研究人员使用普通手机安装在三脚架和背包可装的夹具上拍摄的。Atlas从三到五个摄像机视角重建场景,之后你可以随意重新构图镜头。
幕后花絮:上面的片段仅使用少量手机和运动摄像机拍摄
Atlas为Real-to-Sim的扩展开辟了新途径:./real-to-sim-to-real,适用于导航和操作。
你已经看到Atlas从几张图像明确地重建了空间。对于机器人来说,重建只是工作的一半:当模拟机器人在空间中移动时,Atlas还会生成其传感器沿途观察到的RGB和深度数据。世界和机器人对世界的视角都来自同一个模型。
在这些示例中,我们使用手机视频捕捉了两个大型环境,每个环境使用 24 帧进行重建。像这样的空间扫描传统上需要复杂且昂贵的设备。然后,我们模拟不同类型的机器人沿不同路径导航,并使用 Atlas 从机器人机身安装的摄像头视角生成图像。
Atlas 重建空间并辅助模拟机器人导航
机器人操作更进一步。通过一些随意录制的视频,Atlas 有助于构建一个模拟,这个模拟还捕捉了物体的移动和交互方式。一旦任务被模拟,你可以轻松地改变它:更换物体、它们的位置、机器人的动作、光照、背景。结果是生成多样化的训练数据和大规模机器人测试环境。
浏览器不支持视频播放。
Atlas 仅通过少量真实世界的录制实现从真实到模拟,重新创建刚性、关节和可变形物体的物理交互,同时支持可控的变化。
Atlas 的主要关注点是世界建模,每张图像都是通向可能世界的窗口。尽管图像生成不是其主要焦点,Atlas 仍然是一个强大的图像生成器:它可以遵循复杂提示、渲染文本并生成各种视觉风格。
Atlas 还可以根据文本或图像提示生成 360° 图像,同样可以生成各种场景类型和视觉风格。
Atlas 是一个全能模型,设计用于在单一统一架构中处理多种任务以及多种输入输出数据,并将空间控制置于模型核心。这些目标要求我们偏离 LLM 和视频模型使用的标准架构,设计一个新的基础架构,作为未来世界模型的基础。
Atlas 是一个多模态自回归扩散变换器。其输入以 3D 空间为基础形成空间上下文,并在其上下文条件下生成多模态输出。
Atlas 是一种多模态自回归扩散变换器。它在多模态序列上运行,每次生成序列的一个新元素。每一种架构特性协同工作以实现我们的目标,综合起来,它们使基于空间上下文的新生成范式成为可能。我们依次展开这些想法:
Atlas 是现代大型语言模型(LLM)和视频模型思想的融合。它可以受益于这两类模型中使用的架构、算法和系统进展。
像 LLM 一样,它是一个自回归变换器,因此可以利用用于服务和加速 LLM 的创新,包括 KV 缓存、缓存感知路由、分离式服务等。像现代图像或视频模型一样,它是潜在扩散模型,并可以使用如扩散蒸馏、无分类器引导、偏移噪声调度以及 VAE 设计进展等算法。
Atlas 是一个用于世界建模的综合模型,能够执行许多任务。因此,没有单一的基准能完全反映其通用性。我们重点展示 Atlas 在两个关键任务上的定量评估:基于摄像机的生成和 3D 重建。在这两项任务中,它都优于更专业的模型。
我们与一些在基于摄像机生成中表现优异的视频模型进行了对比。在每次试验中,我们将单张输入图像与一个到三个电影摄像机运动(平移、平行移动、升降机等)序列配对。
我们为每个模型提供一张输入图像和目标摄像机路径。对于 Atlas,我们使用其原生摄像机输入格式对摄像机路径进行编码。其他模型不接受摄像机作为原生输入格式,因此我们在输入文本提示中使用标准电影术语描述摄像机路径。对于某些模型,更复杂的提示工程或创造性的多模态提示可能会改善摄像机跟随,但我们使用文本输入,因为这是描述摄像机运动最常用的输入方式。
一组第三方人工评级人员评判哪个模型更好地遵循预期的摄像机路径。这些结果确认,Atlas 在基于摄像机控制的生成任务中优于近期的视频模型,并且随着摄像机轨迹变得更加复杂,这一优势会进一步增加。
我们还评估了 Atlas 在从稀疏输入视图进行 3D 重建的任务中的表现。在每次试验中,模型会接收一组图像及其相机姿态,并预测每个输入像素对应的 3D 点。这个问题在学术界引起了广泛关注,近年来已经开发了许多专业的重建模型。
Atlas 是一个全能模型,既能进行生成,也能进行重建。尽管具有通用性,Atlas 仍然优于目前最佳的开源专业重建模型。
我们在该任务的多个最先进基准上进行了评估,为了确保所有方法之间有共同且公平的评估协议,我们重新实现了所有基线的结果。
3D重建误差(数值越低越好)
现代人工智能的大部分进展都是由规模化驱动的。模型在很大程度上通过将简单算法扩展以利用更多数据和计算资源而得到提升。
我们看到了强有力的证据表明,Atlas 将随着规模继续提升。我们从零开始在大规模多模态数据语料上对 Atlas 进行了预训练。在开发过程中,我们训练了一系列规模和训练计算逐步增加的模型,并发现每一个新计算级别都会解锁新的模型能力。我们相信,未来的世界模型将遵循这一趋势,随着规模的不断扩大,其能力将显著提升。
Atlas 正在与精选合作伙伴进入早期使用阶段。如果您希望使用它进行构建,请在下方申请访问权限,我们会与您联系。我们期待看到您构建的成果,并与您合作,使 Atlas 成为生成、重建和模拟任何世界的首选世界模型。
我们还在招聘:https://www.worldlabs.ai/careers 跨研究和工程,以推进空间智能。
World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave, and evolve so that we can render imagined worlds for creative users, simulate the real world in high fidelity, and help robots plan actions. At World Labs, we build these general purpose world models in pursuit of spatial intelligence.
Today we are introducing Atlas, our next-generation world model. Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context. Atlas uses that context to generate what comes next, staying consistent in 3D with everything it has seen and imagining what lies beyond it. Atlas is built to scale: its performance improves with increased training compute, and we expect this trend to hold as we continue scaling.
Atlas can perform a broad range of tasks spanning world generation, reconstruction, and simulation:
Atlas will power future versions of Marble:https://marble.worldlabs.ai/ and other products from World Labs.
Atlas takes one or more reference images and generates new views at any camera position and angle you specify. Generated views match the content and geometry of the input images, smoothly extrapolating beyond them to imagine parts of the scene not visible in the inputs.
Atlas handles a broad range of scene types, visual styles, and camera motions.
Videos are generated from one to six input images with manually-designed camera paths
Atlas uses precise camera geometry as a native input type, going beyond coarse text-based instructions for camera control. This lets you frame every shot and control every motion.
In the examples here, Atlas generates a complete scene from a single input image . It uses the content of the input image along with its broad world knowledge to imagine what the scene should look like from new angles. For example, it generates the back side of the robot, and it guesses that there should be a grassy lawn next to the pool.
From a single image, Atlas generates views from any angle. Drag to change the view.
Similar to an LLM, Atlas first encodes its inputs into a context, then generates outputs conditioned on the context. However Atlas is unique because each image is grounded at a 3D position in space; this forms a spatial context .
Managing this spatial context unlocks entirely new kinds of creative control. For example, you can place two unrelated reference images in the context, position them in 3D space, and Atlas generates a world that smoothly interpolates between them.
These examples demonstrate the model's world knowledge and creativity; it imagines doorways, hallways, nooks, and other transitions between otherwise unrelated input images.
Select left and right frames to populate the spatial context, and Atlas stitches them together
Atlas lets you generate long videos with precise control by combining camera movement and spatial context management. You design every scene and every camera angle. This puts you in the director's chair: you are staging the scene, not pulling the lever of a slot machine.
In the example below, we generate a 1 minute video at 1440p resolution using a small number of reference images. We hand-design a camera path through the scene, and Atlas generates a coherent world. The rest of the videos on this page have been compressed to optimize page performance.
Atlas reconstructs real-world spaces from one or more input images. It doesn't require special capture equipment or hundreds of dense views to faithfully reconstruct objects and scenes. We believe Atlas is a major step forward toward solving the problem of novel view synthesis from sparse input images, a decades-old fundamental problem in 3D computer vision.
Atlas can take a variable number of input views of a scene. When parts of the world are not visible in the input views, Atlas imagines a plausible way to fill in the gaps by drawing from its rich world knowledge.
But sometimes you don't want imagination; you might want an exact reconstruction of a real-world location. Passing more input images gives Atlas more context: the more it sees, the less it imagines. Atlas typically gives faithful reconstructions with as few as two or three images, outperforming state-of-the-art results by models specially trained only for 3D reconstruction. However, Atlas can also make use of over a hundred input images in its spatial context, allowing for faithful recreation of real world environments.
In the first example below, Atlas generates an aerial view of the scene from just a single ground-level photo. The garden visible in the single input photo is accurately recreated in the model output, but the rest of the scene is imagined. After adding a second real-world input image of the cottage next to the garden, the model's output shows both the garden and the cottage, but it still imagines the house to the left. After adding a third input image of the main house, the entire scene is accurately depicted.
In the second example, we build up Stanford's Main Quad piece by piece, beginning with the grassy main entrance and ending with the colorful mosaics decorating the facade of Memorial Church. Though Atlas only receives two to twenty-five ground-level input images, it can generate paths from aerial views flying far above the campus.
Atlas can generate many different trajectories through the same scene, giving new perspectives on the same input images. No matter how many times you change the camera path, the scene stays consistent.
In the example below, we show that given a small set of input images, Atlas can generate multiple camera paths through the same scene. Different camera paths can emphasize various parts of the scene, or change moods by varying in speed, length, or complexity.
Atlas can generate many different paths through the same scene using a small set of input images.
In the results above you have seen Atlas output 2D images and videos, which are sufficient for some applications. But workflows in robotics, gaming, design, VFX and beyond often require explicit 3D outputs. Atlas natively operates on both 2D image frames and 3D depth maps, enabling it to output worlds as point clouds or 3D Gaussian splats.
From one image, Atlas generates new views and 3D geometry, then converts to 3D Gaussian splats
From a single input image, Atlas produces a full 3D world by jointly generating new views and estimating their geometry. From a video of a real space, it predicts the depth of every frame and combines them into a 3D reconstruction. In either case, Atlas fills in regions that no camera ever saw.
Atlas can reconstruct 3D point clouds from input videos
Point clouds estimate a scene's geometry, but 3D Gaussian splats make it usable. Atlas fills the remaining gaps and turns the point cloud into a complete splat scene that renders on-device at high resolution and framerates. This is the same representation used in Marble:https://marble.worldlabs.ai/, enabling Atlas to integrate naturally with the rest of our products.
Atlas serves as a world simulator. It understands both the spatial structure of the world and how the world evolves over time. Combining its spatial and temporal abilities leads to new applications for VFX, robotics, and beyond.
Atlas turns a handful of ordinary cameras into a "bullet time" multiview capture studio. With footage from as few as three cameras, Atlas can freeze time and reframe shots, letting you view events from impossible angles.
Real-world videos can be reframed from new camera angles without an expensive capture studio
Notably, these shots did not require professional photographers or specialized equipment. Each of them was filmed by a few engineers and researchers with ordinary cell phones on tripods and clamps that fit in a backpack. Atlas reconstructs the scene from three to five camera views, after which you can reframe shots however you like.
Behind the scenes: the clips above were captured using just a few cell phones and action cameras
Atlas opens up new ways to scale Real-to-Sim:./real-to-sim-to-real for both navigation and manipulation.
You've already seen Atlas reconstruct a space in explicit 3D from a few images. For robotics, reconstruction is only half the job: as a simulated robot moves through space, Atlas also generates the RGB and depth data its sensors would observe along the way. The world and the robot's view of it come from the same model.
In these examples, we captured two large environments with a cell phone video, using 24 frames each for reconstruction. Scanning spaces like these traditionally requires elaborate and expensive equipment. We then simulate different kinds of robots navigating different paths, and use Atlas to generate images from the perspective of the robot's body-mounted cameras.
Atlas reconstructs spaces and aids in simulating robot navigation
Robotic manipulation goes a step further. From a few casual recordings, Atlas aids in building a simulation that also captures how objects move and interact. Once a task is simulated, you can vary it easily: change the objects, their positions, the robot's motion, the lighting, the background. The result is diverse training data and testing environments for robotics at scale.
Browser does not support video playback.
Atlas enables Real-to-Sim from just a few real-world recordings, recreating physical interactions with rigid, articulated, and deformable objects while supporting controllable variations.
The primary focus of Atlas is world modeling, and every image is a window to a possible world. Though image generation is not its primary focus, Atlas is a capable image generator: it follows complex prompts, renders text, and generates a wide variety of visual styles.
Atlas also generates 360 images from text or image prompts, where again it can generate a wide variety of scene types and visual styles.
Atlas is an omni model designed to handle many tasks and many kinds of input and output data in a single unified architecture, putting spatial control at the heart of the model. These goals require us to depart from standard architectures used both by LLMs and video models, and design a new base architecture to serve as the foundation of future world models.
Atlas is a multimodal autoregressive diffusion transformer. Its inputs are grounded in 3D space to form a spatial context, and it generates multimodal outputs conditioned on its context.
Atlas is a multimodal autoregressive diffusion transformer . It operates on multimodal sequences, generating each new element of the sequence one at a time. Each of these architectural properties work together to achieve our goals, and taken together they enable a new paradigm of generation based on a spatial context . We unpack these ideas in turn:
Atlas is a blend of ideas from modern LLMs and video models. It can benefit from architectural, algorithmic, and systems advances used in both types of models.
Like an LLM, it is an autoregressive transformer, so it can take advantage of innovations used to serve and accelerate LLMs including KV-caching, cache-aware routing, disaggregated serving, and more. Like a modern image or video model it is a latent diffusion model, and can make use of algorithms such as diffusion distillation, classifier free guidance, shifted noise schedules, and advances in VAE design.
Atlas is an omni model for world modeling that performs many tasks. There is thus no single benchmark that fully captures its generality. We highlight quantitative evaluations of Atlas on two key tasks: camera-conditioned generation and 3D reconstruction. On both tasks it outperforms more specialized models.
We compare against a selection of top-performing video models for camera-conditioned generation. In each trial, we pair a single input image with a sequence of one to three cinematic camera motions (pan, truck, crane, etc).
We prompt each model with a single input image and a target camera path. For Atlas, we encode the camera path using its native camera input format. Other models do not accept cameras as a native input format, so we describe the camera path in the input text prompt, using standard cinematic terms. It is possible that more sophisticated prompt engineering or creative multimodal prompts could improve camera following for some models, but we use text as it is the most common input modality for describing camera motions.
A team of third-party human raters judge which model better follows the intended camera path. These results confirm that Atlas outperforms recent video models at camera-controlled generation , and this advantage grows as camera trajectories become more complex.
We additionally evaluate Atlas on the task of 3D reconstruction from sparse input views. In each trial, the model receives a set of images and their camera poses, and predicts a 3D point corresponding to each input pixel. This problem has attracted much interest in the academic community, and many specialist reconstruction models have been developed in recent years.
Atlas is an omni-model which performs both generation and reconstruction. Despite its generality, Atlas outperforms the best specialized open-source reconstruction models .
We evaluate on several state-of-the-art benchmarks for this task, reproducing the results for all baselines to ensure a common and fair evaluation protocol across all methods.
3D Reconstruction Error (lower is better)
Most progress in modern AI has been driven by scaling. Models improve in large part by scaling up simple algorithms to make use of more data and compute.
We see strong evidence that Atlas will continue to improve with scale. We pretrained Atlas from scratch on a large diverse corpus of multimodal data. Over the course of development, we trained a series of models of increasing size and training compute, and found that each new level of compute unlocked new model capabilities. We are confident that our future world models will follow this trend, dramatically improving their capabilities as we continue to scale.
Atlas is entering early access with select partners. If you'd like to build with it, request access below and we'll reach out. We're excited to see what you build, and to work with you to make Atlas the go-to world model for generating, reconstructing, and simulating any world.
We're also hiring:https://www.worldlabs.ai/careers across research and engineering to advance spatial intelligence.