World Labs, founded by Fei-Fei Li, has introduced Atlas — an "omni" world model for spatial intelligence that natively works with text, images, video, and 3D. The model generates video up to 60 seconds long in 1440p, where the camera trajectory is specified by explicit geometry as a separate input, not a text prompt, and reconstructs 3D scenes from multiple shots. Public access to Atlas is not yet available: early access applications are accepted through Marble.

What Happened
World Labs announced the release of Atlas, describing it as a universal world model that combines video generation, 3D reconstruction, and simulation. Architecturally, it is a multimodal autoregressive diffusion transformer with rectified flow and a shared spatial context. The model generates video up to 60 seconds long in 1440p resolution with pixel-precise camera control: camera geometry is provided as a separate input, not described in text. Atlas reconstructs 3D scenes from 2–3 to 100+ shots, outputting point clouds and 3D Gaussian splats, and assembles a bullet time effect from three angles. According to the company, in human comparisons, Atlas was preferred over MiniMax H3 in 75% of cases, Gemini Omni Flash in 81%, FLUX 3 in 93%, and Seedance 2.5 in 94%; in 3D reconstruction, its average error was 8.6 compared to 11.1 for Pi3X, with comparisons also conducted against VGGT-Ω 1B and MapAnything. A Real-to-Sim mode was also demonstrated separately: from 24 frames of a phone video, the model reconstructs a scene with RGB, depth, and physics for rigid, articulated, and deformable objects.
Context
Atlas advances the direction of world models — generative models that work not with individual frames, but with the entire scene space. The central idea of the release is the camera as an explicit input type: in conventional video models, the position and movement of the camera are inferred from the text prompt as a side effect, whereas Atlas accepts camera geometry as a separate field. This design turns a generative model into a tool similar to previsualization, where a scene can be "set up" and then shot from any angle. World Labs' bet is that a single world model will eventually displace separate pipelines of "video generator plus separate 3D reconstruction plus separate simulator," and that world models themselves will move from research demos to commercial tools. Product integration with Marble, through which early access is opened, indicates movement in exactly this direction.
Why This Matters for the Industry
For the industry, the main shift is that the camera transforms from a side effect of the prompt into a controllable production variable, and the ready-made pattern of "scene plus camera trajectory" provides a strong lever for reuse: from a single consistent 3D scene, one can obtain both video along a specified trajectory, point clouds, and 3D Gaussian splats. If World Labs' claims are confirmed by independent verification, a single universal approach will outperform specialized reconstruction pipelines, which will be an argument against narrow model specialization. Real-to-Sim from 24 frames of a phone video opens a practical path to large-scale generation of training data for robotics with RGB, depth, and physics for rigid, articulated, and deformable objects. For startups, this reduces the cost of previsualization, VFX, and synthetic data creation, but building commercial products on Atlas is not yet possible: there is no public API, pricing, or data on latency and hardware requirements.
Why This Matters for Users
For those working with AI video, a scene is now set up as in previsualization: space is specified, a camera trajectory is drawn, and the model outputs consistent frames from any angle, including bullet time from just three points of shooting on phone tripods. 3D scene reconstruction from 2–3 photos with the output of Gaussian splats, which are rendered on the device in high resolution, is something that can actually be tried in the coming weeks. A specific action is available today: apply for early access through Marble. Planning costs and commitments to clients is still too early — Atlas has no public pricing.
What Is Still Unknown / Limitations
All quality comparisons are a vendor "blog eval" without a disclosed protocol: the number of evaluators, the set of prompts and scenes, shooting conditions, and the method of selecting baselines are not disclosed in the source. The reconstruction metric, 8.6 compared to 11.1 for Pi3X, is provided without units of measurement, datasets, and test conditions, so it is impossible to judge either the statistical or practical significance of the gap. There is no public access to the model, pricing, latency data, or hardware requirements; independent verification will only be possible after early access is opened.
Sources
Author
Look at AI, editorial team
