🎨 World Labs Shows Atlas: A World Model with Camera Control

Fei-Fei Li's company has released an "omni" spatial intelligence model that works with text, images, video, and 3D. It generates videos up to 60 seconds long in 1440p, with camera geometry set as a separate input rather than text, and reconstructs 3D scenes from 2-3 to 100+ photos, outputting Gaussian splats.

🌍 Video generation, 3D reconstruction, and simulation are combined in a single model, turning the camera from a side effect of the prompt into a controllable production tool. Real-to-Sim from 24 frames of phone video provides a path to training data for robots.

👤 The scene can be "set up" like in previsualization: define the space, draw the camera trajectory, and get consistent frames from any angle, including bullet time from three. Early access will open in the coming weeks via Marble.

Source 1: https://www.worldlabs.ai/blog/atlas