NVIDIA has announced new Cosmos 3 Super models (64B and 65B parameter versions) that outperform the base Cosmos 3 models in video and image generation speed by 25 times, while maintaining high accuracy in physics and spatial understanding.

What Happened
NVIDIA released optimized Cosmos 3 Super models utilizing a 4-step Mixture-of-Transformers architecture. These models combine reasoning (Reasoner) and generation (Generator) capabilities, allowing them to operate significantly faster than standard solutions while preserving the quality of physical process visualization.
Context
The Cosmos 3 architecture represents a unified system for Physical AI tasks. The use of Mixture-of-Transformers allows for the efficient combination of intelligent scene analysis with the visual content creation process, which is essential for creating realistic digital worlds.
Why It Matters for the Industry
For the AI and robotics industry, this means a radical reduction in computational costs for simulating physical environments. Accelerating the Physical AI development cycle from weeks to hours allows for faster training of robots and autonomous systems in synthetic worlds, making the process scalable and suitable for industrial use.
Why It Matters for Users
Regular users and content creators gain the ability to generate high-quality video using significantly fewer computational resources. This brings the use of advanced simulators closer to less powerful hardware and makes the process of AI video creation nearly instantaneous.
What Is Not Yet Known / Limitations
Engineering environments require special attention to operational stability and latency predictability in production, which remains a critical factor when deploying such models.
Sources
Author
Look at AI, Editorial Staff
