What changed
NVIDIA's Cosmos 3 marks a significant advancement in World Foundation Models (WFMs) for physical AI, introducing a unified omni-model architecture. Previously, developers had to manage separate models for distinct tasks such as world generation (Cosmos Predict), controlled generation (Cosmos Transfer), scene understanding (Cosmos Reason), and policy generation (Cosmos Policy). Cosmos 3 consolidates these capabilities into a single model built on a Mixture-of-Transformers (MoT) backbone, capable of processing multiple modalities including text, image, video, audio, and action within a single forward pass. This unified approach allows the model to seamlessly switch between functions like Vision-Language Model (VLM), video generator, forward/inverse dynamics model, or robot policy without architectural modifications. The MoT architecture encodes each modality via dedicated encoders before projecting them into a shared representation space. The input sequence is then split into an autoregressive (AR) subsequence for reasoning and a diffusion (DM) subsequence for generation, with AR and DM tokens interacting through joint attention.
This release includes two model sizes: Cosmos 3 Nano, a 16B parameter model (8B reasoner, 8B generator) optimized for inference on workstation-grade GPUs like the RTX PRO 6000, and Cosmos 3 Super, a 64B parameter model (32B reasoner, 32B generator) designed for large-scale synthetic data generation and research, requiring NVIDIA Hopper and Blackwell GPUs. Both models are available on Hugging Face with model cards and licensing. Integration with the Hugging Face Diffusers library is provided via Cosmos3OmniPipeline, simplifying the use of generation pipelines with minimal code. NVIDIA is also releasing open synthetic data generation (SDG) datasets for physical AI, including Embodied-Robot-Scenes, Physical-Interaction-Scenes, Digital-Human-Scenes, Autonomous-Driving-Scenarios, and Warehouse-Operations-Scenes, all available on Hugging Face.
Why it matters for builders
Cosmos 3 empowers builders to develop physical AI systems that possess a deeper understanding of the real world, moving beyond simple pixel and token recognition to grasp concepts like motion, causality, and physics. This unified model simplifies the development workflow for complex applications such as training robots for tasks like laundry folding, creating realistic simulations for autonomous driving, or generating synthetic data for safety scenarios in warehouses. The availability of two model sizes, Nano for efficient inference and Super for large-scale tasks, provides flexibility for different deployment needs and research objectives.
Practical impact
Developers can now leverage a single model for a wide range of physical AI tasks, including generating physically plausible video worlds from various inputs (text, images, videos, actions), reasoning about physical properties, and predicting future states or actions. The integration with Hugging Face Diffusers allows for straightforward implementation into existing pipelines. Builders can explore the capabilities through provided examples for text-to-video, image-to-video, and single-frame generation. Post-training scripts are available on GitHub for fine-tuning Cosmos 3 on custom datasets, enabling adaptation to specific robots, environments, or tasks. The release of SDG datasets also provides valuable resources for training and evaluating new physical AI models.
Caveats and source limits
The source material does not provide specific pricing information for Cosmos 3 or its associated services. While the release mentions availability on Hugging Face and GitHub for post-training scripts, exact performance benchmarks or detailed comparisons against previous Cosmos versions or competing models are not included. The source also does not specify the exact release date beyond June 1, 2026, nor does it detail the licensing terms beyond stating they are available with the model cards. Information regarding the specific hardware requirements for running Cosmos 3 Super beyond mentioning NVIDIA Hopper and Blackwell GPUs is also limited.
Sources
Claim check: 9/9 supported claims - 9 evidence links - 100% avg confidence
- NVIDIA Cosmos 3 is an open omni-model for physical AI reasoning and action.supported - huggingface.co
- Cosmos 3 combines world generation, physical reasoning, and action generation into a single model.supported - huggingface.co
- Cosmos 3 is built on a Mixture-of-Transformers (MoT) architecture.supported - huggingface.co
- Cosmos 3 supports multiple input and output modalities including text, image, video, audio, and action.supported - huggingface.co
- Cosmos 3 Nano is a 16B parameter model optimized for efficient inference on workstation-grade compute.supported - huggingface.co
- Cosmos 3 Super is a 64B parameter model designed for large-scale synthetic data generation and research.supported - huggingface.co
- Cosmos 3 is integrated with the Hugging Face Diffusers library.supported - huggingface.co
- NVIDIA is releasing open synthetic data generation (SDG) datasets for physical AI.supported - huggingface.co
- Post-training scripts for Cosmos 3 are available on GitHub.supported - huggingface.co
Caveats
- Single-source caution: verify critical details at the linked source.
Radar score 79/100 - how it was calculated
- Reliability 87: Primary official source
- Freshness 100: Fresh official source date
- Novelty 67: Official announcement
- Technical 55: Structured technical source signals
- Developer 56: Builder relevance source signals
- Ecosystem 80: Official source
- Confidence 96: Claims have reliable evidence