Why it matters
Cosmos 3 offers a foundational model for developers building physical AI systems, enabling them to simulate and understand the physical world for applications like robotics and autonomous vehicles. Its unified approach simplifies development by consolidating various AI capabilities into one model.

What changed

NVIDIA's Cosmos 3 marks a significant advancement in World Foundation Models (WFMs) for physical AI, introducing a unified omni-model architecture. Previously, developers had to manage separate models for distinct tasks such as world generation (Cosmos Predict), controlled generation (Cosmos Transfer), scene understanding (Cosmos Reason), and policy generation (Cosmos Policy). Cosmos 3 consolidates these capabilities into a single model built on a Mixture-of-Transformers (MoT) backbone, capable of processing multiple modalities including text, image, video, audio, and action within a single forward pass. This unified approach allows the model to seamlessly switch between functions like Vision-Language Model (VLM), video generator, forward/inverse dynamics model, or robot policy without architectural modifications. The MoT architecture encodes each modality via dedicated encoders before projecting them into a shared representation space. The input sequence is then split into an autoregressive (AR) subsequence for reasoning and a diffusion (DM) subsequence for generation, with AR and DM tokens interacting through joint attention.

This release includes two model sizes: Cosmos 3 Nano, a 16B parameter model (8B reasoner, 8B generator) optimized for inference on workstation-grade GPUs like the RTX PRO 6000, and Cosmos 3 Super, a 64B parameter model (32B reasoner, 32B generator) designed for large-scale synthetic data generation and research, requiring NVIDIA Hopper and Blackwell GPUs. Both models are available on Hugging Face with model cards and licensing. Integration with the Hugging Face Diffusers library is provided via Cosmos3OmniPipeline, simplifying the use of generation pipelines with minimal code. NVIDIA is also releasing open synthetic data generation (SDG) datasets for physical AI, including Embodied-Robot-Scenes, Physical-Interaction-Scenes, Digital-Human-Scenes, Autonomous-Driving-Scenarios, and Warehouse-Operations-Scenes, all available on Hugging Face.

Why it matters for builders

Cosmos 3 empowers builders to develop physical AI systems that possess a deeper understanding of the real world, moving beyond simple pixel and token recognition to grasp concepts like motion, causality, and physics. This unified model simplifies the development workflow for complex applications such as training robots for tasks like laundry folding, creating realistic simulations for autonomous driving, or generating synthetic data for safety scenarios in warehouses. The availability of two model sizes, Nano for efficient inference and Super for large-scale tasks, provides flexibility for different deployment needs and research objectives.

Practical impact

Developers can now leverage a single model for a wide range of physical AI tasks, including generating physically plausible video worlds from various inputs (text, images, videos, actions), reasoning about physical properties, and predicting future states or actions. The integration with Hugging Face Diffusers allows for straightforward implementation into existing pipelines. Builders can explore the capabilities through provided examples for text-to-video, image-to-video, and single-frame generation. Post-training scripts are available on GitHub for fine-tuning Cosmos 3 on custom datasets, enabling adaptation to specific robots, environments, or tasks. The release of SDG datasets also provides valuable resources for training and evaluating new physical AI models.

Caveats and source limits

The source material does not provide specific pricing information for Cosmos 3 or its associated services. While the release mentions availability on Hugging Face and GitHub for post-training scripts, exact performance benchmarks or detailed comparisons against previous Cosmos versions or competing models are not included. The source also does not specify the exact release date beyond June 1, 2026, nor does it detail the licensing terms beyond stating they are available with the model cards. Information regarding the specific hardware requirements for running Cosmos 3 Super beyond mentioning NVIDIA Hopper and Blackwell GPUs is also limited.

Sources

Written with AI assistance from the linked sources; every claim below was checked against them automatically. How we produce articles.

Claim check: 9/9 supported claims - 9 evidence links - 100% avg confidence
  • NVIDIA Cosmos 3 is an open omni-model for physical AI reasoning and action.supported - huggingface.co
  • Cosmos 3 combines world generation, physical reasoning, and action generation into a single model.supported - huggingface.co
  • Cosmos 3 is built on a Mixture-of-Transformers (MoT) architecture.supported - huggingface.co
  • Cosmos 3 supports multiple input and output modalities including text, image, video, audio, and action.supported - huggingface.co
  • Cosmos 3 Nano is a 16B parameter model optimized for efficient inference on workstation-grade compute.supported - huggingface.co
  • Cosmos 3 Super is a 64B parameter model designed for large-scale synthetic data generation and research.supported - huggingface.co
  • Cosmos 3 is integrated with the Hugging Face Diffusers library.supported - huggingface.co
  • NVIDIA is releasing open synthetic data generation (SDG) datasets for physical AI.supported - huggingface.co
  • Post-training scripts for Cosmos 3 are available on GitHub.supported - huggingface.co

Caveats

  • Single-source caution: verify critical details at the linked source.
Radar score 79/100 - how it was calculated
Reliability87
Freshness100
Novelty67
Technical55
Developer56
Ecosystem80
Confidence96
  • Reliability 87: Primary official source
  • Freshness 100: Fresh official source date
  • Novelty 67: Official announcement
  • Technical 55: Structured technical source signals
  • Developer 56: Builder relevance source signals
  • Ecosystem 80: Official source
  • Confidence 96: Claims have reliable evidence
Share
XLinkedInHacker News

Related articles

Benchmarks - Sep 13, 2026MindTopo Benchmark Evaluates Topological Spatial Reasoning in Foundation ModelsResearchers have introduced MindTopo, a new benchmark designed to assess foundation models' ability to understand and reason about topological spatial relations, which are invariant under continuous deformation. The benchmark evaluates models on both 'reasoning' tasks, such as identifying relations, and 'planning' tasks, where models act as agents in an environment to manipulate these relations.Research Papers - Aug 22, 2026DreamHand: Video Diffusion Models for Occlusion-Robust 3D Hand RecoveryResearchers have introduced DreamHand, an offline framework that repurposes video diffusion models (VDMs) as deterministic geometry encoders for 3D hand motion recovery from egocentric video. The system achieves state-of-the-art results on multiple benchmarks, significantly improving accuracy in occlusion-heavy scenarios and for hands that go out of sight.Robotics - Sep 29, 2026Rolling-WAM: World Action Models with Rolling ImaginationResearchers have introduced Rolling-WAM, a novel formulation for World Action Models (WAMs) that addresses latency issues in robotic manipulation. By distributing the joint video-action denoising process across successive replanning cycles, Rolling-WAM aims to improve closed-loop responsiveness.Research Papers - Sep 29, 2026AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive ControlResearchers have introduced AD-WM, an action-discriminative world model designed for counterfactual model predictive control (MPC). This model aims to improve the ability of MPC systems to distinguish between alternative actions from the same state, a crucial aspect often overlooked by models focused solely on factual prediction accuracy.Research Papers - Sep 21, 2026Agile-WAM: Tactile World Action Model for Robot ControlResearchers introduced Agile-WAM, an agile tactile World Action Model designed for contact-rich robot control. This model efficiently integrates visual and tactile data to predict future world states and robot actions, outperforming baselines in success rates and achieving low inference latency.Research Papers - Aug 20, 2026DA-WAM: Decision-Aligned World Models for Autonomous DrivingResearchers have introduced DA-WAM, a novel framework for autonomous driving that unifies predictive representation learning, action-conditioned future modeling, and trajectory scoring. This approach ensures that predicted futures directly inform decision-making by generating distinct future latent states for each trajectory candidate.