Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. NVIDIA releases Cosmos 3, an open omni-model for physical…

NVIDIA releases Cosmos 3, an open omni-model for physical AI (world generation, reasoning and actions)

★★★open-sourceNVIDIAconfidence: high

NVIDIA published open weights for Cosmos 3 (Nano 16B, Super 64B) around 2026-06-01: one Mixture-of-Transformers model that takes text, images, video, audio and robot actions and generates video, images, audio, text or actions, replacing the separate Cosmos Predict, Transfer, Reason and Policy models.

Key facts

What happened

Cosmos 3 is NVIDIA's first single world foundation model that can generate worlds, reason about physics and output actions. Before it, developers had to chain separate models: Predict 2.5, Transfer 2.5, Reason 2 and Policy.

Why it matters

It is the largest open world model aimed at robotics and autonomous vehicles. It also shows the field moving toward models that do both "world simulation" and "policy" in one, a direction GR00T N2 is also expected to follow.

Changelog

  • 2026-09-29: created

Models

Videos (2)

Introducing NVIDIA Cosmos 3: The Open Model That Thinks, Generates, and Acts

NVIDIA · 2026-06-02 · official

Description by Gemini, which watched the video:

Summary
This official launch video from NVIDIA introduces Cosmos, an open frontier omni-model designed for physical AI. Narrated over conceptual diagrams and video demonstrations, the video outlines Cosmos's architecture—a Mixture of Transformers combining an autoregressive reasoning transformer and a diffusion generator—and its applications across reasoning, synthetic data generation, simulation, and robotic policy execution.

What is shown

  • Autonomous Driving Edge Cases [00:01–00:09]: Real-world driving in a Mercedes-Benz test vehicle identifying a rolling ball and a pedestrian child crossing, displaying live "Reasoning" and "Meta Actions" overlays.
  • Architecture Overview [00:14–00:34]: A schematic showing Cosmos processing text, image, video, audio, and action inputs through a "Mixture of Transformers" architecture consisting of an Autoregressive Reasoner connected to a Diffusion Generator.
  • World Reasoner (VLM) [00:39–00:48]: Cosmos analyzing drone timelapse footage of an urban traffic intersection to generate a structured traffic report with observations and actionable engineering insights.
  • Data Generator & World Model [00:49–00:57]: Physics-accurate synthetic video generation depicting an unusual road hazard (a mattress flying off a truck on a highway).
  • Simulator & OmniDreams [00:58–01:12]: Cosmos operating within simulation runtimes (AlpaSim) and NVIDIA OmniDreams as an action-conditioned world model, generating predictive sensor output for extreme scenarios (an elephant crossing a residential road, cone navigation at night, and heavy snow).
  • Policy Model / World Action Model [01:13–01:29]: Integration with Alpamayo 2 Super and robotic manipulation, demonstrating multi-step tool grasping (picking up a screwdriver and placing it on a rack) with live step-by-step reasoning and motion planning.

Claims & numbers

  • The narrator claims real-world physical data cannot scale on its own, asserting that "compute is data" for physical AI [00:07–00:12].
  • Cosmos is described as an "open frontier omni-model for physical AI" [00:16].
  • Cosmos utilizes a "Mixture of Transformers" architecture where an autoregressive transformer plans and instructs a diffusion transformer that generates downstream frames/actions [00:19–00:33].
  • Cosmos serves as the underlying foundation for NVIDIA OmniDreams, an action-conditioned world model predicting future sensor outputs frame by frame [01:03–01:10].

Notable quotes

  • [00:11]: "For physical AI, compute is data."
  • [00:16]: "An open frontier omni-model for physical AI, built on a new Mixture of Transformers architecture."
  • [01:37]: "Cosmos: the foundation for developers of the age of physical AI."

Assessment
This is a polished official marketing and architecture announcement from NVIDIA. While it showcases real video samples, simulated robotics rollouts, and software interface mockups, it is heavily produced and cut for promotional impact rather than providing live unedited developer workflows or technical benchmark disclosures.

Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.

Meet Cosmos 3: Our Latest Frontier Model for Physical AI

NVIDIA Developer · 2026-05-31 · official

Description by Gemini, which watched the video:

Summary
Ming-Yu Liu, Vice President of Cosmos Lab at NVIDIA, announces and details the release of Cosmos 3, NVIDIA's foundation model for physical AI. He explains that Cosmos 3 unifies prediction, transfer, physical reasoning, and policy generation into a single "omni" model architecture available in two sizes: Nano and Super.

What is shown

  • [00:00] Ming-Yu Liu introduces Cosmos 3 from NVIDIA.
  • [00:09] Visual recap of previous Cosmos components: robotic arm tea/powder preparation (Cosmos Predict), simulation-to-real domain transfer (Cosmos Transfer), drone inspection of wind turbines with text Q&A reasoning (Cosmos Reason), and tabletop manipulation ("put purple eggplant on plate", "put brown chicken wing on plate") (Cosmos Policy).
  • [00:32] Diagram of the Omni model interface handling text, image, video, audio, and action for both inputs and outputs.
  • [00:40] Architecture diagram detailing the "Mixture-of-Transformer" framework featuring an autoregressive Reasoner tower and a diffusion Generator tower sharing multimodal attention.
  • [01:08] Physical AI downstream robotics and autonomous driving clips, including dual-arm manipulation, tool sorting, race car telemetry, and night driving lane prediction.
  • [01:28] Robotic bread-toasting demo evaluating next-best action and generating step-by-step reasoning tokens.
  • [01:48] Leaderboard benchmark tables shown: VANTAGE-Bench, Traffic Anomaly Reasoning (TAR), PAI-Bench (Physical AI Bench), R-Bench, RoboLab-120 Overall, and Artificial Analysis Image-to-Video Leaderboard.
  • [03:03] Announcement of open availability via Hugging Face and GitHub.

Claims & numbers

  • The presenter claims Cosmos 3 is NVIDIA's strongest and most versatile model built to date, unifying previous discrete models into a single architecture.
  • The model is released in two sizes: the smaller Nano model (tailored for edge device deployment) and the Super model (optimized for high accuracy in physical AI tasks).
  • The architecture is a novel "Mixture-of-Transformer" with two towers: an autoregressive tower and a diffusion tower.
  • Benchmark claims highlighted:
    • Ranked #1 on reasoning benchmarks including VANTAGE-Bench and TAR (Traffic Anomaly Reasoning).
    • Top performance on generation benchmarks including PAI-Bench and R-Bench.
    • Ranked #1 in RoboLab (RoboLab-120) for policy evaluation (Cosmos Nano-Policy shown at top with 476/1200, score 73.1).
    • Ranked #1 for open-source models on the Artificial Analysis Image to Video Leaderboard (Cosmos3-Super-Image2Video shown with 1,212 ELO).
  • The presenter states Cosmos 3 is open, with weights available on Hugging Face, code examples on GitHub, and training scripts and datasets provided.

Notable quotes

  • [00:24] "In Cosmos 3, we bring all of them together in a single model. The latest Cosmos 3 model is the Omni model."
  • [00:40] "And it's based on a novel architecture called Mixture-of-transformer, where you have two towers. The left tower runs autoregressive, the right tower runs diffusion."
  • [02:43] "At NVIDIA, we want to help accelerate the physical AI revolution. We are doing our part to build high quality, open physical AI foundation models to unlock all the developers."

Assessment
This is an official NVIDIA product launch presentation featuring an executive walkthrough accompanied by motion graphics, benchmark tables, and pre-recorded robotics/driving test clips. While the performance metrics are backed by standard third-party and community benchmark leaderboards, the robot clips and simulations are curated highlight reels rather than unedited live interactive demonstrations.

Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.

Related events

  1. NVIDIA GTC 2026 robotics: GR00T N2 world action model previewed, Cosmos 3 and GR00T N1.7 announced ★★★
  2. NVIDIA posts $96.2B quarter; Vera Rubin in full production and deploying at major clouds ★★★★
  3. RoboArena: crowd-sourced, double-blind real-world evaluation of generalist robot policies ★★
  4. NVIDIA unveils Isaac GR00T Reference Humanoid, an open humanoid research platform built with Unitree and Sharpa ★★
  5. Xiaomi open-sources Xiaomi-Robotics-U0, a 38B unified world model that generates multi-view robot scenes and training data ★★

Sources (5)

id: 2026-06-01-nvidia-cosmos-3-open-release · updated 2026-09-29 · open in the interactive timeline