6 Sources
[1]
How Cosmos 3 Helps Physical AI Think Before It Acts
The new, open NVIDIA world foundation model brings vision reasoning, multimodal generation and action prediction together to help robots, autonomous vehicles and vision AI agents think before acting in the real world. The real world is always in motion. To operate autonomously, physical AI systems
[2]
NVIDIA launches Cosmos 3, chip-fab tools and humanoid robot platform
NVIDIA unveiled a broad set of technologies aimed at accelerating the development of physical AI systems, expanding its push into humanoid robots, autonomous vehicles, semiconductor manufacturing, and industrial automation. Announced at GTC Taipei, the company's latest releases include Cosmos 3,
[3]
Nvidia's new world model helps robots navigate the world
Why it matters: Nvidia is continuing its move beyond chips into AI models and software, positioning itself to become a foundational platform for physical AI development. Driving the news: Nvidia says it trained Cosmos 3 on 20 trillion tokens of multimodal data, including nearly a billion images,
[4]
NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI
* NVIDIA Cosmos 3 is a new leaderboard-topping open physical AI foundation model, built on a breakthrough mixture-of-transformers architecture for physical AI reasoning, world simulation and action generation. * Cosmos 3 is the world's first fully open omnimodel with native vision reasoning and
[5]
Why NVIDIA's Cosmos 3 is a Massive Leap for Multimodal AI
NVIDIA's Cosmos 3, introduced at GTC Taipei, represents a significant leap in multimodal AI by unifying five distinct data types, text, images, videos, audio and actions, into a single framework. This integration eliminates the need for separate models, streamlining complex tasks like text-to-video
[6]
NVIDIA Calls Cosmos 3 The World's First Fully Open Omnimodel, As Robots And Autonomous Vehicles Get A Powerful Brain Grounded In Physics
NVIDIA has just announced its Cosmos 3 world model at the ongoing GTC Taipei, giving us a glimpse at what it calls the world's first "fully open omnimodel" that is capable of vision-based reasoning, while supporting multimodal output in the form of text, image, video, and ambient sound. NVIDIA's
Share
Copy Link
NVIDIA launched Cosmos 3, an open world foundation model for physical AI that combines vision reasoning, multimodal generation and action prediction. Trained on 20 trillion tokens including nearly a billion images and 400 million videos, the model helps robots and autonomous vehicles understand causal relationships and predict outcomes before acting in real-world environments.
NVIDIA unveiled NVIDIA Cosmos 3 at GTC Taipei during COMPUTEX, marking a significant expansion in the company's push beyond chips into physical AI systems
1
2
. The new open world foundation model addresses a fundamental challenge: enabling robots and autonomous systems to operate in real-world environments where capturing and recreating scenarios is slow, expensive, and often impossible to repeat at scale1
. "The big bang of physical AI is just around the corner thanks to breakthroughs in multimodal reasoning language, vision and world models," said Jensen Huang, founder and CEO of NVIDIA4
.
Source: Geeky Gadgets
NVIDIA trained Cosmos 3 on 20 trillion tokens of multimodal data, including nearly a billion images, 400 million real and synthetic videos, ambient audio, text and action data from humans and robots
3
. This massive dataset gives developers a powerful pretrained foundation for building physical AI systems with less data and lower training costs4
. The action data distinguishes Cosmos from regular video generators, as it's designed to model how machines move, not just how scenes look, according to Ming-Yu Liu, VP of NVIDIA's Cosmos Lab3
. The multimodal AI model can generate rare or dangerous scenarios such as robot collisions or unusual road events that are difficult, expensive or unsafe to capture repeatedly3
.The foundation model is built on a breakthrough mixture-of-transformers architecture that pairs a reasoning transformer with an expert generation transformer
4
. This dual-tower design enables vision reasoning and multimodal generation across text, images, video, ambient sound and action in a single system1
. The architecture allows Cosmos 3 to understand object interactions, motion and spatial-temporal relationships before generating video and action trajectories4
. Developers can use the model as a vision language model, a world model for simulating physical environments, or as the backbone for world action models that help train robotics systems to perform specific tasks4
.
Source: NVIDIA
Cosmos 3 is designed to generate action data such as robot joint angles, gripper positions and trajectories that can help train machines to navigate and manipulate the physical world
3
. In a warehouse, a robot may encounter object configurations it's never seen before, while on the road, an autonomous vehicle may need to respond when a pedestrian steps out from between parked cars1
. The model delivers leading results on physical AI benchmarks, ranking first among open models across Artificial Analysis, Physics-IQ, PAI-Bench and R-Bench for world generation accuracy, RoboLab and RoboArena for action policy, and the VANTAGE-Bench and TAR leaderboards for vision understanding4
.NVIDIA is releasing two versions immediately: Cosmos 3 Super, a 32-billion-parameter model for tasks requiring high physics accuracy such as training robots and autonomous vehicles, and Cosmos 3 Nano with 8 billion parameters per tower for faster inference that can generate results in fractions of a second
3
5
. An edge model that can run locally for real-time, on-device processing is coming soon3
5
. The model reduces physical AI training and evaluation cycles from months to days4
.Related Stories
NVIDIA launched the Cosmos Coalition, a global collaboration between world model builders and AI developers including Agile Robots, Black Forest Labs, Generalist, LTX, Runway and Skild AI to advance next-generation world models in AI
4
. The company also introduced the Isaac GR00T Reference Humanoid Robot, an open reference design combining a Unitree H2 Plus humanoid robot, Sharpa dexterous hands, Jetson Thor onboard computing, and the Isaac GR00T software stack2
. Research organizations including Ai2, ETH Zurich, Stanford Robotics Center, and UC San Diego plan to use the platform2
.NVIDIA is bringing AI for semiconductor manufacturing deeper into production through its collaboration with TSMC
2
. TSMC is using NVIDIA CUDA-X libraries and AI models for computational lithography, transistor simulation, process control, wafer inspection, and fab scheduling, achieving improvements in computational efficiency while using NVIDIA Metropolis and TAO Toolkit to improve detection of nanometer-scale defects2
. The announcements highlight NVIDIA's strategy to build a full-stack ecosystem for physical AI covering everything from synthetic data generation and simulation to real-world deployment in industrial automation2
. Physical AI developers across industries are building on the Cosmos platform, including companies like Li Auto for autonomous vehicles and Samsung for robotics applications4
. NVIDIA's bet is that the next wave of AI won't just answer questions or generate images but will need to predict, simulate and act in the physical world, with AI agents capable of understanding causal relationships and executing complex tasks3
.Summarized by
Navi
[2]
[5]
07 Jan 2025•Technology

17 Jul 2026•Technology

12 Aug 2025•Technology

1
Science and Research

2
Policy and Regulation

3
Technology