
The Story
For the last couple of years, the robotics story has been a pile of separate parts. You had a vision model to see the scene. A world model to imagine what happens next. A vision-language-action model — a VLA — to actually move the arm. Three models, three teams, three sets of headaches getting them to agree with each other.
관련해서 a model lab buying into a robot company도 함께 참고하시면 좋습니다.
For the opposite end — small models running locally — see Liquid AI’s on-device Nanos models.
At GTC Taipei on May 31, 2026, NVIDIA shipped a “physical AI foundation model” that tries to be all three at once. It’s called Cosmos 3, and the framing they went with is the interesting part: they’re calling it an “omnimodel.”
Here’s what that means in plain terms. Cosmos 3 is one model that takes in text, images, video, ambient sound, and actions — and generates the same range back out. So it can watch a scene and reason about it like a vision-language model. It can imagine a plausible future — a video of what happens if the robot pushes that cup — like a world model. And it can serve as the backbone for predicting the actions themselves. NVIDIA built it on a “mixture-of-transformers” architecture, and they’re pitching it as, in their words, the world’s first fully open omnimodel. Don’t let the jargon scare you: the whole idea is that perception, imagination, and action stop being three separate boxes and become one continuous thing.
It comes in a few flavors. Cosmos 3 Super is the high-accuracy version aimed at robots and self-driving cars, where the physics has to be right. Cosmos 3 Nano is the fast one — it does video and action reasoning in fractions of a second. And there’s a Cosmos 3 Edge coming for real-time inference on the device itself. Super and Nano are already out, and — this is the part that matters — they’re on build.nvidia.com, Hugging Face, and GitHub. Open weights, not a locked API.
On the numbers, NVIDIA says Cosmos 3 ranks first among open models across eight physical-AI benchmarks grouped into three categories. For world-generation accuracy, that’s Artificial Analysis, Physics-IQ, PAI-Bench, and R-Bench. For action policy — how well it drives a robot — it’s RoboLab and RoboArena. And for vision understanding, VANTAGE-Bench and TAR. Take a company’s own leaderboard claims with the usual pinch of salt, but the benchmarks themselves are the ones this field actually uses. NVIDIA says it trained the model on billions of samples spanning text, image, video, sound, and action trajectories.
And they didn’t ship it alone. Alongside the model, NVIDIA announced a “Cosmos Coalition” — a group of AI labs and robotics outfits building on open world models together. The names on the list are worth noting: Agile Robots, Black Forest Labs, Generalist, LTX, Runway, Skild AI. That’s a mix of robotics companies and generative-video shops, which tells you something about where NVIDIA thinks this is heading.
The Takeaway
So why does this matter beyond another product launch? Because it lines up with the one thing this field keeps running into: data.
When I wrote about NVIDIA’s “scaling law” for robot dexterity a couple of days back, the whole argument was that robotics is bottlenecked not by motors or hands but by training data. You can’t teleoperate your way to billions of examples — a human demonstrating tasks one at a time is painfully slow. The way out, the field increasingly agrees, is to generate the data in simulation. And a world model that can imagine physically plausible video is exactly the tool that makes that generation trustworthy. Cosmos 3 is NVIDIA planting its flag on that bet: fold the world model into the same brain that reasons and acts, and you close the loop between imagining data and learning from it.
That’s what makes the “omnimodel” pitch more than marketing. When I covered Google DeepMind’s two VLA models for humanoid control, the architecture there kept perception, reasoning, and low-level control as distinct stages that hand off to each other. NVIDIA is arguing the opposite — that stitching separate models together leaves value on the table, and a single unified backbone is the cleaner path. Both bets are live. Nobody’s won yet. But it’s a genuine fork in how the field thinks about building a robot brain, and Cosmos 3 is the clearest statement of the unified side so far.
The open part is the quieter power move. There’s a version of this world where NVIDIA keeps its best physical-AI model behind an API and rents it out. Instead, Super and Nano are downloadable, and the Coalition exists to get outside labs building on the same foundation. My read is that this isn’t charity — it’s the same playbook that made CUDA impossible to dislodge. If every robotics startup and every carmaker learns to build on Cosmos, the model being free barely matters. The GPUs it runs best on are not. NVIDIA doesn’t need to sell you the model. It needs you to standardize on it.
The honest caveat: benchmarks aren’t robots. Ranking first on Physics-IQ is not the same as a machine reliably folding your laundry, and the field is littered with demos that looked incredible and generalized to nothing. Cosmos 3 Edge — the on-device piece that actually runs inside a robot in real time — is still “coming soon,” and that’s the part where physical AI usually gets humbled. So the right frame here isn’t “robots are solved.” It’s that the shape of the tool is settling. A year ago, the answer to “what runs a robot’s brain” was three models in a trench coat. NVIDIA’s answer now is one model that sees, imagines, and acts — and it’s betting the whole industry adopts that shape. Whether it holds up is the next chapter. But it’s a real chapter, and it shipped.
This article is for informational purposes only and is not investment advice.
Photo: Gabriele Malaspina / Unsplash
댓글 남기기