
How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
Machine Learning Street Talk (MLST) is made possible by:
“So, Hilton called to me the superstition concierge to make your fan rituals a reality. Want to make sure our team doesn't wash your lucky jersey? Hilton's unmatched hospitality can keep up with any superstition.”From the transcript
The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.
He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different rates. Ming-Yu treats "world model" as a set of tools, not one definition: forward dynamics, inverse dynamics and policy, trained together under a capacity limit so that each helps the others. He also explains why plentiful first-person human video carries over to robots, which have far less data of their own, and why a Cosmos model post-trained on the DROID dataset is a good starting point for pick-and-place policies.
The most practical thread is testing. A neural simulator does not need accurate success rates. It only needs to rank policy A above policy B the way the real world would, so a team can narrow down which checkpoints deserve a real trial. Cosmos Dreams applies that closed-loop idea to driving and robotics, and Ming-Yu argues that humanoids around children and pets make safety matter even more than it does for cars. The conversation ends on the Super, Nano and Edge sizes (Edge targets Jetson Thor, Orin and DGX Spark) and where to find the open weights, code and data.
This episode is a paid partnership with NVIDIA.
Learn more about Cosmos: https://nvda.ws/4cJoY1S
Explore Cosmos Lab: https://research.nvidia.com/labs/cosmos-lab/cosmos3/
---
TIMESTAMPS:
00:00:00 A road that was never filmed
00:02:28 Inside Cosmos 3: reasoning and generator towers
00:05:02 World models: dynamics, policy and one clock
00:08:59 Learning robot skills from human video
00:11:06 Ambiguous tasks and system 2 planning
00:12:53 Neural simulators for policy verification
00:16:41 Cosmos as a starting point for robot policies
00:19:00 Cosmos Dreams and robot safety
00:22:04 Super, Nano and Edge model sizes
00:24:24 Open models, the Cosmos repo and feedback
---
REFERENCES:
tool:
[00:00:13] Cosmos 3 (NVIDIA Cosmos Lab project page)
https://research.nvidia.com/labs/cosmos-lab/cosmos3/
[00:18:27] NVIDIA Cosmos GitHub repository
https://github.com/NVIDIA/cosmos
[00:22:05] Cosmos3-Edge model card
https://huggingface.co/nvidia/Cosmos3-Edge
[00:22:15] Cosmos3-Super model card
https://huggingface.co/nvidia/Cosmos3-Super
[00:22:16] Cosmos3-Nano model card
https://huggingface.co/nvidia/Cosmos3-Nano
[00:22:50] NVIDIA Jetson Thor
https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/
[00:22:52] NVIDIA Jetson Orin
https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/
[00:22:53] NVIDIA DGX Spark
https://www.nvidia.com/en-us/products/workstations/dgx-spark/
[00:24:42] Cosmos 3 collection on Hugging Face
https://huggingface.co/collections/nvidia/cosmos3
other:
[00:01:07] Cosmos-Dreams closed-loop simulators (NVIDIA SIGGRAPH 2026 blog)
https://blogs.nvidia.com/blog/siggraph-news-2026/
paper:
[00:08:54] Cosmos 3: Omnimodal World Models for Physical AI
https://arxiv.org/abs/2606.02800
[00:17:43] DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
https://arxiv.org/abs/2403.12945
---
RESCRIPT: https://app.rescript.info/share/e2385948cf465f0d6a2c0930150fc3ab
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
276 searchable segments. Every word is indexed and playable.
Full transcript
Machine Learning Street Talk (MLST) — How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu. Machine-transcribed; use the interactive transcript above to jump the player to any line.
College football is back. So, Hilton called to me the superstition concierge to make your fan rituals a reality. Need a room to match your lucky number? We got you. Want to make sure our team doesn't wash your lucky jersey? Oh, that smells lucky. Hilton's unmatched hospitality can keep up with any superstition. Even a marching bandwick up call it 555 and 55 seconds. Hit it! When you need a team that will do whatever it takes on game day, it matters where you stay. Hilton, for this day. This car is about making a lap turn. Can you envision the trajectory the car is taking and then the video is generated? Guess what? That road was never filmed. It was made by NVIDIA's Cosmos 3. It reads video. It simulates roads and it's learning how to handle objects. Mingyu Liu leads the research and this is a paid partnership with NVIDIA. Now, Cosmos, it takes text, video, audio and actions as input.
Let's start with the simplest case. As a model, what happens in the videos? How many cars are in the video? Whether the robot succeed in complete action in the video. Okay, there are three cars. And the robot succeed in pick up the apple and put in the basket. When you have a video and test as input and generate test, it becomes a vision language model. So why don't we flip this thing on its head? The model stops describing the scene and starts generating it. The device Cosmos dreams is a set of skills to build a cloud loop simulator, not just for sale driving. I can employ a hungry driver, drive a hungry car, go to different intersections and figure out whether it works better. Or I can use this simulator to verify the accuracy. You are not limited by the size of your fleet.
Anytime you need to do a lot of testing, you'll be at the computer, just launch it. Okay, so what if we took the same idea but on an even harder problem? Like robots, for example. You know, navigation, you don't want your physical device to touch any other scenes, but manipulation tasks require all kinds of interactions. When there's an interaction, there's a cushion. There's a potential deformation, you know, when scenes are manipulated. So it's more challenging. We will reach your state where we will be able to use new simulation to handle this complex manipulation well. Once we reach the status, you can test your robot policy with the simulator. Okay, so one model, Cosmos 3, handling three jobs, handling objects in robotics is the future. It's very exciting. So you can see that you're looking at a very diverse, versatile model architecture. And we now have action as a first class citizen,
which means, you know, the model might actually take an action and then its next observation could be caused by the thing that it did before. We start with a language model. We either turn its front scratch also or take a city in open model, right? Start with the language model. It has a basic understanding of the language structure, you know, understand language. Then we have a vision encoder that can take a video or a set of images input. Then we connect this vision encoder to the large language model. Connecting this vision encoder to the language model, we kind of build a vision language model, right? And sometimes you can also start with a future and a vision language model. Okay, now this language model, you know, operating in the auto-regressive manner, we did a token at the time, right?
Okay, in learn to connect, you know, visual output to the language space, it thinks in language, it's spent what it's seen, right? So you have a vision language model. Okay, now we take this vision language model, you know, the pretender weights and then initialize a generator, right? So one key thing is that this vision language model is auto-regressive and our generator is actually, the diffusion is bidirectional. Now this bidirectional generate power none to generate video action audio. Okay, in this tower, every token, attend to every other token. So when you generate the video of the action, there's a coherence among the chunk it generates, among each generation, the signal within this chunk is coherence and it also dirage the token
from the vision language model part, the original tower, which is because the reason tower. And so it also understands the instruction, you know, what you want to generate. Okay, and in that generator training, there are two stages. We training, none the basic, and mid-training, we add action in. In your estimation, what is a world model? This is a very challenging question and I think earlier, there's a sound-syncing signal, people tried to give a definition to AGI. I think several years ago, they started it, but I think still today, there's still no kind of common agreement what AGI is, right? People have their own definition. And I think a world model gonna be the same, right? So I think a world model is a collection of useful tools, right? We model something because we are trying to achieve some goal, right? For example, it could be pretty the future, or it could be
understand why this happened. There are three important modeling in robotics, four dynamics, you know, given what the starting point and the action taken, what would be the future? Pretty what gonna happen? Even if you're not a model, inverse dynamics, given the visual transition, you know, infer what was the action taken? And policy, what should the robot do to achieve the task? We're combining these different modalities like audio and action and vision and so on and they might run on different timescales. How do you make them all run on the same time scale? Different syncing signal have different frequency, right? Even for video we have, come with a different frame rate, high speed video, low speed video. And the audio, you know, there's also different hertz and action also different.
So we have a temporal position in bedding skin where we put it much normalized all the signal into the same essence, the same scale. So that's when the token represent different time chunk or different special time chunk of the signal, better know which other token are in the same time instance, at the same time instances. And also know the relative distance between different time instances. So that is critical part to make this only model work, you know, handling audio, video and action all come with different frequencies. You were talking earlier about these different modalities, forward dynamics, inverse dynamics and policy in one second model. How exactly do they reinforce each other? When we change the model, we put them together, right? And we give the information bottleneck.
You only have these amounts of capacity to expand all blend. These are all trying to connect a video visual observation to action. So in the world action model, right, you, you have a test instruction where you want to achieve. And this is why you observe. And you start to generate action and possible outputs, right? And possible output is similar to forward dynamics, right? When you have action in what could be the outputs. In the dynamics, when you have this visual sensation, what could be the action? Right? Also similar to what happened in the world action model, right? You have the video and action in the correlation between them. Yeah. So I think these three, they are trying to fundamentally trying to capture the correlation between the observation and action. Right? It's just a different, different slice, different perspective. Yeah. And once you put the information bottleneck,
you put the capacity content, you know, the model known. And our result in a pair of sure that there's a synergy. One does help the other. And I think it's fair to say there's a bit of an asymmetry in the training data, right? So there's lots of human generated video data. And then there's there's far less action data. And but but you get this virtuous transfer between the modalities. But could the model ever know when not to transfer, when, when transfer could be harmful? I think the model doesn't really know. It's so I think talking about we have a larger amount of ego centri video, right? The first person view and the two hands trying to complete certain task. And we have really small amounts. And you must smaller amounts of the robot video, you know, robot also have a head camera, you know, group of camera. And you know, see it's complete certain task, right? So I think it's more about our shared vocabulary for different embodiments, right?
So, you know, human or robot looks like human, right? And how human hand manipulate one object, the visual pattern, the visual action correlation, human, how human manipulates an object. The pattern, you can see it's very similar to how a robot would manipulate an object. Even though the action space doesn't really correlate precisely, once you correlate one set of action, the pattern is similar, it's easier to generalize to the other. And it also give you, when you incorporate a couple different embodiments in, it also helps you to generalize to unseen embodiment. Because I mean, human or robot, they're gonna look like human, maybe some have bigger hands, some have longer hands, but it can have a similar, right? So, and group of camera can be different, but, you know, the function are also similar.
The key is really the pattern, the correlation between the visual observation and action, right? And the translation from one action in body, one in body, one to another in body, one, I think it's really easy. Apparently, one of the failure modes is when tasks are not well enough specified. I think a system two is the key to make ambiguous tests more concrete for system one to execute. Yeah, it's fascinating. So, I'm thinking that we want to use this technology in safety, critical environments, right? And some tasks are just quite ambiguous. But presumably, we could add in layers of, you know, maybe we don't understand exactly how to specify the thing that we're doing, but we could add in layers of tests and checking and we could red team it and make it improve over time. You give a task, it will be similar to LN agents, right? You give a task. First initial call of the agents or the model will do find a second of task, need to be complete, right?
And it will execute one by one and after each execution, it will check whether this is complete. And at some point, namely, it might be choices, need to be made, right? So, okay, should I do this or the other, right? I think this will require us to think this robotic task more in the system level, a layer above the model level. You have some harness to kind of get the resources together, either you pass memory or the tools that you can use, right? So, I'm seeing this is going to be a more more complex way to do get the physical task. One thing we've not spoken about enough is is using Cosmos as a simulator essentially to train a policy.
And I'm just thinking whether there would be some potential problems with that. You know in reinforcement learning, we have the SIM to real gap. And could you have a situation where the policy might learn to exploit features inside the simulator? Yeah, it's possible. The way I look at how this neural simulator progress is that I think it will be first, it will be more useful for passive verification in the beginning. Passive verification in the task that you know, if you are a model builder, any model wrong, you're going to have a lot of chip points. And you might do a version study and you're going to every, you know, variance, you're going to have many chip points. And how do you know which one is better, right? The ideal case is to deploy each policy in your real robot, in your real environment and make sure the complete rate, right? Success rate. In all the hope happened to sell driving car, right?
Every time you update your driving policy, how do you know it was? So you need to have a fleet of driver and go to different challenging scenario and check whether this policy is desirable, right? It's so costly. There's so many chip points. So if we can use a war model as a wall simulator, as a replacement of the real world and real driver, right? You have the policy that was the interact with the wall simulator. And the wall simulator in the end of the row out, you can measure whether this test is completed or not. Then you have the success rate, right? If the success rate of the wall simulator, the new simulator, co-elect with real work, you know, testing, the ranking preserve. You don't need them to be precise. You just need to know if policy A is better than policy B in neural simulator.
And most likely, policy A is going to be better than policy B in the real world. You just need this ranking preserve. Then with this, you can quickly narrow down on the number of chip points you need to do real developments. You're going to largely improve your development velocity, help you to find a better policy sooner. Okay. And in this policy verification, you don't directly trend the model using the visual output, the row out, right? And the model, I mean, if what happened in the last year, you're not going to happen in war model, we're going to see rapid improvements of war model, right? So even before it's mature enough to feed the policy model, the training data to train, wall simulator can be used as a policy verification. Yeah. That's what I think. Okay. So now we go back to the question you asked. There's a hacking. And I'm thinking going to happen in, you know, even if you use deep learning, you know, if there's some pattern, some scene that can explore it, right?
That's going to be other words of samples. So I think a new simulator, you know, going to be sports and other similar going to also be sports. And I think as we reach the status, I think there will be ways to regularize the model training to measure those cannot be used. Yeah. So I don't know precisely the approach. I think this is also a very exciting area. And I just believe that's, yeah, people are way. Yeah. And on this thing about using Cosmos as a teacher, essentially, for policies, you guys have got a set of recipes around this and you should talk to that. But one interesting thing is the ability to imagine, you know, like the, the, the notional elephant in the room. So something that is quite unnatural that would be quite low frequency in the training data, you can imagine those, but you could also have this adaptive system where you could recognize edge cases.
And then you could improve the policy over time. The Cosmos as the teacher. So in Cosmos, right. So we try to help the ecosystem in three ways. Better data, better environment and better starting points. Right. So we believe that a Cosmos model, only model model, model, let's link video action test together to learn the shared representation is a great starting points to be a policy model. It not just a theory, we actually post chain, Cosmos model on the story that I said she still the art, a big and place policy results. And the reason, I mean, our intuition is that if we can predict the pixel dynamic well, and we can learn the correlation between the, the role action, then you're going to help you to build a policy model. When, when we kind of keep the venting Cosmos model, when we bring more embodiment in it, it's gonna become a better and better starting points. And the amounts of effort required for adapt to a new embodiment will be less.
Well, maybe, maybe you should comment on the recipes because this will really help folks at home get to get up to speed with it. As we build a Cosmos as a platform, we also provide post training recipe in our kind of a Cosmos report. So that's a user, a developer can give you the recipe to reproduce the results we have starting with the Cosmos model. Right. So, and we do forward to add a more recipe there, maybe also skills to help you, you know, provide data and then agent help you to post chain Cosmos and to do something you like. Yeah. So that is trying to serve the community this way. And can you tell me about Cosmos streams, we divide Cosmos dreams is a set of skills to build a cloud loop simulator for not just for sales driving. It's for all sorts of embodiment. Right. So it's a four into the conventional traditional definition, one model in robotics, right, action in, you know, future observation out for dynamics.
You know, you're not limited by the size of your fleet. Right. And anytime you need to do a lot of testing, you'll be of the computer just launch it. And it can help you the demand velocity, you know, sales driving very useful. And I think it is going to also be true for robotics. I think one model now is good enough for navigation task in robotics. The other challenge is the manipulation task. You know, navigation, you don't want your physical device to touch any other scenes, but manipulation tasks require all kinds of interactions. When there's an interaction, there's a cushion, right. And there's a potential deformation, you know, when since I manipulated them. So it's more challenging. I'm very optimistic with the excitement the whole field has in one model and the, you know, the continual events of the deepening research and also computers.
We will reach your state where we will be able to use new simulation to handle this complex manipulation well. Right. Once we reach the status, you can test your robot policy with the simulator. Right. And you know, sales driving car requires a lot of policy verification today because cars moving around in our human space. Right. So we need them to be safe enough. Right. So when robot with human robot reach the status where you're going to be deploying on many people's houses, safety is even more important. You don't expect to see a car, you know, in your house. But once human always everywhere, you're going to expect to see humanoid. So only you and maybe you're kid, you're pets. Right. So you do want them to be safe enough. And how do you know? Right. So a developer of human robot, how do you measure your every iteration of your new your policy, you know, all the whole system is safe enough. You need a policy verification. And you won't have enough space for setting up all kinds of questions, all kinds of different tasks. And so I think you see new simulator is the only way to give you the development of the system.
To give you the development velocity, you know, we need this. And you also have this Cosmos 3 edge model. So this runs on a single device. Yes. Cosmo come with models come with three different sizes. We have a super, nano and H. Right. Super is the, you know, the give you the highest fidelity is a frontier model. If you want very high accuracy and you can, you have to come. The resources are definitely go with super. Nenol is in between smaller, easier to to portray with all sorts of GPU resources you have. We want to measure it wrong really well on edge devices like the Justin Thor or an old DJ spark. Right. So the reason we build this one is that we want the world model deep where robot leads. Right. So that's you don't need to have a round chip to data center. We want to take certain action. And as it's important, you know, when we come to real deployments, right. You may not be able to come down the, you know, the network connection good enough to do all the tasks. Right. And sometimes it's sensitive critical. Right. You have to complete the task at the time. So we did we need to do a lot of things.
So we did we develop each so that it's small enough, but also powerful enough to run on those edge device. And because the model is smaller. It you know, take this compute resources. So it can be easy to find you. And earlier we have a recipe on fine to the customers age model in one day to boost its visual understanding capability. So I think this all three version cannot be super important. Yeah. And I also envision that let me need to work together in some form for future robotic development. Very challenging task. You probably want a very powerful model. You know, it's like co out, co for help. And for simple task, you know, you have to finish it right away. You know, you want to depend on the model. This is on the edge device. And for folks at home, where can they get their hands on on these models and find out more about it.
Yeah. So all our models are open and all our co and even some of the training data we created are open. So the models and the data live in Hagen face and the code live in GitHub report. Yeah. So we have a GitHub.com slash media slash Cosmo as our landing report. Finally, you can see links to the Hagen face models, the data sets and also the training code, the post training script. Yeah. And we also have skills in the report to help your agent quickly pick up the offering in Cosmo. We need to help the ecosystem to help the a wide range of physical eye developers. I have my email box, you know, I as you know, I'm watching every issues come to this.
Cosmo repo and we do want to see your feedback, you know, good or bad both are very welcome. And we are committed committed to make it better and better. Ming you it's been a pleasure and an honor having your own MLST. Thank you so much for joining us today. Sing a theme.
More episodes
More from Machine Learning Street Talk (MLST)

When AI Research Starts Moving Faster Than Human Research - Zhengyao Jiang
Machine Learning Street Talk (MLST)

How Deep Learning Finally Cracked Messy Tables - Frank Hutter
Machine Learning Street Talk (MLST)

Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick
Machine Learning Street Talk (MLST)

Speech Recognition Is Not a Solved Problem — Pavan Kumar Reddy
Machine Learning Street Talk (MLST)