
Robot-Use Agents: Why General-Purpose Models May Win in Robotics
Get every episode summarized
Each time Y Combinator Startup Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now Frontier researchers are showing that this includes controlling robots.”From the transcript
One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training.In this episode of Decoded, we're joined by the founders of Waddle Labs and RoboCurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents.
Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.
Get every episode summarized
Each time Y Combinator Startup Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
140 searchable segments. Every word is indexed and playable.
Full transcript
Y Combinator Startup Podcast — Robot-Use Agents: Why General-Purpose Models May Win in Robotics. Machine-transcribed; use the interactive transcript above to jump the player to any line.
One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now Frontier researchers are showing that this includes controlling robots. This led MIT professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents, where general purpose models could make different robots more capable. So today, Francois and I invited the founders of Waddle Labs and Robocurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs. Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your companies focus on. I'm Hamay. I'm from Waddle Labs together with Vincent. We work on building LLMs that control robots. And we do this by doing two things. Building a harness that allows LLMs to do this very effectively, selecting data and using that data to train better LLMs. Hi, I'm Jay co founder of Robocurve. We are an e-vows company for physical AI.
We measure everything, any robot, any model, including LLMs and also classical approaches such as VLAs, VGN language action models and World Action models. And we evaluate all kinds of involvements like hands, papers, arms, human rights, quadruplets, all kinds of stuff. So in the last few weeks, videos from both of you guys went quite viral on Twitter. And I think inspired, I think both of them are actually quoted in that fill up by solar essay talking about robot usations. The videos showed things like LLMs being able to unscrew caps, being able to communicate between multiple robots, being able to do tasks like uncapping a pen, for example. And so I thought maybe what would be fun is for Francois and I to dig in with you guys on some of the research that led to this moment to begin. And then maybe we can also show some demonstrations of what this can actually do to start. Why don't we talk about some of the early research around transfer between language models and robot control. So I want to talk about some of the early work on pre training language models for decision making and then also the RT two paper, which established the VLA.
So maybe Jay, do you want to tell us a little bit about some of this early work involving being able to use language models to do any kind of robot control. One of the earliest successful approaches of using AI on robots is the RT two paper where they use a pre trained language model on web, text and images and use that to control robots. And the interesting thing about that is that it's a fine tune version of a language model. And instead of outputting English, for example, they just output what we call an effective pose, which is decoding is that you can then translate into joint commands that can control the robots. In some cases, in some sense, it's very similar to what we see now with flat language models, just that instead of fine tuning, the models are good enough to just do it after box. And this is actually quite similar to like the mapping that I have in my head is kind of like a the COT moment. And so like when we did GSMA K and you said, Susie had had $8, she spent five, how much does she have now.
It had to output like four hash marks and then the answer and then EOS, there was no ability for it to chain a thought. And so then we allowed it to actually have like, OK, let me go eight minus five is three. And so like, hash, hash, hash, three. And so you allowed it to do this kind of thinking before it actually gave an action or an answer. And similarly now like VLA is where basically like in our RT 2 is basically like it has to output an action. There's no like I can't allocate more compute for a more complex task. And then now I have this code chain of thought thing that I can do and I can say, OK, even even if you were using the example where you're not actually outputting the code, just giving it to Astra and let it think, think, think, think, think, and then output an action similarly now like and then with code, it's a more callmogov complexity or lower callmogov complexity code length. To just use code and really compact code to just like here's code just go. I totally agree with you on that. I feel like at the end of the day, it's a lot about a bit of less. Right. If you give the agents or you give the AI more or more economy and you if you unsackled a bit more and give it more resources, it can actually do a lot of the things that we find you in it to do.
I'm just going to hop in here as well. I think the really interesting thing about the bitter lesson here is that like VLA's by the sign architecturally they're built on top of language models as well, right. They can potentially reason they can write code. So perhaps the bitter lesson here isn't necessarily what architect you build necessarily, but it's what data is most useful. Right. Like people have been sort of nagging at this VLA data bottleneck for years now and we're seeing very, very slow progress. And so the less that here is maybe we have a modality of data that we know works very well. And this idea of transferring across different modalities, which having are super excited about like how do you take something that is traditionally out of distribution, lover of all things and make it something that's in distribution. Like maybe the way of harnessing the bitter lesson is saying that's pick it data for which we know this modality is pretty blessed. There's a lot of data. These language for these models understand it well and then use this as a way of unlocking a lot of other domains as well. Yeah. And to add to that, the reason why I do RT 2 is such a good model is because it is using a language model that taps into the modality of all the data that a language model is trained on.
So all those web images and web tags actually improves the VLA compared to just training a specific robotics foundation model without any pre training. So you're benefiting in the case of RT 2, these early VLA approaches from the from pre training. I guess what's the limitation in the, I guess, action taking fine tuning approach that you can now get around if you can directly write code like you're referring to like a difference in data there. What exactly is that difference in data you're referring to? There's a few things that models have got much better since already to like one, I mean, they're much better tool use, for example. So, you know, one kind of data is now these models can, I mean, they're right code much better. So now they can write complex policies as code. So, you know, it's much better using tools to explore the kinds of environments there in like arms they have access to and a lot of this comes from in context learning like the key difference between like these VLA's and we're like GPT 6 or LLM isn't necessarily architecture, but kind of approach you take towards training.
We want to be better less impaled, right? We want to benefit from all kinds of data. We want to pour in computer computer use data into our robot models. We want to put in coding data into our robots models. So when you do that, you just end up getting what's when we think of as like these general purpose LLM agents. And then so to me, it's like, okay, why not just build a really good foundational LLM and I use that control robots rather than training like a model that's specifically for robotics and that's more dependent on just robot data on that. So maybe one of the things that I think if we could rewind the clock a couple of years may have been a good sign that we were training in a good direction to this is research on code as policies. Right. Do you guys want to talk a little bit about when the research community refers to code as policies? What exactly that means? And I think it's worth putting this in the context of when this paper came out, which was coding agents just starting to work. I think this was still in the era in which people were like putting comments and auto filling, you know, Python code blocks, which now we think of as Asian history, but this was like two years ago. So, yeah, how does code as policies work and how did that inspire some of what you guys are now seeing as possible?
I probably rewind back to Voyager where like Voyager was like the one of the first. What is what is a coding agent mean? Coding agent requires good tool use and on the fly tool creation. That's called code. Okay, you have Python built ins. Let's just say I have only by 100 built ins. Those are tools and I have sort and I have like if and I have four and I have while I have all these things and I need to like use those tools to assemble a new tool called a new Francois dot pie. Right. And like now that's a new tool and I get to use that. That was Voyager and like yeah, I don't know. I'm not going to say Voyager was the first one to do this Voyager was the most popular one to do this for Minecraft. And then they created tools that they can invoke on the fly to help them play the game better and compressed thinking and experience into a new tool that I can later invoke. And then I think because everyone in 2020, 24 had this insight was like, okay, if we just pour if you know, Darian Sam both all realized if I pour all of my heart and soul and resources and compute and intelligence into getting better at coding.
We will have liftoff then I can automate the ML engineer. Then I can have liftoff and I'll have all the things I'll solve all the things. And so then that's what everyone did and I don't think that they had the insight. I don't think they're so insightful to know that, oh, if I did this, then we will have LMS that we good enough to write policies for robotics and I can replace all of robotics with actual coding agents and I can displace artists with like, you know, coding agents that are writing JavaScript to make beautiful like art. Like I'm not sure that they had that insight is the honest. Could we talk a little bit about what some of these early code as policy methods were even able to do. So even, you know, this is well ahead of us having coding agents that are widely available and people intuitively understanding how they work. It's well ahead of people using or L methods to actually train those coding agents to be really good at tool calling what did some of those early ones, maybe even before Voyager like code as policies, which is, I think the paper came out in 2022 at the end right around my chat should be T came out. What was the capability of those convertible you can see now. Yeah, those code is policy papers, especially once the Google deep mind team were incredible because what they did is they created these kind of functions like pick up an object, lift up or like move to certain posts.
Like they provided this list of functions to a coding agent in a form of like, they repite on functions and then decoding agent will be able to write code involving these functions to then control the robot to do very complex hacks. And what was very surprising or what this excelled was that code agents can do this very one shot data on need additional robot data to in order to work with this code because they're already trained on so much coding data. They already have a sense of like how to do first, what do you do second in order to like move a block into a ball, for example. And I think it's just one shot ability, this in context exploration ability that really motivated a lot of later work to continue exploring, including us to continue exploring how can apply LLM's to robotics. I know, friends, why you have this framework we've talked about of you know the various ways that learning can happen. There's going to be learning that's embedded into the weights versus being in context learning. So and do you want to quickly talk about that? And maybe we can think about where all the research over the last few years has fit into that framework, especially the direction it seems to be going. I mean, I've been working on this experiment and actually I haven't really crystallized it into a paper yet, but basically it's like what is the most efficient way from us intelligence,
intelligence, first sample perspective to input a learning a let's say that's in this context, a state action results are state action reward, tuple back into the policy. There's ICL there is the context learning where I just like append it. This is how most people are using LLM so that just like, oh no, don't do it like this do it like this. And I'm like, okay, and then it's still in the context and then I can remember I can iterate and that really only works if you train the LLM or at least post train the LLM to be able to learn and improve. And that self refines a reflection like literature from way back that like actually allowed more if you turn into the training set. If you don't have that, it doesn't learn doesn't not and kind of learn. But even then there's a limit to how good I see how it can go. So I've done this experiment where like you take an LLM that we trained on and I held out a task, let's just say GSMAK for simplicity.
And then I I see L it and I want to measure on the Valset how much it improves on a per sample basis and it improves greatly very cheaply it doesn't cost much I don't have that there's no SGD so I don't have any flops. So very quickly I can adapt and improve on my Valset example example example one is non monotonic improvement which is wild so it gets worse it gets better gets worse it gets better like pretty graph aggressively number two is it caps out very quickly. And so after like 20 30 maybe 40 examples it is basically saturated and more examples back into the context not improve. You're like you're constrained by the models ability to intelligently use all of its context exactly from the post training how many multi turns did it actually get to be able to improve and do self reflection and then the most important one for sure they can't do is after it hits the context window context length of the model that it was trained on if it was training 100,000 really that means you have a context window of like 50,000 once you exceed 50,000 you don't improve anymore you actually just get worse because the model can't attend over everything.
There's a really example you said which I didn't really think about is like well if I can rag or I can I can load in my act into active memory and this is very prime agent continual harness kind of like thinking rag and pull stuff in on the fly example similar examples then it's like me like I have a exam or I have an exam with a textbook I'm going to do better if I have the textbook to look up as reference and so that is. You know definitely going to work and it's still and also scales would I would imagine much better and then the last two paradigms would be Laura with rank one or rank two or whatever Laura with rank 10 or rank 100 and then all the way to full SFT RL and if you're Tesla and you have all the data infinite data and you're learning doing self driving car by ICL what are you doing right. Like you can it's like clearly we wouldn't do that right and if you're figures like that but it's amazing how good in the in low data regimes you can do with ICL and then even cooler we can compress the ICL into tool use into tools or like compress it into learnings and that's that's the hierarchy that we haven't really figured out you guys are probably on the forefront of.
Yeah, do you guys want to elaborate on that a little bit because I imagine that's a central part of how you guys think about building a lot of there's a few things that are quite interesting here I think to unpack. Yeah, the first thing is we've been sort of thinking about this idea of a harness almost as a form of like domain specificity like when you for example deploy a robot in a new environment maybe it's in a wet lab and it just needs to do a lot of test you picking up for example you can learn a lot of this in context but then one way of consolidating this context for future agents is sort of packaging each of these skills that you learn into specific programs for example like writing out these skills and learning. Writing out these skills and writing out these memories is a form of consolidating it's it's like a form of distillation from past experience for your future agents to use so this is quite interesting the other thing that's quite interesting is it almost seems like there's some relationship to the broader meta learning literature throughout in machine learning history like you have this sort of bigger model that then programs a small model to do things and one of the very interesting things is that yes these smaller models are going to learn in context there's sort of wrapped inside the harness and it's going to be a little bit more interesting.
So you can wrap it inside the harness and they're doing the task you can make them maybe smaller you can make them run faster as long as the bigger model can sort of do this domain specialization of your harness quite well so this is something that we're pretty interested in right now and last point as well I think the really interesting theoretical question is how far in context learning can get you and I think it's really anyone's guess as to that like this paper that show that you know these models sort of can approximate great in descent during in context learning as well so it's this line between weight space and symbolic space I think is a little bit blurry. Why see this next batch is now taking applications got a startup in you apply at ycombinator.com slash apply it's never too early and filling out the app will level up your idea okay back to the video. So what exactly are we watching here? We're watching ashrah controlling the arms to pick a block off the table and put it into the bowl. So here ashrah is using just these camera inputs and presumably it's aware of the robot that it's controlling it's going to write code that controls the robot at its individual joints.
Yes, so it does look at all the cameras it has all the camera feeds instead of code it's more like a tool call. So it sends commands to the robot to to control and put the block into the bowl. So in this situation we're using ashrah directly but if we were using say waddles harness what's the kind of difference between directly using ashrah to control this versus doing this through waddles API sometimes directly like having that we commend the post robot should go to is not direct to optimal tool to use. If it's repetitive task you don't want aster in the loop maybe you want to write code that can run repeatedly very very fast or if it's a task you've done a similar task before you should be able to call a skill that has compiled and use that skill to do the task faster and handle edge cases better. So here we can see it seems to be approaching in on this guy. You can look at you can notice that the latency is a bit slow because we are bottlenecked by the latency of ashrah but if you look at the trends the latency of these models are improving very rapidly.
So one thing that we saw is that for people class LMS their latency is improving by around two X a month is very fast. If the trends continue we could get real time control by end of the year. So now that we saw this task go from end to end why don't we walk through what that actually entails so there's a coding agent in the loop here running this yes. What are the steps that the coding agent would have taken and let's kind of contrast the direct coding agent version versus the waddle harness in the loop. Yes. So what just happened was that in a series of turns ashrah received the images from the cameras and it output the n effect opposed for the robot to go to. And it does it over multiple turns to complete the task. This is less of code as policies but more of two calls. The unscreen part is pretty repetitive so you can actually automate a lot of that using code as well as actually approaching the bottle picking it up like a lot of these things are pretty deterministic once you've done a task quite a few times.
The interesting part is where you build in the variation into the code as policy graph. And so we have a few points of variation this can come from for example when you detect an object maybe use a VLM as far as the tool call. Or when something fails for example how do you check that it's failed how do you then potentially do something else to pay on the failure like this kind of more flexible responses we tend to put a VLM inside a loop to ensure that you know what while the coding graph is sort of deterministic there's points of variation that allows generalize the biggest insight I think I've had a change in world view on how to perform machine learning and how to get the rest of the distance on a GI. And so it was like a lot of conversations I've had with Francher Shuley and in 2020 maybe even 2018 when he did on the measure of intelligence he talked a lot about transduction versus just wrong versus program induction and transaction just means I'm learning a function theta that maps from X is to wise and why is that wrong it's just slow and it's information inefficient and so to go from X is to wise I need a lot of pairs of X is a wise.
A small amount then you need a lot of inductive bias and what you need to act is a good a generator to map from X is to wise so now my theta takes in three and end pairs of X is and wise and he mits the function that maps from X to Y's and that's what code is and so that's it like if you give me a coding interview whiteboard you say okay here's a putting problem here's some examples okay cool right the function. Yeah, we are emitting an F that will map from X to Y and this is quite interesting right because I think I mean this was a lot of the traditional program synthesis literature and I feel like a lot of reason why those methods didn't work as well was because that inductive bias as you said right it's your shifting difficulty of the problem from finding that mapping to find the right inductive biases to map it to this like smaller code to then do the mapping but finding that set of inductive bias is super hard but maybe that's what aster is buying us like even a cox either such a job. I think that's actually a pretty good segue to a kind of final topic which is around inspired by this paper also from the details the platonic representation hypothesis where platonic because as a reference played us cave and you know the idea that you know they're using the same kind of a system that's not really good for us to do that.
shadows of various realities and I think here the point he's making is that there's extensive evidence to show that language models and various representations of different data is actually learn these distance mappings between similar things under different training policies that are overlapping and so like the idea there would be that as you train these systems on increasingly large amounts of data they converge to our consistent mapping of the world and maybe this implies that we would expect language models to get better and better at things over time that would make them useful for new tasks like robot control but maybe I imagine this representation hypothesis is actually pretty central to all of your guys as world view or your view of your companies and so maybe one of you guys tell me a little about how you think about this and also maybe we can use that to make some guesses as to why we think models like asteris seem to be so much better at robot control than previous models and where you make some predictions for where we might be going.
Yeah, I think one way to make this concrete for the robotics models was language models debate is that the very very strong language models will have very similar representations of the world with very strong robotics models and if that's the case then if you have a really strong language model you also have a really strong robotics model using if we believe in the platonic representation hypothesis and if that's the true that's then the lesson that is the best manifestation of the bit of lesson right because you just need one really strong model regardless of architecture and they would be up performing any specific models that is slightly weaker. The interesting thing that that's maybe not entirely intuitive and maybe it'd be curious to hear you guys all think about is why specifically it seems like these newest model specifically asteris seems to be so much better at spatial intelligence. We've all seen the demos online of controlling blender and making great 3D images now it seems like you showed J in a benchmark that is like a meaningful step function improvement on certain tasks.
Where do you guys think the pre training and post training I guess of those models what is likely changed about open eyes approach there that made some so much better compared to models even six months earlier like you know that's still were benefiting from their lesson and we're still extremely good at code and we're still solving math problems but having quite crack the special intelligence needed here. I think what asteris does incredibly well is it's like vision to keep abilities it was probably pre train on way more computer use data than ever before sorry pre-trained so much like cad data and it's like all of these kinds of data probably teach the model like similar like understanding of like physical world as like a lot of robot data might. So the computer use data is an interesting point is not totally intuitive I think why computer use data is useful for understanding the physical world that it tends to be like a computer where you're clicking around like say more about why you think that is a big unlock is I totally by the train the way more computer you stay than ever before right I think immediately if you were just training a model for computer use that would not be apply for robotics but if you feed a computer use data into a big model like asteris and by computer use data I mean like you drag the cursor around on a screen to orbit some cat object in order to design and.
So I think that's why this data helps so much more making these lm so much better a robot use I think like from the other direction like a lot of people in a robot is community we're trying more more different kinds of data as well like egocentric videos kind of went from just tell the operation to like you know broader range of this data because it's not just robot data. I can teach a model how to use a robot so if we take this to the extreme it's like why not feed every kind of data coding computer use egocentric into the same model I think that's how we get to the most capable like robot use agent so funny that like you know if you go to nineteen eighties your ex park or whatever like we literally made the graphical user interface to be more like the physical world so that we can interface with it and like it we ended up doing is building an environment that was actually helpful for me. For robotics learn how to use the physical world right we have file systems we have like files we have folders we have you know the guise for like solid works and like auto desk to like spin around stuff to make it similar similar to the physical world and then like we couldn't get robots to work in the physical world so we just trade on that and now it works.
Right exactly there were incredible papers I think from Princeton and also copies robotics touches on this is like we just if you design a right harness where you make the tools that you face with the robot look like computer use tools like you have a agent drag cursor around to control what a robot goes that improves how well to LLM is able to perform on these physical tasks where do you guys see you know given like a reasonable guesses as to where the base models are going to continue improving and your guys own investments in either evaluating these models. And also the models or building harnesses around them what do you think is going to be capabilities that would maybe now see as challenging to do but that are going to be increasingly possible or even trivial very few months from now like you know six months ago even the demonstrations that I've seen you guys posted your Twitter I think would have been kind of mind blowing to imagine from an all. There's some consensus within the frontier labs and also in the robotics foundation models companies that we will have general purpose robots within the next two years or even earlier and this is something that society is probably unaware of or even unprepared for and when we say general purpose robots we mean something like if you give any natural language instruction it can do what a competent teenager could do with their hands it's kind of like the chat GPT moment but for robotics in terms of
capabilities we can generalize to unseen task and unseen environments for you guys for the world's guys what does that mean for what you guys are building I think we're to want to take steps to get there in about two years time I think there's so many challenges that are very visible for example why love belong line talk about latency this is a big problem if you just have astra being in the loop thinking at every step this is really slow is not going to be economically like useful so how do you you know consolidate kind of first pass by astra like this is this context we're learning into like a faster skill or policy that you can then run repeatedly and incredibly high throughput I think it's the the mappings to humans will get more and more like this we're like the the optimal thing to do maybe to to put a new SAR state action the word back into context is very quick but then there needs to be some like of go to sleep for while almost everything that is intelligent sleeps like tell me an intelligent system that doesn't sleep right in some way and then
at during sleep compression happens the weird thing that happens from your hippocampus and shortwave ripples to both lobes and like there's weird you know pass from memories that were compressed throughout the day to train the weights and similarly maybe what the right thing to do is similar to dagger data segregation framework in our classic rl where you're going you're collecting a bunch of data and then you're somewhat reflecting on it and then you're using it to update your weight file and then you have distillation in from those experiences back into an updated weight file that maybe isn't ashtray but maybe is your own models sounds very dreamcoder and I think I want to say about that and I think I think that the next I think I'm going to ask but a lot of those specific tools. I think that people used to build to take dreamcoder for example right it had this library of skills and then during the sleep phase it basically refact everything into more compact representation and so on. I mean we're seeing sort of similar things just happen not as rigidly as before in the space of programs but now it's maybe like maybe you want to refactor traces maybe you want to refactor skills a lot of these robotic
how you organize them and things like that. So yeah, I think as water goes forward, like how you manage our growing contact of skills of deployment data, all that is going to be quite interesting. I think with that, I think this was an excellent discussion. Thank you so much Vincent, Tanming and Jay for being here. I'm very excited for all the incredible robotics advancements. I think we're gonna see over the next few years and I think you're totally right Jay, that I don't think broader society is totally aware of how much is coming. That's gonna be a very incredible few years to come. So thank you so much.
More episodes
More from Y Combinator Startup Podcast

The State of Startups in 2026
Y Combinator Startup Podcast

8 Ways To Improve Your Outbound Sales
Y Combinator Startup Podcast

Open Models Change The Economics of AI
Y Combinator Startup Podcast

Paul Graham On Startups, Ambition, and Great Founders
Y Combinator Startup Podcast