
About this episode
World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence.
At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world.
They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data.
Resources:
Follow Fei-Fei Li on X: https://x.com/drfeifei
Follow Justin Johnson on X: https://x.com/jcjohnss
Follow Ben Mildenhall on X: https://x.com/BenMildenhall
Follow Martin Casado on X: https://x.com/martin_casado
Learn more about Atlas: https://www.worldlabs.ai/blog/atlas
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get every episode summarized
Each time The a16z Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
378 searchable segments. Every word is indexed and playable.
Full transcript
The a16z Show — Fei Fei Li: The Race to Build World Models For AI. Machine-transcribed; use the interactive transcript above to jump the player to any line.
On the past two spatial intelligence generating pixels that are truly spatially contextualized and grounded, that is the very hard step that atlost stick. You know, LLM's are built on next token prediction. We've seen video models as being built on next-frame prediction. Atlas is really new view prediction. This is the real place where AI can actually unlock a ton of value for people in their process. For saying like 50-100x reduction, there's a famous shot in the first Matrix movie where Neo is falling down. Exactly. They had hundreds of cameras viewing that angle on a green screen. On Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. No one has ever seen this result. When you set out to do this, did you know what was going to work? I was pretty sure. Each time we made the model bigger and each time we trained it for longer, it got significantly better. Does that mean we're going to get 40 video? They're like, go walk around. Language models are built around predicting the next token. What happens when a model instead learns to predict the next view of the world? In this episode, our team Casado sits down with world labs co-founders Feifei Lee, Justin Johnson, and Ben Mildenhall to discuss Atlas and the broader challenge of building AI that can reason about physical space.
They explain how Atlas combines generation and 3D reconstruction, why the team chose new view prediction as its underlying primitive, and what they learned trying to scale an approach that hadn't been tested before. They also get into what's still missing, including richer dynamics and interaction, and how world models could eventually connect simulation with robotics and planning. Underlying it all is a bigger hypothesis. Could new view prediction play a similar role for spatial intelligence that next token prediction has played for language? So, a big day yesterday you launched a new frontier model, which has got an amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So, maybe Justin, you want to talk about what was launched yesterday, why it's significant. Yeah, so Atlas is our new next generation world model, but has three basic things that can generate, reconstruct, and simulate the world. So, within that, there's a couple different major capabilities. It has really good camera condition generation, so you can input an image together with the camera trajectory,
and steer the model and have a generate video frames, any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to 100 frames that are the views of the real world, and use those to reconstruct the real world, and that reconstruction can take the case either of novel a video flying through the space, or an explicit 3D reconstruction of the space. Then finally, it can be used for simulation. And for this, we show off these awesome bullet time videos, which got a lot of attention online, and then also robotic simulation. What's up in bullet time video? A bullet time video. This comes from the Matrix. There was a famous shot in the first Matrix movie where Neo was like, I like that. Oh, yeah, that's amazing. So, and then remember in that flame, the shot, he's falling down, it's in slow motion, and the camera flies all the way around. The way that they did that shot is they had a ring of hundreds of cameras. So then he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that famous shot in Matrix. But now with Atlas, we can do this with as few as three cameras. So no studio capture, no green screen, no expensive calibration. We can literally stick three iPhones on dry pods, use these to take sort of a video of something happening, like someone shooting a basket,
someone dropping a strawberry into a bowl of milk. And then from those three iPhone videos, we can then refraining the shot, and imagine like freeze time, have the camera fly in, as the milk is splashing up, and get these amazing frozen time views. And we can do this with just a couple cameras. What is the simplest description of what Atlas does, what goes in and what comes out. Yeah, so one of the really core principles of Atlas, the most fundamental thing is it does new view prediction. And this is a really fundamental primitive that we think is super exciting, a super new primitive for base models that no one's ever done before. So we know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction, right, that given some number of views at a scene or a description of the scene, those go into what we call a spatial context that describes implicitly what is the world that we want to talk about. Then you can point a virtual camera at an arbitrary point in space and time. And Atlas will understand what that world is supposed to look like from that position in space and time. Then that video models out there all claiming to be world models and all claiming to have novel views. And can you maybe tease apart kind of more concretely how this is different from the myriad models that have come before.
Yeah, I think what Justin was saying about the spatial context aspect is super important here. So there's many video models, a lot of video models actually got their claim to fame from their single image input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omnoreconcing with 20 30 50 images. But what's key with Atlas is that it actually has a kind of like spatially grounded meaning to every frame you put into it. So it's not just an image that the model is going to interpret whatever way it wants or you can kind of try to argue with it and the text prompting and get it to do something specific with Atlas. Every image actually has an associated three dimensional camera pose and that means that you can perform this task of reconstruction with an extremely high degree of precision right. So if we had four views of this room one at each corner, you can put those into the model and then get an exact replication of everything you see in this room. So it's not going to guess what's in the other corner like the relationship between things it's just going to reproduce exactly what you give it. And you can also do that in a kind of creative or imaginative sense too. If you take two photos from different AI generations or real world locations, you can actually position and stage those to build these kind of intentionally directed fly throughs that are really governed by exactly the precise place that you put the content you want and where the camera is going to look and travel.
So it's going to be very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models. Is this just kind of an obvious scaled up version of a traditional video model or is it a new architecture? I think it's a pretty new thing for a couple different reasons. So we talk about is it does both generation and reconstruction jointly in the same model like Ben was saying this thing can take a couple of views of this room and then reconstruct everything in this room exactly as you see it. And historically reconstruction has been its own subfield and computer vision with its own specialized task is own specialized models and generation is what all the text video models are really good at like with all the big diffusion models we've seen the last couple years. And those are great for creative applications I want to imagine something that's never been there before but now with at list for the first time we're putting these two different parts of visual intelligence together in one model so can do both 3d reconstruction and generation together in one architecture. So to do that we had to make a couple changes one is we had to make it multimodal from the start so this thing natively works on text it works on images it works on videos it also works on camera poses as a native input to the model which I don't think anyone's ever done to pre training phase before.
And it uses 3d as a native modality that it works on so this thing from the beginning was designed to be natively multimodal in a way that no one else. Sorry, I just have I don't know the space super well by 3d is this like depth or yeah models are like what does that mean yeah so the formulation we use so far is depth maps. So right now when you have a frame that has a virtual camera telling its position in 3d space that camera position and camera parameters are a native input to the model and then what attach to that camera position you can have both like RGB telling you what is that position in space look like and you can have a depth map that tells you what is the spatial structure of that position in 3d space. So then text image video 3d cameras are these modalities that this thing all does jointly in a multimodal way. I want to add something to this I think what Justin just said is actually so important and also what Ben said that it's under appreciated it's the first time we have a unification of pixel generation and pixel reconstruction in the world of computer vision this field has been around for more than half a century.
City here having been in this field for decades I cannot tell you how many PhD thesis have been written on the problem of reconstruction or novel view things synthesis and also our field traditionally have multiple tracks you go to a computer vision conference you have the pixel generation track you have some recognition track and you have 3d reconstruction track. This is a elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and the viewpoint estimation and that's just incredibly powerful can we take a step back and then maybe you just feel something out so when you start with the company I remember you saying you want to tackle spatial intelligence right and now we have this new model. And so as a lay person feels very general to me that next particular prediction and this is next new view prediction new prediction right so like you can get one virus out of the years needed a new view can you pencil out like how this is a significant step to this general problem spatial intelligence maybe by starting to describe.
What spatial intelligence is or spatial intelligence eventually must enable us to both generate what the space is reason within it and being able to edit and interact within it now we can argue is it 3d or 4d ultimately it's 4d with the time dimension but even just 3d these are the fundamental tasks that one has to do or spatial intelligence has to enable and then we talk about with that you can render. You can simulate and you can plan actions but to do that a fundamental problem to solve is to understand the geometry and structure in the physics of the space and I do believe atlas is a significant step forward because now with every single frame you can generate an estimated important piece of information which is the viewpoint the camera pose and that is the most critical information one needs about the geometry.
Of the space and that can lead to all the emergent behaviors we see in the downstream of the model which we showed in the blog so on the path to spatial intelligence generating pixels is definitely a early step which we have seen with what you call. Gisilians of models but generating pixels that are truly spatially contextualized and grounded is absolutely another major step and that is the very hard step that atlas has taken. We definitely have we can just keep going here right like there is the fourth dimension of time which will bring in dynamics and there is more higher fatalities simulation and delineation of the space so this is part of the road map of spatial intelligence. Great yeah I mean I definitely want to like dig into like where this is going but first maybe let's talk about getting here how long is world that's been to existence to have to have to have to yeah and so you actually release models before so what why did you just jump right to atlas.
It's so magic right yeah just as team needs a lot of chips yeah you need a lot of to use actually skills so what what last year we released our marble world model that was the first kind of big major world model that we put out on that powers our car marble product and marble is really cool marble can take images it can take videos it can take text promise and use these to generate 3D worlds. The biggest differences between marble and atlas is exactly what is that output reality so marble was really focused on gouching splats is an output representation so whatever you're inputting on it's going to output a 3D world represented as a gouching splats and gouching splats are really useful right they're really nice they're easy to render they are they can they can render efficiently on mobile devices on the art devices they can interoperate with other with a game engines simulation engines is a lot of nice things about gouching splats but you know that I think that was kind of a bottleneck in the previous marble model. So we did with atlas is redesign the thing a little bit and we realize that we need to by fear Kate these modalities earlier and actually have these things all these modalities working in a more unified way in the model.
So now with atlas the fundamental primitive is not like a good generate a gouching splat world the fundamental primitive is as we said new view prediction. And that can generate RGB frames that can generate 3D and we can use those to generate a beautiful gouching splat worlds when you need them but we don't need to bottleneck our outputs through the gouching splats when we don't need to and that was a that actually took a lot of you know blood sweat and tears to understand like what are all the pros and cons of these different representations that that's one part of it. The other part is you got to like climb the scaling ladder right you got to like work your way up and like do smaller experiments do smaller models like to build your conviction on what's going to work and what. What's going to scale and there's there's the only if you could instantly know the right thing that's going to scale you know you should just do that but. When we started to look at the world was a very different place like there's no scaling law of spatial intelligence right so like when we started the company like the world is in a very different place the tech was in a very different place we had a lot of ambitions about where we wanted to go but it took a couple it took a couple iterations for us to hit upon this formulation that we thought is actually is actually like this is the one this is the one that can scale up.
You know been you know being the creator of nerf and doing a lot of 3D and reconstruction so it's not so obvious to me that like if you have multiple views that you actually end up with a 3D thing but like you've kind of like made a career of any up with the 3D thing so maybe talk a little bit about like kind of that step yeah yeah. Yeah I mean as you said I spent many many years of my career as a mass majority my career actually working on producing 3D things from images. And this is actually something we talked a lot early on in the company even if like is this going to be the approach to produce a 3D right are we going to synthesize multiple views and then build 3 out of that are we going to try to go direct to 3D like there's been a lot of uncertainty in the field around like which of those approaches kind of will we now are will kind of like read the best advantages earlier on. But I did have a lot of conviction just from seeing the kind of power of what I would almost call the brute force scaling scaling in a very very very small baby scale not like real model scaling but the scaling of dense reconstruction that we had seen happening over the past three years before.
So basically we what is density construction yeah dense yeah. Yeah. Because I know we're going to talk about sparse and I want to make sure that the builder stand what is dense and what is the sport. Yeah so this is actually even on the kind of like distance and commercial side I think been one of the challenges of productizing 3D reconstruction technology like at a fundamental level right people kind of don't. In a in a casual sense like you think I took three photos of this object right took six photos of this room like I look at the photos I can understand in my mind like how those piece together I can kind of fill in the gaps and get it but there's just never been really any kind of reconciliation between. Those like really data driven priors and then the kind of brute force density construction which it actually is much more akin to almost like scientific or medical imaging what we did and density construction right you basically have to say. Every single thing I want to appear in this reconstruction I need at least three or four views of it and if you think about that like even just in this room right there's like under the microphone under the table between every different crack and crevice and the plant is right to actually truly get a picture that covers every.
One of those spots it's just like very tedious and exhaustive effort to walk around the room I think you've all seen me running around various places like capturing them it takes you know for for someone who's. Well trained per se like it can take minutes but if you hand a casual consumer or even some like kind of professional trying to do this for the first time an average. Like cell phone camera capture device like it's going to take them probably an hour I've seen someone for the first time trying to scan a multi room environment spend like two hours walking through. Again of coverage and that's just this like very very exhausted tedious loop and say I want me say dance we really mean that it's like this room. We just been like lots of so many photos right I want like hundred two hundred three hundred photos of this room to capture it and what we're trying to do is bring that down to like. Sorry right well we're saying like 50 100 extra adoption and then that's at that scale where it just completely like flips that calculus on its head of like what type of captures you reconstruct you can go back to existing imagery you have you can go to. Stuff you find on the internet and even build seems out of that you can go to casual videos and like kind of inertial lot of footage in the past we never have treated as reconstructable.
I go back and like bring it to like this really potentially this is something we've been playing around with a lot with Atlas right like taking old clips like I take a bunch of my own old captures that never worked before and then put them to the system and I kind of seen a reconstruction for the first time or taking my old captures and thrown away 95% of the photos I took. And you know imagine angles that I never would have gotten from a traditional kind of like nerfer splat type reconstruction one thing that's under appreciated on the website of the demos is the Stanford demo where. Ben showed anywhere between three to 25 images you can reconstruct that entire Stanford quad but the thing is we had to show it from a real view. But every single input image is been standing on the ground taking a picture from the ground so everything you see are generated but according to the laws of reconstruction and this is really magical. And this is where like generation and reconstruction need to interplay and a really fundamental way to solve this problem because under the classics kind of reconstruction stuff that Ben was talking about like the reason you need so many views is because.
I need like multiple images and I need to triangulate this point in three space and see it from multiple viewpoints so that means like that's required in the traditional version and on the flip side anything that wasn't captured in these views like any pixel that was not visible in one of the input views will be a whole in the three to the structure because fundamentally like if a thing wasn't visible in the input views you know you need to imagine it to fill in the gaps and that's fundamentally a generative process so even in this room even if we set Ben loose with a DSLR and like let him like capture like hundreds of views of this room even the world expert on doing these dense captures is still going to miss some spots like he's not going to get like underneath all of the microphones or underneath all the tables are like in between all the chair likes you're always going to miss something no matter how many views you get so that that's where you need generation as another mechanism. In the model because you're never going to get everything so you need to have some generative capacity for the model to imagine based on what I'm seeing then like first triangulate what I what I can see but then fill in the gaps of the stuff that inevitably inevitably was not capture.
Yeah and there's something like super cool about this that LLAM to really understood this for a long time right there was kind of on almost these like context wars like the first of the years I was like oh we got to 128 to 26 to like 5 to 5 we got a million right and everyone kind of understands now it a pretty tangible level to value if you know you crank your context like the high when you're using your coding model it's a hard problem like everyone has a feel for that but like no one has pushed that at all on the image and video most I didn't the same kind of like principled way like no one's out there trying to like put a like an hour long video through into a needle in the haystack retrieval of like a frame at the 3037 minute mark. Whereas with reconstruction generation you actually have the same exact thing like reconstruction is just like generation with a really long context and you put a lot of stuff in it. Yeah right like that's the way to actually build this continuum where you kind of bridge between those two things and like atlas like being able to like this is we can ever do with marble marble had this kind of fundamental blocker of like you couldn't really jam more than honestly like a couple images in. But at least I can go on I can actually take like a 64 image capture and do like a fly through an entire house and everything is grounded by being you know seen or like almost seen or like slightly extrapolated from what's not there.
But you're just getting these you know I'm taking captures I did with 2000 images of a multi room house and taking it down to like 3040 inputs and the fly through looks like basically the same and this is just like totally inconceivable before and it's all enabled by building this gracefully scaling kind of context window that you can dump stuff into. And so the way to think about it is like the sparseness or the pictures that you physically took. And then atlas is a model creates the rest of the views and use classic reconstruction techniques that roughly the way to think about it or in some sense yeah yeah I mean that's a video of atlas is like you can take however many inputs you have down to like a single view and then you can always use atlas is this this rendering engine to produce anything else you're on right you can you can navigate. Like the virtual camera exactly you can say like okay I have a picture here I have a picture there there there you can make a couple of those then you can stage a dense fly through you can do this in sequence because it's an order of massive model it's up to you right to kind of pick and choose what you add interactively into the context as you generate. I mean the thing that I just blows my mind is I just have a very simple number not the model I I four pictures and then I've got to like have the model extrapolate between them and then it has to fit when you're
like it's got to be 3D like and I will think of these diffusion models is like being visually great but not accurate and so like In other things though there's a question here but like how is like the room fits like how is it that it's 3D consistent is it just lots of data or you know I mean it's a part partially it's a belief in the scaling of offices right like you know did you by the way I have to ask when you set out to do this did you know I was going to work I was pretty sure I think three of us have total conviction about the skin in all that that I think we do I do think the exact architecture choices and data mixtures is where the devil's are in the details I you know have watched Justin and his team going from we really don't know how long this is going to take to maybe sign of life to while this is going to work so
No one what no one has done it but I think the hypothesis two hypothesis one is scaling all hypothesis the other one is next the viewpoint prediction we had And comics and of these two things are really out so I think I was very convicted that it was going to work I was not sure it was going to work this well at this fast right like I thought there's a chance that we do this maybe it's not clear that like the first cycle of pre training a new model with a new architecture a new paradigm like the first cycle of that working is insane so I thought there was a chance in which we had to we might have had to do a couple more turns of that of that pre training cycle before we got to the level of quality we were so we wanted is it is Are we kind of like at the end of like the scaling for this architect for approach when you know the breakthrough or is there no no we're at the beginning really Yeah, without changing the architecture yeah, we're basically the beginning and we're basically limited by compute at this point All right, like data is very important is faith a likes to point out but like everything has a bottleneck and I think the main bottleneck on continuing to scale this thing is actually training compute right like during development we trained a sequence of models we
Read up at this in the blog post a little bit, but we trained a couple models that like the first couple of wrongs of the scaling ladder Um and each time we made the model bigger and each time we trained it for longer each time we put it on more chips like it got Significantly better and the model of size that we like the model that we showed in the blog post is obviously the biggest and best one that we trained But the thing that was limiting it was not the scale or the data or anything like that it was literally like we had a deadline of when we wanted to release this thing and therefore we backed up what we could afford to train in that deadline But here's a little bit of an insider story right like Justin and team are are training from the smaller and the slightly bigger you know are training having these road maps And then there was one day in summer early summer that it's not even this the current Atlas Uh, model size it's a smaller model and then Ben Justin Ben feed it into you know the viewpoint generation And remember that famous table the garden table for nerf paper and many papers that
Over and I I got a slack I mean we also the slack from Ben that The fire to the floor through under the table with the soccer ball Yes, with the soccer ball emergent or is that in the real right? Okay, that morning the three of us looked at each other in the eyes and say that's it This is we're gonna build this like we made a decision within five seconds. How this is just say No one has ever seen this result then um Can you talk through maybe more specifically the use cases that so so Uh world lapses historically had a lot of users that were creatives and they use it for like consistency and you know like Whenever 2d images and for movies and for 3d and for games etc So maybe can you talk about how this extends use cases are catered to the existing ones and then we'll like talk about robotics actually Yeah, sure Yeah, I mean it's kind of funny actually one of the kind of main ways we even solve people using marvel plays exactly into this new view prediction case
Like a lot of our model being the previous sorry a marvel artist previous product Like people would take that product put an image in get a full 3d scene as the gouching spot Take a couple screenshots of it from different points of you leave I am we're like We can just So I think like you know there's a lot of degradation there They're like just flat could look better and it's like okay. What if we just generatively model those viewpoints with that exact Value control so I think like even that core capability of just like views synthesis Um, it's sort of been this academic problem for a long time They in the sense of oh, you're gonna do this really dense capture like like generative use synthesis is a relatively quite a new problem And we just see so many people who In this creative pipeline right people have a multi-stage work flow rate I don't think there's a single person out there using one modelistic model not even see the answer whatever for for their entire Task people will have this like you know kind of bunch of storyboards and mood boards of images They pull out from like their favorite collection of image models and then they'll go to different video tools I'm like build those together as keyframes
Then they'll go on like clip and edit those later right? So we were seeing this like sort of you know Nege but very specific use case for marble is just providing that like Sanity that you can ground your generations in some kind of 3D consistent world right people You know, I know I've I thought was with various image models to ask them to like give me different viewpoints of a room And every time you can just look and see all those things kind of moved around like it's not stable And like even that one seat of a use case I think kind of signals that there's this value and there's hiding under the surface there like there's just Decades of people being used to persistent 3d state like virtually modeling what they would be doing in the real world And having you know a stage and props and like elements there whether it is for A movie or a show or marketing show or like building out game environments like this this Statefulness and persistence is so key in how people think about spatial reasoning and like developing an environment over time Like people don't think in this a thermal like Generated thing journey thing like just throw it away Keep my text prompts like people want to build this like collection of assets and like model a world in that way
So we're trying to provide like again with with this facial context mechanism I know those things like we're trying to provide that level of control and precision And the ability to ingest different modalities of input starting with the post images But you know, we want to get people more control over the elements of the things in the scenes They're looking at and editing and interaction and all that as we go forward And I think that that it unlocks like further use cases in those areas where we're seeing but Also expanding out into kind of any place people want to create a virtual replication or or like you know a pre-imagination of a real world Space they need to build right for architecture and construction like I talked to a guy at some point building Booths for conferences right there's just so many things in the world you don't think about need to be fabricated And every single one of those basically goes through this like pretty painstaking virtual design phase And of that process like the part where you go into 3d software is kind of one of the most like Arduous and like labor intensive parts right now like like taking Feedback on a 3d design from kind of like verbal commentary or sketches or really really quick
Stuff you got from like a creative director or like a design director or an architect or whatever like mapping that back Into the 3d representation is like 95% of the work right you can have a meeting get feedback And then you go back into a weaker versions and that's just because like our software is kind of decades old at this point And it's just never became as intuitive as you know playing with Legos or like pottery or doing this stuff with your hands Or sketching with a pencil and this is the real place where AI can actually unlock a ton of value for people in their process Whether it's a creative application or something more industrial or design or or whatever And that like really motivates me to kind of build Different flavors of our model to cater so those kind of people Yeah, I can understand how it helps with the creatives because like marble did it and also how that extends to things like designer architecture But say say you acquired a robotics company. It's less we just talked about it too But it's less obvious to me especially in the context of the Atlas like how that maps to robotics So if you wouldn't mind dispensaling that out yeah actually Atlas is a key part of the puzzle so
um We acquired this company that was formerly known as Cnex and what is the key technology right now the key technology is a system that Goes from real to some and then some to real and what does that mean in robotics situation you want to train a robotic arm to you know figure out how to Do cabling let's say in a in a industrial setting Well, you need a whole bunch of data to first train a robotic policy to do these cable cables cabling activity And then you will evaluate if the robotic policy is doing a good job And then you deploy the robot into the cabling Environment yeah in order to train what you what this company Cnex and now our robotics team used to be doing is Do you exactly what Ben was saying dense reconstruction you take you know pictures of a situation And then and try to reconstruct that environment is excruciating the painful takes a long time
Liberious and it really blocks the velocity of robotic simulation real to send right so Atlas really is the next Generation technology for for that and this is not just for robotics cabling or anything we should zoom out and recognize The biggest problem right now in robotics is actually data one day it'll be chips, but for now is data because it's so hard to collect real world data where robots you know are operating on and and In order to Not only you need to Collect the data of let's say the cabling situation or dish washing situation or whatever There is also a very important step called randomization is that you have to take the same environment and then randomize the condition So the cable doesn't literally only you know bend this way can bend a different way or the box can have different sizes colors
Different lids and all that and and or in different parts of the scene So you have to go through a real to some situation in order to get enough of that data in addition to other data You can get from internet. So this real to some step will be You know really helped by by Atlas. That's just the first part of this is meeting the the the robotics needs in the current technology Because we don't yet have a frontier foundation model that's robust enough for for robotics But asless is a omnimodel is a multi multi-modo model It takes a different kinds of input and generates different kind of output You can totally imagine the next step is at least taking in data that's in the in that's dynamical And that can really start to bridge the gap between you know action planning and uh and um
And robotics and and the Atlas output. So that's on that roadmap. Where are you gonna say? Yeah I was gonna say there's something fundamentally different about training or robotics policy compared to really any other application in AI We've seen before um and that's like if you're generating a piece of code like you're generating an image you're generating a video I'm the model is fundamentally creating this this artifact Yeah, and that artifact like there's a lot of examples of artifacts that you can go out on the web or somewhere and collect Right, you want to generate images. There's a lot of images out there. You want to generate videos. There's a lot of videos out there You want to generate a code base. There's a lot of code bases out there. You could art from yeah a robotics policy is something fundamentally different It's not producing a static thing It's instead a policy that's gonna go out into the world make actions and like try to achieve a goal And the world's not always gonna respond the way you expect right unexpected stuff is gonna happen So robotics policy is like fundamentally an agent that is out in the real world interacting with the real world and stuff happens So you need like a critical part of that is the those those policies during training need to be exposed to every possible thing that could go wrong during their deployment
Um, and that's where simulation is really key for robotics right so there's the then there's two angles on that like one is the kind of classical simulation You can go out you can go and like go to your favorite physics engine and like try to imagine creatively as a human designer What are all the scenarios that might happen in this in this in this uh when achieving this task and then try to write explicit code that models them all That that's one angle um, and that's an interesting angle with coding agents like that actually gives supercharged too But there's another angle which is try to more data driven simulation right like maybe we can't can we have a learned model that can Understand how the environment how the world is gonna respond to actions and maybe right might respond in unexpected ways sometimes Then could we build these neural simulators that are trained on as much data as we can then use these neural simulators these learned neural simulators You know as a simulation bed to train robotic policies Um, and that's that's a really interesting future direction battles But but then it doesn't stop there right So um, but once you have you know is learn simulator like this learn simulator
Kind of already has in its like mental brain like it understands the world it understands how the world is gonna respond to actions and Why doesn't the simulator itself become the plan? Yeah, right like the same and that's kind of the core thesis that we've had are on world models and their generality that There's some core stuff that a model should understand around generating worlds simulating them understanding how they appear in different search situations and You know understanding how the world's gonna respond to an action is highly related to Imagining what kind of action I need to take to make the world respond in a particular way. Yep what um Well piece of feedback that I got I've been by the way Congrats on the launch it was overrunly positive I think it was probably the most significant model once this year and yeah I've just said glowing things for one person who's an expert in the space who I texted was like what do you think? As great person is fantastic. It's amazing, but there needs to be more dynamics And so it seemed to be at least the robotics case, but generally it was kind of ideal to actually have a world that moves and so
Maybe talk a little bit about that and then any other future directions that a you're comfortable sharing But you think we're talking through yeah, I mean like dynamics is clearly gonna happen like actually What have baby done that we actually do have baby dynamics already and this is something I think people didn't quite appreciate Yeah, I didn't really highlight in the blog post, but like the previous marble world model It was like fundamentally static Yeah, like the the model just like could not handle any dynamics at all And that was just like baked into the model architecture baked into the training like the whole thing was fundamentally static Um, we already knew that was a big problem post marble and we already fixed it in Atlas right like yeah The Atlas architecture is already fundamentally supports dynamics Um, the Atlas training data fundamentally has dynamics And if you look carefully in some of the videos that we've posted actually Yeah, the way it's the waterway So like some of the examples there's like waves in the water like in some of the like air generated aerial views There's like little cars moving around So like dynamics is actually already in this model, but it's a by the way Dynamic seems very problematic to me if you're trying to reconstruct 3D for multiple views
It is like are these things like ads or no so it's actually one of our feces here is that like you know If you're gonna do fundamental 3D reconstruction You actually have one have no dynamics like you want to be able to model like exact views of the scene with exact frozen time But then like this is actually kind of a problem with our previous marble approach right like they're like You can try to find data that's fully static But that's really hard to scale and really hard to get more of and the thing we realized is that even in the case where I want Static output in the end the best way to get it is actually expose the model to dynamics Right like expose the model to as much dynamic stuff as you got as much static stuff as you got and let the model figure out how to factor out the dynamic stuff So again, especially in like the go interesting This is like and this is actually so again in the Atlas pre-training already like it it saw a ton of dynamics in the pre-training already Then the post training that we did specific to this checkpoint and this release was focused a lot more on on static stuff Focused a lot more on spatial movement and less on temporal But like we already had we already like I'm pretty sure this the the pre-trained checkpoint already has a lot of latent dynamics in it
And this is something we're gonna improve quite a lot going forward. So Ben does that mean we're gonna get 40 video You like go walk around You can see the smile on their face So I actually can see if like you just stopped now and you only did kind of bigger You know faster better you could build almost an entire industry like it feels like a very horizontal primitive and if you did nothing else But are there other things that are not just kind of bigger faster that you're excited about the applications You're focused on which tend to be kind of more on the kind of content created 3d side Yeah, I'm really excited about pushing that kind of multimodal aspect I think different modes of control is so critical here. I think like it's super underappreciated Especially in the academic community how critical it is to add control conditioning to these models to kind of get out what's inside I mean honestly this dynamic personality I'm not sure I even understand what those are Let's just leave layman's language is Editability yeah, I think editability is is the key here Yeah, so I mean this is only we seen in like sort of single image models And starting this year in video models is starting to be unlocked in terms of oh like
Getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person in this object And this thing to happen and kind of combining those all together in like one pastiche without having to do a lot of like manual work But the system like it just interprets it like kind of frontier Image models are kind of there right for in terms of editing But we haven't seen that propagate out as strongly into video yet and then into world models right We've seen some really kind of toy examples of oh, I can like put in a sentence and like you know a dinosaur appears or something with these like sort of real-time models But I want to like turn that up to really industrial strength and make that because like the trick here is You've got to add control but not compromise equality with the model or just becomes a party trick basically like it's like No one is going to seriously think about swapping there like cutting edge frontier video model usage for your model If you give them extra knobs, but the quality degrades So I think it's really that game of like how can we maintain like the high bar? We've set was the outputs were able to get in the current model and then add all kinds of interesting stuff that people ask us for in terms of like
I want to interact with the scene or control the layout or control like the identity of the objects and the things that we're seeing within their or control time Right And I think that's like an access where it opens up like a ton of really interesting product and interface work The more complexity you add their enrichness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful You know 3d worlds in the computer like that's that's really the end goal here Is getting like all the capabilities you need to build that kind of system awesome Anything you'd add to that as far as new functionality that you'd be excited about that's not just bigger better I think for me let's go back to the first principle of intelligence Intelligence is not sitting there stack and just seeing something or interpreting something when it comes to space and physical space It's really this Closing the loop between seeing and experiencing and interactions So just thinking about going off that ladder is exactly what Ben said Oh, I think one interesting motion there is this notion of AI completeness. You read this report? Yeah, yeah
So like everyone I hear about AI complete by the way in terms of LLM Which is like you have to be basically You know like the smartest LLM to answer the question what the smartest LLM will need to answer or you have to solve general intelligence No, no, it's obviously it's a it's a connection to Turing completeness right like the idea of being that like A task is Turing complete like in classical complexity theory if like I can take any class at any problem this category Reduced to that one problem Yeah, yeah, yeah, yeah, yeah, yeah, right yeah, so you can take any NP hard problem and reduce it to 3 sat Therefore you can use it to solve any any problem. Yes Yeah, yeah, so then like that the kind of like soft definition of AI completeness is like there's this fundamental primitive That's let's an AI task But if I could solve this AI task in its full broadest generality you saw that it would solve any intelligence Yeah, and like the classic example at LLM's is like next token prediction is a complete because I could like You know, there's the classic example I think from Ilya where like I there's a mystery novel and like the thing has to read the whole mystery novel and the final sentence of the mystery novel is like And the killer was a pretty good next
So like you could basically like frame any kind of intelligence task in terms of that So clearly next token prediction is something that people believe is AI complete But I think there's something we're kind of realizing and Ben was talking about this earlier today is like new view prediction This primitive that we have an atlas especially generative new new view new view prediction. Yeah, this is also AI complete Right and because I could take something like you can you can have the the movie and you do all the frames of the movie And then like the killer walks out and then you predict a shack to watch out I want to have a world where like martine is like writing a proof of the rewind I have on You have to answer the next one and he saw So I mean so to take an evolutionary view right that new viewpoint prediction is exactly evolution had to solve by making animals move You nature give animals eyes But later didn't give trees eye Eyes why because when you move you see a new viewpoint
And that is the the the weather you call it AI complete or intelligence complete So so we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction amazing Well with that congratulate since all of you in a phenomenal model launch we're very excited for future model launches and thanks for coming Thank you Thanks for listening to this episode of the a60t podcast If you like this episode be sure to like comment subscribe leave us a rating or review and share it with your friends and family For more episodes go to youtube apple podcast and Spotify follow us on x by a16z and subscribe to our sub stack at a16z.substack.com Thanks again for listening and I'll see you with the next episode As a reminder the content here is for informational purposes only Should not be taken as legal business tax or investment advice or be used to evaluate any investment or security And is not directed at any investors or potential investors in any a16z fund
Please note that a16z and its affiliates may also maintain investments in the companies discussed in this podcast For more details including a link to our investments please see a16z.com forward slash disclosures You
More episodes
More from The a16z Show

How AI Is Rewriting the Power Law of Venture Capital
The a16z Show

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
The a16z Show

OpenAI Researchers on the Future of Mathematical Reasoning
The a16z Show

Can Open Source Keep AI Power From Concentrating?
The a16z Show