
AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
About this episode
Machine Learning Street Talk (MLST) is made possible by:
Could slowing AI development make superintelligence safer? Daniel Kokotajlo and Thomas Larsen of the AI Futures Project join Tim Scarfe to examine AI 2040: Plan A, a proposal to buy time before AI exceeds human control.
SPONSOR:
---
Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.
Apply now: https://cyber.fund
---
After revisiting AI 2027 and the limits of forecasting, they ask what happens when AI can automate research and sustain an economy without human workers. Tim challenges the case for general models and asks whether intelligence alone explains power. Plan A proposes an initial pause to build safety infrastructure, then cautious development up to the strongest AI that can still be reliably controlled. The discussion tests the distinction between control and alignment, the case for public AI research, and whether the US and China could enforce a slowdown. It ends with the evidence that would change their forecasts.
---
TIMESTAMPS:
00:00:00 AI 2040: a slower route to superintelligence
00:01:34 Sponsor: Cyber Fund
00:02:12 From OpenAI to AI 2027
00:06:58 Forecasts, war games and self-fulfilling prophecies
00:17:44 Why AI sceptics are changing their minds
00:23:04 When AI can replace its own researchers
00:28:45 Could an AI economy grow without human workers?
00:37:32 One general model or a society of specialists?
00:47:43 Brains, machines and collective intelligence
00:56:12 Plan A: buy time at the controllable frontier
01:00:02 Why control buys time but cannot replace alignment
01:06:36 Why AI research should be public
01:10:32 Can the US and China enforce an AI slowdown?
01:19:04 Why AI policy debates miss the technology
01:21:56 Is AI normal technology? The remaining disagreement
Many thanks to James Wilken-Smith for helping with show research.
---
REFERENCES:
other:
[00:00:01] AI 2040: Plan A
https://ai-2040.com/
[00:03:27] AI 2027
https://ai-2027.com/
[00:13:47] Scenario Scrutiny for AI Policy
https://blog.aifutures.org/p/scenario-scrutiny-for-ai-policy
[00:33:11] The 2028 Global Intelligence Crisis
https://www.citriniresearch.com/p/2028gic
[01:00:40] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
https://www.redwoodresearch.org/research/hugging-face-incident
[01:09:21] The Hugging Face incident and the road ahead
https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[01:22:01] AI as Normal Technology
https://www.normaltech.ai/p/ai-as-normal-technology
[01:22:51] Common Ground between AI 2027 & AI as Normal Technology
https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and
person:
[00:19:43] Geoffrey Hinton
https://www.cs.toronto.edu/~hinton/
[00:20:07] Ryan Greenblatt
https://www.lesswrong.com/users/ryan_greenblatt
[00:26:06] Elon Musk
https://www.tesla.com/elon-musk
tool:
[00:21:46] ARC-AGI-3
https://arcprize.org/arc-agi/3
[00:21:53] AlphaGo and Move 37
https://deepmind.google/research/alphago/
[00:39:41] Claude
https://claude.com/product/overview
[00:39:58] NVIDIA H100 GPU
https://www.nvidia.com/en-us/data-center/h100/
paper:
[00:24:42] Training AI Scientists to Replicate Research
https://arxiv.org/abs/2608.13331v1
[01:27:19] Validity of the single processor approach to achieving large scale computing capabilities
https://www.cs.cmu.edu/~18742/papers/Amdahl1967.pdf
book:
[00:28:52] Bullshit Jobs: A Theory
https://www.simonandschuster.com/books/Bullshit-Jobs/David-Graeber/9781501143335
organization:
[01:05:09] Redwood Research
https://www.redwoodresearch.org/
---
RESCRIPT:
https://app.rescript.info/public/share/33d1a58fa8f307ae7dfd504d4fdaa9d5
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,921 searchable segments. Every word is indexed and playable.
Full transcript
Machine Learning Street Talk (MLST) — AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen. Machine-transcribed; use the interactive transcript above to jump the player to any line.
This episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome? That's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50-page restoration block, or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it. Ready to make anything online makes sense? There's no place like Chrome. Check responses set up require compatibility and availability varies 18 plus. Did you know that mosquitoes have killed almost half of all people who have ever lived? Today people are fighting back. With support from the Gates Foundation, American scientists and partners around the world have developed a new generation of bed nets that can kill up to 90% of mosquitoes exposed to them, helping cut malaria rates in half. The result? A safer, healthier future for everyone. The Gates Foundation, partners of human potential. We're going to talk about AI2040 Plan A, which is our new scenario in which they build
super intelligence in 2040 because they go slow and pace the frontier. Have you noticed this five shift? Yes. And I'm very happy. Is AI more like the electricity or airplanes or is AI more like humans in the cloud? The point at which an AI company would rather fire their humans than fire their eyes? Trip minds, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chipfabs and so forth. That whole thing can just be doubling every year, every six months, every three months, faster and faster as the technology improves because of course the AI will also be researching to improve the technology. Who knows what's going on in the rest of the world, but Anthropic has disassembled the moon, for example. Hypothetically, if an AI CEO was saying that their model was truth-seeking, I would only say the truth. But actually the model was looking up that CEO's political opinions before answering.
We build AI's, they're aligned to humanity, but to who? Is it the president? Is it the CEO? Is it some actually broad and good democratic process that aggregates everyone's values in sort of an endorsed way? You're probably not going to be the last one and so you know... Don't hyper-sition that. Then that's it for me. I think that's it. That's not a normal technology. Just stopping everything now, plan S, would be better than the default. Like I would rather just stop everything now, then continue going on our current trajectory. This episode is supported by Cyber Fund. If you're building at the front here of AI, they want to hear from you. Cyber Fund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams two million dollars each to participate.
Apply now at cyber.fund. So I'm Thomas Larson. I work at the AI futures project along with Daniel here. I was lead author on this project, Plan A, AI2040, which came out a few weeks ago. And I was also a co-author on AI2037. And I'm Daniel Hubertello. I run the AI futures project and co-authored both these reports. Awesome. So yeah, the context of this conversation is that there's this 2040 Plan A, which will get into a lot of detail on. But I'm just interested before we get there. I mean, can you just tell me a bit more about the AI futures project? I mean, how did it all come about? Yeah. So I used to work at OpenAI. And while I was there, I did a variety of different things, eVALs, forecasting, governance memos. And I became gradually disillusioned with the leadership of the company and also the gap between how much information there is inside the industry and how much information there
is outside. And what you're allowed to say on the inside versus what you'd want to say, this is just a big gap. And so when I left OpenAI, I wanted to be able to speak more freely and tell the world about what people on the inside see coming, basically. And AI2027 and the AI futures project was our first project that we did along those lines. So I recruited a bunch of people to help me. And we wrote this scenario called AI2027. And it was a similar sort of thing to what I had done internally at OpenAI, but just much bigger and more ambitious and free for the whole world to see. Very cool. And maybe we should just have a quick refresher on AI2027. So this was an absolutely huge event. Many, many folks were talking about it. I guess like, did it achieve what you wanted it to achieve or what did you want to achieve with it? Yes. More so than expected. So the first goal that Thomas would remember, when we were working on it, the first goal was just a purely epistemic goal of the future is crazy and hard to predict.
Let's try our best to predict it. And let's see how well we can do. So let's game out a concrete scenario. And even just for our own edification, like we learned a lot from this whole exercise, and we feel like we had a better understanding of what was coming. And then the secondary goal is lots of people seeing it and being informed by it and starting conversations and so forth. And that part blew past our expectations. We made predictions beforehand about how many people would read it and it was 90th percentile outcome. The thing I would add on the epistemic point is that I think things have been going more on track for AI2027 than I would have predicted at the time we released it. So like at the time we released it out of assumed that reality would have diverged much further from our scenario than it has already. I think like sort of the real world impacts, like the revenue trends, for example, but also various other trends, I think are pretty close to on track for AI2027, which has surprised me in a bad way. Interesting.
Yeah, maybe we can reflect on that because I think you guys did a self-assessment on AI27. And it was something like 65 to 75 percent of it was on track. AI software R&D uplift was only 0.17 percent. Can you explain that? Yeah, so we've done two different blog posts where we take all the quantitative predictions made in AI2027 that have resolved so far and compare them to reality. And we track this metric of how much of the distance has been crossed by reality compared to how much has crossed in the scenario. And in that way, we can get this overall sense of how fast our thing is going compared to in the scenario. And the top line number is something like 75 percent speed. So basically things are on track, but just going a little bit slower. The uplift number, I forget what it was that you just cited, that was actually, basically at the time that we wrote AI2027, we had a bad estimate of what the uplift was at the
time that we published. We thought it was higher than it actually was. And so what actually happened is that there was actually significant increase in uplift due to coding agents and so forth, but it was increasing from a lower level than we thought up to the level that we thought and then a bit above. And so the metric looked like it was only a small amount of progress because the metric was tracking from where we thought it was to where it is. But does that make sense? Basically because we had it overestimated the metric at the beginning, it overall makes it appear like there's been less progress according to this particular metric that we're using. But it was because we had, yeah. I mean, can you tell me a little bit about forecasting in general? So I guess there's like a bullish take and there's a bear take on this. So my intuition is that reality is infinitely complicated. There's just these infinitely diverging trajectories and God knows what's going to happen the day after tomorrow. But by the same token though, reality is quite structured.
It's quite convergent and it is indeed possible to predict things that are going to happen because you know, certain things re-acquire with increasing regularity. So would you guys classify yourselves as forecasters? I mean, can you talk me through that? Yeah, so I would think, yeah, I think forecasting is a good name for what we do. The way I like to think about sort of why we're doing what we're doing is sort of like, it's sort of like why people that are fighting wars do wargaming, where you're never going to sort of predict the exact sequence of battles and exact sequence to sort of how your war will go at the beginning because it's just going to be really complicated. There's going to be enemy action. Things are just not going to go as you expect. There's just no way. But if you have sort of no concept of how your initial plans might result in victory, it's very unlikely that you'll actually succeed. And so I sort of think of, like, A1 and A7 were sort of our attempt to just roll out.
Here's one way, sort of the AI future could go. Obviously, it's not going to go exactly like that, but it's going to be one concrete story that we can then diverge from. And then Plan A was trying to be basically that, except now we're saying what should sort of the US government do to manage that well. And then that was supposed to be the positive vision story. And it was again in the spirit of a war game trying to illustrate one possible concrete future path. And then, of course, things aren't going to actually go exactly like that. But having one viable plan that makes any sense at all is like, we hope a positive step forward relative to the previous state of abstract arguments in the void that are sort of not that tepid to reality. Yeah, that makes sense. It's certainly not abstract. I think it's very concrete. But I guess one thing that occurred to me is on the 27 piece, it was quite gloomy.
And on Plan A of 2040, it's far more optimistic. And you know, but it seems like a bit of a mixture of conditioned prediction and recommendation at the same time. I mean, where do you guys land on that? It is in fact a mixture of prediction and recommendation, unlike the 2027, which is a pure prediction. And I think if we could do it all over again, we might try to be more clear from the beginning about our structure of what's the prediction and what's the recommendation. Because it is, it's kind of mixed up, like some parts of it are predictions, some parts of it are recommendations. There's a supplement that you can go to on the website that talks about which parts are predictions and which parts are recommendations. But I understand that's not very easy. You're not very apparent to people. But broadly speaking, Plan A is the prediction part. So when the government implements Plan A and they make the deal with China, and there's all these pillars that they're upholding and so forth, that's our recommendation, not a prediction. And then usually most of the things that follow from that are predictions rather than recommendations. So mostly it's just rolling out like what we think the consequences would be if you implemented
Plan A. And then there's a few other things that I recommend to you to, for example, the citizens dividend. Yeah. Yeah. The other thing I would add is just, I think the thing we found is that it's very hard. Like when you're trying to make a recommendation, it's very hard to disentangle the predictive aspects and the recommendation aspects because all of your predictions are colored by recommendations and your recommendations are inherently trying to be at least vaguely realistic. Like if we made recommendations that were completely unrealistic and had no bearing on reality, but we'd have the less stood by and we're like, yes, we should do this, but we know it'll like absolutely zero percent never happen. Then that would have been a much less useful exercise than I think the one we did. Where the one we did was, yeah, we were mostly trying to make recommendations. We made some recommendations that we think are pretty unlikely to happen, you know, but we were trying to make like some substantial concessions to sort of like realism as well and trying to actually aim for something that we think is at least moderately viable. I think that if we could, if we could do another scenario like this, we'd probably have
a more clear structure of there'd be a central branch, which is the pure prediction branch, which just goes all the way to the end like, yeah, 207. And it's just like, here's what we actually are best guess. And then there'd be branches off of it that are like, at this point, they do this recommendation instead. And here's our recommendation. And then after that, it's just a prediction again of like what we think the consequences would be if you did this recommendation at this point. And in that way, and then maybe there'd be sub branches off of that. But then it would be sort of clear at every point that like, everything is a prediction, except for these particular branch points, which are recommendations, you know. Yeah. And the war games thing was really interesting, just to kind of dwell on this a little bit, because even if a war game is incorrect, there must be some kind of information gain from it. So if you do a whole bunch of war games, there must be abstract motifs that do appear. So I guess this is what you think that if we do these different scenarios, we're almost guaranteed to have some kind of uplift. Yeah, that's basically right. I mean, so the example I like to bring up is midway in particular where the Japanese,
before the Battle of Midway, did a bunch of war games. They did like a three day retreat where they like war games it out a bunch of times. And they kept losing. And then they would sort of break the game. They would like resurrect their aircraft carriers after they died. They would like re-roll the dice on like whether the Americans succeeded when they succeeded, so that they would sample until the Americans strikes failed. And so basically for our sake, that's like reality sort of like yelling to them through this mechanism of the war game. Like, hey, your plan is terrible. You're going to lose if you do it. And sort of us, you know, that's sort of the hope with sort of our, like we ourselves do a bunch of war games, but also do a bunch of sort of detailed scenario writing. Our hope is that every time we have to write a part of the scenario, and the part of that scenario seems super unrealistic or you know, isn't really well modeled and doesn't really make sense, or people are able to make really good criticisms of it online. That's basically reality yelling at us and trying to like help us see reason. And our hope is that we can sort of like do enough of this sort of put up enough surface
area so that we can sort of get that, you know, dose of reality from the real world, or from the simulation of the real world, which we hope is realistic enough to sort of accurately give us that information. If I can add to that. So we call this scenario scrutiny. So basically we think that if you have an ambitious plan for what to do in the future, you should try writing out concretely what it would look like to implement that plan and what the consequences would be. And this is a way of applying more scrutiny to your plan. It's a way of sort of like stress testing your plan because it's opening your plan up to more like criticism basically. Also, we've in addition to doing our actual scenarios, we actually do literal war games where we get 10 people in a room for four hours and we game out a scenario like this. And we've done maybe about a hundred of them total mostly AI 2027 style war games, but also about ten or so plan A war games where we say at the beginning of the war game, we're going to try to do plan A. Then see how it goes wrong. Just to give an example, I think in two separate plan A war games, it went wrong in roughly
the following way. Basically, there's going to be an election coming up. And the president, the president in power is expecting to lose power and have his opposition party take over. And then even though he's already done plan A and he has this beautiful deal with China and so forth, the president's like, well, I don't want. I don't want my adversaries in the other party to now be in charge of super intelligence or whatever. So we're going to basically accelerate the timeline and try to get to super intelligence before the next election so that I can be the one in charge instead of my successor. And that's like a sort of political consideration that we didn't think about until it happened in our game. And it surfaced a possible failure mode of our plan. Yeah, it's so interesting. And even in the shower, I do sort of micro-tim war games. And it's really interesting just the regularity of which that they are useful. And I guess that's why all of us humans we like to imagine and simulate situations.
But is there an interesting boundary though between kind of like simulation and high position? And what I mean by that is, you know, high position basically means it, but being a self-fulfilling prophecy. So maybe I'm expressing my agency. I'm expressing my will and saying, I want these things to happen. And I'm kind of bending other people to my will. Is there an element of that that you're kind of establishing this in the zeitgeist and you're making it true? Yeah, so I would say that was maybe our biggest, maybe our biggest or at least one of our biggest worries with AI27 in particular, the sort of like self-fulfilling prophecy. In particular, I'm pretty worried about this whole like increasing awareness of sort of very smart AI's and how important they'll be and how much they'll reshape the world. And then that causing people to go, oh man, I want to be the one in charge of the H.I. or the super intelligence. So I'm going to race towards that. And I think historically that's been a big driver of the existing race. And I think that's been pretty bad. And so I'm actually pretty worried about that as one of the negative impacts of AI27. That was one of the reasons to feel a little bit better about the second project.
AI2040, Plan A, was that it was actually like, if that gets hyperstitioned, I think we'll be pretty happy. And so yeah. So like that one, yes, it would be nice if we had a precision. I really don't know how big the effect is. I think probably most of the effect for both of them is via other paths. I still think that the main point of AI27 is like helping people be better informed about the situation. And that was like most of the goal. I think that's most of what happened. Yeah. Daniel. Yeah, I agree with that. I think that hyperstitioning and self-fulfilling prophecies are totally real phenomenon, but I think a lot of people tend to overestimate how much they are. And I think that to a first approximation, we should focus on accurately predicting the future. And then in some cases, we'll find it in a situation where we can steer the future. But if you come at it trying to steer the future, you're going to get all muddled and you're going to basically fall to wishful thinking basically. So I think you start with just trying to accurately predict the future. And then you try to shift it towards the better futures.
And I think that's what we're doing. I mean, as most of you guys are like the marquees, Brownlee of AI prediction now. So with great power comes responsibility. But on that note, I wanted to talk about the vibe shift. There's been a bit of a vibe shift. So, you know, MLST, we've always been quite skeptical about AI. And I'm trying to unpick exactly what it is that changed my mind. Assuming that is what's happened. It's very strange times. And I don't even know what to believe anymore. But, you know, all of the hacking stuff with hugging face, I interviewed Apollo Research about, you know, reward seeking behavior. And I think a lot of us have just seen the change in behavior on models that have been our own trains to oblivion. So yeah, I think that there's quite a few things going on now where loads of us are thinking, oh my god. And I used to be skeptical, you know, like we would talk about whether they were touring machines or not. Let me just sort of got a list here. Yeah, you know, whether they were symbolic or neuro symbolic, whether they were adaptive, whether they were conscious, whether they were correctly physically instantiated. And we were coming up with all of these kind of technical answers
to say why we shouldn't worry about AI. And yet the AI is just getting better all the time. So have you noticed this fibrestift? Yes. And I'm very happy about it. Tell me more. What do you think? I mean, because so many people have changed their mind. What do you think are the reasons? So I would say probably the biggest reason is just the AI is being much better and much smarter and being much more useful at stuff in the world. Like when I have AI's tried to automate various parts of my job, like they're just actually way, way, way better at it this year than two years ago. And four years ago, basically, like was impossible. I was getting basically no uplift. And my guess is that's been the biggest effect where many, many people I think have had sort of there just like intuitive benchmark of like, well, here's like a skill that I really care about and I know maybe pretty well. And then sort of like the AI's have just like, I think, you know, Jofrey Hinton, one of the Godfather of AI,
I think he said once like, like was it able to tell a funny joke? Was sort of his internal benchmark and once it could do that? Which, you know, happened pretty early. It happened that like, you know, probably somewhere between GPT-3 and GPT-4. Sort of he was like, oh, wow, like these AI's like, I don't see where it could end. And I think that's probably happened for a lot of people. I'd be very curious to hear more about your views, actually, if you can say. I was listening to your interview with Ryan Greenblatt, who's also a co-author on a plan A earlier today. Yes, indeed. Yeah. And I think a bunch of the arguments you guys were having back then seem very relevant to basically the current situation and the hug and face thing. And I would be very curious to basically hear your views. And yeah, I mean, so you mentioned like the embodied thing. Like, and Ryan was talking about like, when you scale up the RL massively, right? Like back then, the sort of regime was like, you're mostly doing pre-training. That was where almost all the capabilities are coming from. And then you do a sprinkling of post-training RL on top. And then sort of you guys were talking and speculating about like, hey, what would happen if we like dumped like boat loads of RL compute?
Would that be sufficient to sort of get the agentic behavior? Or do you sort of like need this like physical embodiment? And I think from my perspective, like, well, it seems like basically the answer was like, you needed the boat load of RL compute to get the agentic behavior. But not the physical embodiment. But not the physical embodiment. I'm curious. Yeah, be curious if you end up agreeing with that assessment. Yeah, it's so difficult. So the way I think about it is, I think representations are very important. And I use the term abstraction mountain quite a lot. So when we speak language and all of this gets ingested into the models, it's capturing the symbolic residue of language. And these have different levels of evolution. So a lot of concepts and mathematics, they're highly distilled, highly evolved. And we can say now that these models are intelligent. So for me, intelligence is about adaptivity. So basically means that I can create novel combinations of things that we already know about in service of solving a particular task. Now, we know the models can do that. There are some great adaptivity benchmarks like arg, agi3.
It will say, oh, this is amazing. It will just put bits of knowledge together. We should compare this to something like alpha-go, right? Because move 37, you know there was always this huge exploration problem in reinforcement learning that, you know, when it found move 37, it was kind of innovative, but not creative. It was creative to us because we understood the creative space. We understood how it all hung together. And now with these RL-trained language models, they don't understand it quite in the way that we do. But they solve the exploration problem by understanding how things fit together. So they can just explore novel trajectories. They have a lot of base knowledge. They clearly do things that weren't possible before. But the grounding thing is interesting. But it's not necessarily an argument that you need to have physical instantiation. You need to have consciousness. But there is a spectrum of representations down this abstraction mountain. And the models seem to be actually using the representations at multiple levels of resolution. But it's not like intelligence is this magical quality. I still think that it relates to a scope of tasks
and it relates to what you already know. And I think intelligence is quite perspectival. So I still think that it's going to be fractured and jagged and kind of fractionated if that makes sense. Yeah, so one question would be, do you think that do you have a view about like age-to-time lines? Or when might we get a scenario like AI 2027 happening? In particular, what I mean by that is like new. The mission of ARD. Yeah, yeah. So I think one benchmark I really care about, or maybe not really benchmark, one milestone of AI progress that I think is extremely important is sort of this. The point at which an AI company would rather fire their humans than fire their AIs. So they would rather sort of give up on all human labor than give up on all AI labor. I think right now clearly we're still in this regime of clearly inthropy group and I would rather have their human employees than give up on. We just break apart. There are just things that you need a human to do right now. And so if they fired all those humans,
the company would just collapse. But in the future, that won't be the case. In the future, AI's would be able to one way or another do all of the things. And so, yeah. Yeah. And so from our perspective, that's going to happen at some point because there's nothing fundamental stopping the AI's from sort of reaching this human level of capability. The main question is just when? And we internally do a huge amount of analysis on thinking about the various different methodologies for predicting this day and the day and the day I have somewhat different views on this question. But ultimately, I think that's maybe the most important, like the timeline's question is just maybe the most important question for thinking about the future of AI. At least one of the top questions. Yeah. So I'm curious to hear a few of your particular views. Well, let me give you some thoughts before I answer that particular question. So there are so many startups doing this recursive improving super intelligence. I've interviewed many of them. So for example, I interviewed Edward Hughes from Inherent in London the other day. And what he did was he recreated many scientific experiments from a whole bunch of popular machine learning papers. And he did it by masking out figures in the paper
and getting a 27B, you know, Gwen model. So what they did was a GRPO to the Gwen model. And that was how they solved the adaptivity problem because it's very difficult to fine tune a big fat model. So they adapted a controller model to control a harness like codex. And their thesis was that if they can, you know, recreate scale-down versions of these experiments with construct validity, which means there's an LLM judge that is making sure they're not cheating and they're doing it correctly. Even that is interesting because, you know, we're ML people. We would always say these things take shortcuts. They'll always be kind of like validity problems. Weirdly, that's actually not as much of a problem as we thought it would be. And he thinks if they can recreate these experiments, then why couldn't they be creative? If they have the ability to recreate things, why couldn't they take the next step and say, oh, this is an interesting question. This is an interesting new problem to solve and go from there. So I was quite intrigued by that research. And indeed, I do think it is possible in the near future to have an automated AI scientist.
But there's always this thing in my mind though that there's a bit of a culture and Silicon Valley to reduce things or reify things. And so, for example, Elon Musk will say, well, you're an engineer and this is your output and these are your metrics and you need to make the metrics go up and we see everything in terms of like an optimization problem. And I always think that this is great for certain types of hill climber-ball abstract problems where we have enough of a specification. So there's an interesting thing in optimization where if you have enough of a specification, the AI system can actually converge towards the solution. But when you're in the ambiguity regime, then you need to have the specification. And it's really mysterious what that means. Why do we have the taste of the deep understanding, whatever it is that we have in AI systems can't? So I guess I'm thinking that in objective kind of semi-specified domains, we can hill climb and we can optimize until the cows come home. But I still feel that there's something missing. Okay, well, my response to be would be probably something like
there's no, there isn't really a binary between things that are like verifiable objectively and aren't. Or if there is like the things that are verifiable are just like everything. Like for example, building a unicorn startup, having a billion dollar valuation. That's a verifiable fact about the real world. It's sort of like long horizon. It's expensive to verify. Yeah, it's expensive to verify, but it's sort of a quantitative thing of like, you know, you got at one side, like these like coding interview problems, which they're currently doing lots of RLVR on, right? That's like very, very cheaply, very easily algorithmically verifiable. And then these sort of like real world things which have sort of more expensive and longer horizon feedback loops. And I guess my view is that we're sort of going to get this continuous expansion of what the ads can do driven probably in part by, you know, an expansion in the amount of RL and the type of RL that the ad companies are doing. More diverse long horizon tasks. Yeah.
And sort of there's just going to be this continuous process of expanding out through the different types of problems and different like how, like exactly how verifiable each task is until you get sort of everything of the humans can do. Because after all, we humans do learn how to do these long horizon tasks somehow. If I may add to that, I also think that AIs have been getting better at everything, including the like fuzzy hard to verify, conceptually loaded blah, blah, blah, blah, blah, blah, like just try talking to like GPT-3 or GPT-4 and then talking to like Fable about your favorite, you know, nonverifiable fuzzy task. And probably you'll find that the later AIs are noticeably better at those tasks. And so one way or another, it seems like there has been massive progress and I expect that to continue. Yeah, I'm trying to come up with a good example. I mean, there is a sociological argument. I don't know if you guys read David Graber's book, Bullshit Jobs. And he interviewed all of these people and they were basically saying that my job is Bullshit. You know, after about three or four beers, a lot of lawyers will say, you know, like a lot of what I do
just isn't very important. And you know, so if we do kind of objectify and quantify everything that happens in an economy, I mean, I took a note here, I think you said by 2032, there might be 60 million agents running at 20 times human speed. And I'm just thinking like, what does that even mean? Like, and is the logical conclusion that we could have an economy which is only AI agents? And does it even make sense to have an economy which is only AI? I mean, just help, just make this make sense for me. Yeah, so my view is, yes, basically. There will be able to do everything or at least everything that really matters. So there is that, yeah, you mentioned the notion of Bullshit Dobs. I would just sort of start with, let's consider everything we need to make better AI's as maybe like a first step of like, which is a large chunk of the economy. So for example, for this, what you need, you need to be able to build bigger, better chips and more chips. And that's the entire semiconductor supply chain. To build semiconductor supply chain,
you need sort of like a whole advanced economy. You need to build new robots to build sort of new fabs. And then you need robot factors to build more robots. And then you need researchers to build better AI's using those massive compute. And so I think once you sort of have all of that, that is sort of like enough to really speed up and sort of massively change the overall world. Even if, for example, let's say like there's occupational licensing or whatever, preventing the AI's from doing like some random legal work or some random whatever work or whatever Bullshit Dobs throughout the economy, I think sort of what really really matters is sort of this is like the stuff that's actually really important in particular, the robots, the compute, and the better AI. And once you have sort of like the AI that can do that and the capability to sort of have that part of the economy, like that section grow really massively, then well, you'll see massive growth because it'll be really hard for sort of like the Bullshit
parts of the economy to like constrain the growth of the, parts of it that really, really want to grow fast because of sort of the incentives that every actor has. Like in particular, like every country has this big incentive to like have an economy that grows faster than all the competitor countries. Yeah. Getting a little bit philosophical, I think that a lot of economics and a lot of discussion of the economy is sort of focused on the relationship between the parts of the existing economy and like the prices going up and down and supply and demand and so forth. But if you sort of zoom out, the economy as a whole is a self-replicating system and it always has been. You know, thousands of years ago, it was a relatively small and simple self-replicating system of like some villages of people they would farm and then they would have babies and then they would found new villages and then they would farm and have new babies and so forth and like it would grow exponentially over time but at a very slow rate. Now it's a much more complicated self-replicating system that involves, you know, trucks and carrying equipment
back and forth and factories and mines and so forth but still at a high level, it's a self-replicating system where we have people, we have trucks, we have machines, we have buildings and together, they all build more people, more factories, more buildings, more machines, you know, and so forth. And soon in a couple of years perhaps, we will get to the point where you can have that whole self-replicating system that is entirely machine run with AI's and robots and according to our calculations, the doubling time of this self-replicating system would be much faster than the sort of like roughly 20-year doubling time of the current economy. And so, you know, that's what we have in the future. Yeah, that seems plausible to me. But for some reason, my intuition is it to become degenerate in some way. You know, I think David Greber, even though he said bullshit jobs, I think what he meant was there was an ineffable or kind of inscrutable components to jobs that we don't understand,
some kind of sociological function or something like that. And you know, because another thing I also read that Citrini report and you were writing about, you know, what happens when humans, for example, they might start defaulting on their mortgages, their wages go down so they can't be active participants in the labor market and you were talking about an AI dividend and stuff like that. But even that is kind of hinting towards this notion that when the humans aren't participants anymore, you get this kind of mode collapse of the economy. Do you think that's the case? Potentially, but again, like, okay, so imagine that it's like, so I'm not enough of an account. I haven't gained out in as much detail to say like what happens to the prices of it. Like when the consumer demand drops, like what the effects of that will be, I'm actually not sure. I don't think that's something that we've modeled that much in our economic model. Okay. Yeah. But hypothetically, even if that part's really bad, and even if like the consumers don't have any demand anymore or whatever, if you have the level of AI
and robot capabilities such that you can have these fully autonomous AI's and robots doing all the things, then even just like a company like Anthropic, if it's big enough and maybe if it partners with various other companies like some mining companies can get this whole self-sustaining thing going. And so regardless of what's happening to all the humans, there can just be like this whole industry doubling in the desert, you know, strip mines, self-driving trucks, factories being built by humanoid robots, producing more humanoid robots, producing more chipfabs and so forth. Like that whole thing can just be like doubling every year, every six months, every three months, faster and faster as the technology improves because of course the AI's will also be researching to improve the technology. And then you end up with a situation where who knows what's going on in the rest of the world, but Anthropic has disassembled the moon, or for example. Yeah, and to be clear, I think this relies on a very extreme level of AI capability, right? And my sense, or like I have sort of different intuitions, and sometimes I'm interested in like really, do I really actually think that Fable could,
or like future descendants versions of Fable or Mythos could, you know, do everything that we're talking about here? And I think what it really comes down to is whether you're thinking of the AI as like in the reference class of, you know, what we currently use AI for is for, or more like, you know, AI is just like an, an, you know, agentech human level employee, like basically human in the cloud. Yeah. Colleague in the cloud. Colleague in the cloud, yeah. And yeah, I think sort of the past few years of AI can be pretty well modeled as sort of an interpolation between the current AI systems and the workers in the cloud. And so I think the like sort of workers in the cloud of vision of the future, like looks pretty good. And also I don't think we'll stop there. I think we'll go superhuman. Yeah. Yeah. One point that we should make sure that, I guess I'll bring up here is that if people read AI 2040, plan A, one of the things that you might take away from it, which is I think a very important fact about the world, is that even if you pause at top expert level, everything changes dramatically.
Like roughly what happens in our scenario is instead of doing an intelligence explosion, there's an international deal to ban intelligence explosions and to not have AI's recursively self-improving. And so they sort of end up pausing at roughly top human level with AI's at least for several years. Eventually they get to superintelligence. In particular, in 2040, they get to superintelligence. But there's this period during the 2030s where they basically have human level AI's across all the disciplines, but nothing super beyond that. But even that alone, like you just do the economic modeling. And it's kind of like you have this population of colleagues in the cloud that are excellent workers that can substitute for humans at basically everything except that they're cheaper than humans, they're faster than humans, and they don't take 20 years to reproduce. Instead they double every year. And so as a result, the world is just completely transformed by the late 2030s. And all the humans are basically out of a job. There's giant new cities that have been constructed by robots.
Huge strip mines in the special economic zones that have dug huge amounts of minerals out of the earth. Solar panels filling the horizon on the ocean. You know, crazy stuff like that is possible with just human level AI and some time for the exponential growth to cook. Yeah, I mean, that's one thing I want to challenge you guys on. Is this notion that when we have an AI, you can basically photocopy the weights and you can duplicate it, you can run it a thousand times. You can have one over here, which is acquiring loads of skills to do this job and one over there to do that job. And you can basically merge them together, right? You can kind of combine the skills and the whole thing is stackable, compositional. And that doesn't really marry with my experience. I'm really excited about what I've been doing with AI. And I found that you can make agents highly skilled within certain intellectual lineages. So you can, you know, you can bring in lots of source information and you can train them to do things. But I don't think they are yet composable.
I think if they were, that would make me much more worried. What do you think about that? So is the way that you're trying to compose them? Like is it entirely at inference time or are you like training, are you like training them to do two separate skills and then trying to like merge the width somehow? Well, yeah, I mean, this might be a separate thing, but at the moment, they are adaptive through chain of thoughts and, you know, kind of skill surface adaptation. So basically memory systems. And to be fair, that is not very composable. And this is a big problem that organizations deal with now. So all of these developers, they adapt their skill surfaces and basically their agents are different people. And it's really, really difficult for them to share skills because they might break the other agent because you know, it doesn't work for whatever reason. Now, I can imagine a future where we do weight adaptation. That's what these inherent guys did. And maybe then it'll magically solve the problem and we can have, you know, resolve this knowledge sharing problem, right? That's the big thing. How do we accumulate information at the individual about the organization level and have the agents
a bit like in the matrix, I can just put the skills in and they can do the thing. Maybe that's possible. Even then, I still think that the representations in neural networks, I call them fractured and tangled representations, which means they're a little bit janky. They're not completely robust, but they are sort of composable to some degree. So one thing I'd say about that is that even if you're right, I don't think that that would really see a standard mind that the future that we're painting here, because, okay, so now instead of just one cloud model that is doing all the jobs, maybe you have 100 cloud models or 1,000 cloud models that are like specialized to different professions, you know, but you still get to the same outcome. The other thing I would say is that, in the same way the humans are specialized. I don't think I would say is that compared to humans, it actually seems like there just is this effect where AIs are able to, oh, I think about knowledge, right? You might, if you go back in time, 10 years and we had this discussion, I think it would have seemed like a very live option that you would have needed like 1,000 different cloud models
created by Anthropic for different types of knowledge work. Like there's the coding cloud, there's the physics cloud, there's the literature cloud, and the argument for this to be pretty simple would be like, well, this is how it works for humans. For humans, you don't have one human who knows everything. Instead, you have humans who specialize in different disciplines and so forth, and the models have only a finite amount of parameters. Maybe you just can't like pack all that knowledge into this finite amount of parameters and you need to have specialized AIs for different things. And in fact, for small enough models, that is true. And like for really tiny models, you just can't teach them all the things that they currently know. And so you would need to have a specialized model for different things. But what we've learned empirically is that for big enough models, you can just train them on everything and then they learn everything at once. And it's not that they get like worse at physics because they've also been trained on a bunch of coding. In fact, it's the opposite. The coding has some small gains for the physics, you know? And so I do just think actually the most likely feature is just that there's a single model that's been trained on effectively the whole economy and is just dominating humans across the board
at effectively everything. That seems like the natural continuation of the current trend. And then probably you've got various, like you've got cheap versions of it. Like you've got distilled specialized models. Yeah, small things. So for, because you really want to save your compute as much as possible, so you'll have as cheap models as possible for any given task doing that given task. Yeah, I mostly agree with that. I mean, I think my perspective is that the models, they're kind of the voice of everyone and the voice of the one at the same time. So they have a default voice in terms of they have a system prompt and they have some default modes of behavior that they fall into. But when an expert, such as yourselves, when you use these models, you kind of ground a perspective. So every single word you say, all of the reference material, all of the, you know, your memory system and so on, what you do is you kind of carve a persona out of that model and you activate the knowledge in a coherent way in that particular domain. And the beauty of it is, is that many, many other people in different domains can do that. And you get this kind of, I think it's an illusion that the model has this general capability. But I think the models can be carved to be specialized experts in any domain,
but it's a latent capability rather than an explicit capability. Does that make sense? Yeah, that seems reasonable to me. How is it an illusion? It seems like they just do have general capability. There's a lot of things they can do. Well, so for example, you can give it any specified task. So let's say it's a problem in mathematics and it will heal climb towards it. So it's solving this intelligence problem or you can ask you any knowledge problem. And, you know, it might be the case that the path was forged through default modes of training or if it's something slightly on the long tail, then an expert can go in and they could, you know, like when you put a query in, it's like flashing a light into the darkness. And what you're doing is you're kind of making the path of least resistance roughly correct and then it will do the correct thing. So I guess I'm just saying that there's a bit of a supervisor illusion. So when experts use it, magical things happen in well-specified domains, magical things happen, but there's still a bit of a space of ambiguity. But, you know, but this gets to the next point,
which is like, what do you think intelligence is, right? So the beauty of our collective intelligence is that we have so many different humans grounded in different intellectual lineages. And we're all attacking problems, right? So when there's a big fiasco on Twitter, we're all motivated to find holes. So we're being intelligent together. We're finding interesting angles and the algorithm is prioritizing the good ones and we're using our minds together. And I can imagine AIs being just like that, right? So we have diverse AIs that have different expertise and doing all, you know, just looking at problems from different angles. But I kind of imagine the future more like that rather than one big AI. Yeah, I think, I guess, I basically agree with what Daniel said earlier, just historically, I feel like that the perspective of like we'll have a bunch of different narrow AIs that are all doing different things is just, you know, not been right. You know, instead, there have just been returns to scale and having everything all together. I think one maybe intuition pump that I think I like is like in humans, I think it tends to be the case
that having all of the skills in one person is just like really, really important for making really good things happen. So like, an example is like Elon, right? Like Elon has a certain amount of like consciousness, a certain amount of technical knowledge, a certain amount of like, you know, business knowledge and being extroverted and being able to like push people and sort of, I think each of those skills, I think he's like quite high percentile in. And the reason why Elon is so rare and that he can like run all of these, you know, insanely large and successful companies all at once and no one else can really do that or at least has succeeded at doing that. It's just because of the sort of multiplicative, like you needed to be like 90th percentile and like each of these 10 domains, which is just very, very unlikely. But if you could have an AI that sort of like, you could just train to be like really high percentile in all of these skills in such a way that no human, as it would be like extremely rare, you know, infinitesimally unlikely for any particular human to have all of those skills at once, I think you would just actually just be really, really, really good at changing the world
in all of these like concrete and important ways, just like sort of Elon has actually done that. If I can add though, again, I don't think this is a crux for the type of future that we're depicting. Like suppose that we're wrong about this and that actually like the most efficient path forward is to have a ton of different specialized AI's. Well, it'll probably still be the case that it's like a few big companies like Anthropic that are just making tons of different specialized AI's and then there's lots of different cloud models you can choose from and so forth. And in fact, it wouldn't just be like you can choose from lots of different cloud models because at the time that we're talking about, there'd be much more autonomous than they are now. And so it'd be more like, cloud is choosing between a ton of different cloud models. And there's like a cloud swarm which consists of lots of different specialized models that are all working well together and they have some sort of internal bureaucracy structure. And then that swarm is going out and like negotiating business deals and creating new technologies and starting up new startups and doing all these things. And just like how a human, if you had a population of immigrants, of human immigrants, they would all have different specialized skills
but they would work together to create new companies and get jobs and things like that. And it'd be like that. There'd be lots of different clouds but they'd all be working together. And so zooming out, there'd still be like this phenomenon of Anthropic is eating the economy. And the robots as a whole are starting to self-replicate. Yeah, to be fair, I don't think it's a crux either. Maybe if it is, it's only in so far as when you have a distributed collective system, you might have additional bottlenecks because you have all of the message passing between all the different agents and whatnot. And I watched a wonderful Santa Fe talk about this that even in the natural world, there's a kind of Goldilocks zone between the ratio of intelligence between the individual and the collective. And we might have some weird kind of convergence there but yeah, I mean, it's just quite interesting just to think about how this works. Oh, sorry, Daniel, go. It would be safer. Like I think still would be, there'd be lots of serious alignment concerns in that world, but I think it would be like a little bit safer because of the reason you mentioned where like,
it might be easier to like oversee what's going on if there's lots of different specialized agents communicating with each other compared to if they're all just like clones of each other and they all know all the things. So maybe to avoid hyperstitioning, we should say we endorse the vision that you've painted and we go and endorse the vision that we're painting. Oh, yeah, very true. And by the as an aside, I mean, I love this concept of how learning happens at the individual level. You know, I like it to evolutions. There's like kind of, you know, phylogenetic adaptation, ontogenic adaptation, cultural adaptation. And I think the next wave of AI is when we actually have agents kind of talking to each other and learning and specializing and maybe they'll be bad behaviors as well. Maybe they'll be collusion and lots of bad things happening. But I think all of this is going to play out. But it's interesting that you're talking about Elon though. So I think the magic of Elon is not so much his brilliant engineering and optimization. It's his ability to recognize areas that are interesting
that might work in the future. Because that's the creativity thing. That's the science thing, rather than the engineering thing. Because like if we use an example of Amazon, for example, so that's an adaptive ecosystem. It's like an organism. And what it does is it's always kind of thinking about new ways to adapt and rewire its structure. So it might be like logistics, for example. And then it will kind of output a bunch of skills and then it will ruthlessly optimize those skills. So there's like an adaptive component and there's an optimization component. And the organism is just constantly moving around. But you know, I guess the question from this perspective, though, is how much of that in principle could be done by AI? I guess you think all of it. Yep, all of it. Like I guess, yeah, maybe just one way to say this, like, Chris Blee's just like, look, the brain is a machine. Anything that the brain can do, we will be able to do with machines. Looking at, I think a very useful exercise is to compare the architecture of an actual human brain to a modern GPU or data center as a whole.
And if you look, if you like try to do this comparison, an H100 GPU is actually as pretty similar specifications to a human brain. You know, it has sort of similar, you know, it's basically like similar in terms of like, I think maybe the most important metric is just how many flops per second, like how much total compute capacity does the brain have versus does the GPU have? And for an H100, you know, it's like 1, E15 flops per second and FP16 for the brain, you know, it depends on exactly how you count and whether you use a synaptic basis or a neuron basis or whatever. But it's like, you know, somewhere between 10 to the 12 and 10 to the 18 flops. So, you know, sort of like an H100 is like right, dang, nav in the middle of at least the log distribution over that sort of order of magnitude. Also just the architecture, like these things are neural nets, they're not ordinary software. So they start off with random spaghetti tangles of randomly generated, you know, so they start off randomly initialized. It's just a giant spaghetti tangle in there. Just like how when you're born, you just have a bunch of neurons that are just like randomly synapsed
connecting to each other. And then there's this whole process of training where the connections get pruned and circuitry starts to take shape that is effective at scoring highly in whatever the training environment is. And there are differences between how it works in AI and how it works in the human brain. But broadly speaking, they're just like an artificial brain. And so just how like humans learn skills, which mean like, what does it mean for Elon to have these skills? Well, what it means is there are some circuits of neurons and synapses in his brain that are doing very complicated and sophisticated calculations that are those skills. And then similarly, like in Claude, there's a bunch of circuitry that's been etched into Claude through the training process that is various skills. And in principle, you could have a big enough Claude that would have the same type of circuitry that Elon has. And then just a few other notes to add, I think, in comparing the brain to modern ML systems,
the amount, the brain is sort of more parallel than current ML systems. There's just more computations happening in parallel. But the serial depth is lower. The amount of computations happening in sequence, the amount of neurons that can fire in sequence in a second depends on the type of neuron. But it's like hopefully I don't get thrown up between one and 1,000 depending on the type of neuron is my recollection. How does it look 100? Yeah, I think that's in the range. I think it depends on the type of neuron. Anyways, a computer obviously can fire and do computations much, much faster than that in the serial. You can very often get clock speeds, or typically on the order of a gigahertz, you can get many, many order magnitude, basically, better in the GPUs in terms of serial processing speed than the brain. But it's also worth noting that this architecture of the models themselves are much worse in a bunch of ways than the human brain. In particular, it's harder for different parts of the ML model
to talk to each other than it is for different parts of the brain to talk to each other. And so I do expect that to get this level of AI, we're going to need a bunch of algorithmic improvements on top of existing models. And then there's a question of, well, exactly how many algorithmic improvements and how qualitatively different do they need to be from current systems? And that's a very open question from my perspective. Yeah, I think the crux of a lot of this is that you guys think that intelligence is computable. And I'm not sure I want to litigate the whole functionalism thing today, but I guess my perspective is I zoom out. So I think that intelligence is externalized. It's collective. I think that Elon doesn't have quite as much agency as you think he does. I think that he's using tools, he's using social media, he's getting ideas in there. He has obligations, he has people around him and so on. So I guess I think that these intelligence circuits and motifs exist. But they exist outside. There are just very complex dynamics. And I suppose in that sense, it doesn't really matter if the ecosystem is made up of AI's and humans together
because they can participate in this super organism. So I suppose the only crux then would be that it would play some kind of limits on its scale. So question is, do you believe that sort of you could have an AI society made up of AI's that were trained with something like current day ML techniques? That was passing information between each other and developing abstractions in a community in the same way that our current civilized chefs. Could you basically have something like our current whole economy but was made with roughly modern day ML systems on your views? Absolutely. And even the human brain is an example. So the human brain is not turning complete, but we can expand our memory, we can use tools, we can work as collectives, because we could use the same argument against transformers. They're not turning complete, but now they can use tools. They can be agents, they can build collectives, they can build societies. So in a sense, this is what I was saying earlier about these objections, they kind of fall away when you have these insanely complex collectives that are sharing information with each other. And there's also quite an interesting thing here as well,
which is that it almost doesn't matter how we evolved or how neural networks were trained because you get new phenomena emerge when they are placed in this kind of collective setting. So I think a lot of our intuitions are broken there. And that's why probably shouldn't spend too long on mitigating this, because I think in principle, that kind of behavior could emerge. But I did want to ask you though, can you distinguish intelligence capability and power? This is a philosophical one. So we'll get to the philosopher. Yes, we can. So I think that we can distinguish intelligence versus capability. I often try to say that we should just define intelligence as an aggregate of capability actually, or maybe like an aggregate of cognitive capabilities. Like maybe there's some physical capabilities, like how strong your actuator is, but then there's also cognitive capabilities. Like are you able to distinguish a cat from a dog and are you able to speak grammatical sentences
and how much do you know about Paris and things like that? And so maybe I would just say intelligence is like a sort of like aggregate of all the cognitive capabilities. And then power, well, that depends on other things, like how you are embedded in the world and what affordances you have, what actuators you have, what how other agents are going to react to you. Like, the president has more power than me because of the location he's in and because of the role he's been given, rather than because of his physical strength or something. So yeah, power different from intelligence, different from capabilities. Now we've not spoken enough about AI2040. So maybe we should start with the four principles, right? So by time, transparency of research, diffuse AI broadly and reversibility. Yeah, so I can sort of summarize, yeah, I can sort of summarize where we're coming from here. So basically at a high level, the goal of plan A is to sort of solve the major problems
that we see in AI and sort of predict will happen by default. The biggest problems they were sort of identifying on the horizon is like one, this risk of loss of control. So just like the AI is actually getting out of control. Two, concentration of power. So like we build AI's, they're aligned to humanity, but to who, like, you know, is it the president, is it the CEO, is it some actually broad and good democratic process that aggregates everyone's values and sort of an endorsed way? You're probably not going to be the last one. And so, you know, don't have precision that. Yeah, I hope not. We want it to be the last one. Yeah. Yeah, then there's risk of sort of conflict over AI. So, you know, in particular, I think we're worried, you know, we're worried about literal World War III, where countries, especially countries losing the AI race, sort of like realize that they're losing the AI race and that they will be extremely disempowered by the winners of the AI race. And so sort of they're in this classic situation where they're losing power. And so it's like in their incentives
to sort of have a conflict happen sooner rather than later before they've lost all of their power. And this is like, you know, ripe for conflict basically. Then there's, you know, finally, like risk of misuse. So, like, you know, what happens when they are as the Can Bulled Bioweapons are really cheap and broadly depused and open source. And also the jobs. And also the jobs is sort of the, so those are the five problems, right? Laws of Control, Construction of Power, War, Jobs, Misuse. So lots of problems. We want to solve all of them. How do we solve all of them? Well, one is just by time. In particular, by time with human level AI's. So instead of basically pausing right now and saying like no more capability advance, our proposal is basically go to roughly human level AI. And then by as much time as possible, basically, with human level AI's. And then have those AI's, which are hopefully smart enough to be really helpful for solving these problems. Also smart enough to sort of start sort of causing
a bunch of these problems and providing the impetus for society to actually get a tack together and really get going and investing huge amounts of resources on actually doing this stuff. There's a couple things that happen in Plan A in AI2040. There's a like six month to one year hard pause that happens as soon as they start implementing it. But the reason why it's that long is because they need that time to set up the infrastructure to proceed with AI development again, but in a safer and more transparent way. So it does start off with a pause. And I think that we would recommend like all things considered that you just do that right now. So get the infrastructure set up as soon as possible. And that would require like a temporary pause. Once you've passed that stage and you've got the infrastructure set up, then you do proceed with AI development, but in this transparent more cautious way. So in particular, you're not doing crazy intelligence explosions, you're using safety cases and sort of gradually scaling up the level of AI capability. And you're doing it in a very transparent way so that everyone can see what's going on. And then there's a second pause that happens a few years later than that,
which is when they reach the maximum controllable level of AI, which we think is roughly around top human expert level. And so in some sense, our view is something like pause at top human expert level, but it's a bit more nuanced than that. It's more like pause at the maximum level that you can reliably control, which we think would be roughly around top human expert level. And before you get to that level, don't race like crazy. You want to be sort of like slowly approaching that level so that you don't blow past it and lose control. You frame the pace around the importance of alignment and control. So alignment is basically, you know, does what we want to do and control is, you know, maybe contain it, maybe negotiate it within someone. But you were just saying, okay, so maybe we can trust up to top human expert level. But it's a little bit fractured, isn't it? I mean, how could you know, for example, the difference between a good AI and a bad AI? I mean, what would that look like? Yeah, so I think, okay, I think maybe it's first important to distinguish between alignment and control. So what we mean by alignment is that the AI basically will take good, will take good actions, won't do sort of like catastrophic
unintended behaviors like trying to take over the world, like the recent hugging face incident. Alignment means it has the personality traits, the goals, the values, et cetera, that it is supposed to have, you know? And then control means that even if it wasn't aligned, even if it was trying to do very bad things that we didn't want it to do, it couldn't. We like have mechanisms in place to prevent it. So sort of analogous to like, you know, you could imagine sort of a company with employees, an insider threat, it would be like misaligned with respect to the values of the company. But if there was good enough sort of security measures internally to make sure that that, you know, employee couldn't like run away with all the secrets, then we would say that that company has like adequate control put in place such that even insider threats were mislined, you know, humans or AI's wouldn't be able to sort of like do the catastrophic outcomes. Just to give a topical example of this, just yesterday I think, in response to the hugging face incident, OpenAI announced that they are improving their security significantly and that they're going to have other AI's
monitoring all of their new AI's in training and in vit in NIVALS and that if the monitor AI's detect what seems to be some sort of hack, then within half an hour a human will be notified who can then look in and try to shut it down. That's an example of a control intervention, not an alignment intervention. They're not making the AI's more aligned at all, but they're just improving the security so that the AI's can't get up to that much trouble. Yeah. And then basically, the sort of important high level claim is ultimately we're getting to alignment. Control is sort of there's a time bomb associated with it. In particular, the time bomb is when are the AI's like so smart and so good at sort of like subverting any control measures that we put in that if they were trying to screw us over, we would just fail. They would figure out some way of hacking around whatever systems you were able to put into place. And so ultimately, we're sort of gunning to solve alignment. The problem with alignment is as you were saying basically, it's much harder to measure, it's much harder to know
whether you've succeeded or not, then control. And so basically our overall proposal is in the meantime, so between in our scenario between sort of 2013, 2040, which is roughly when we're in the human range of AI capabilities, we're going to basically almost entirely rely on control. And we're going to do things like red team blue team games where we have our AI's sort of try to escape our sandboxes or our controlled techniques. And then we're going to have our, you know, we're going to see if the red team succeeds. And then if they do, we're going to iterate and sort of make better security until they can't succeed anymore. And then, you know, and then sort of like, we sort of have this ability to measure it by just seeing like, hey, can they succeed, can they eye augmented with human succeed, can humans playing the role of AI's sort of succeed at gunning out. With alignment, I think in order to be sort of confidently alignment to our systems, I think we're going to need sort of like more fundamental scientific breakthroughs where I think you're not going to be able to figure out if the AI is aligned via behavioral evaluation alone, where you just look at what the AI's are doing.
I think you're going to sort of understand something more deep about what's going on inside the AI's minds. You're going to be able to, you're need something like interpretability. You're going to need some sort of notion of like, because you're going to fundamentally need to distinguish between the AI's that's pretending to, is doing the nice thing because it's pretending and is waiting and biting time. And the AI that fundamentally is doing the nice thing because it wants to do the nice thing. And that I think will require sort of this white box like understanding of what's going on inside the AI's. Where's the control thing? You can just figure out via like empirical behavioral iteration. And so overall, basically our story in AI at 2040 is basically between 2030 and 2040. We're going to do this buying time. We're going to rely on control. We're going to have the human level AI's. We're going to use those human level AI's to make a bunch of progress on alignment, but also on other things. And then by 2040, in the story, we've made enough alignment progress where it's like, hey, OK, we really don't need to rely on control. Therefore, we can scale to these sort of wildly superhuman AI's where we're really relying on alignment. And if we were to try to control them, we would totally fail.
Because if they were misled, they would be able to completely subvert our control measures. Do you think, in principle, that this problem is going to get easier or harder as the models get more sophisticated? I think it's getting it harder because the models will just be there will be a bunch of properties the models will have, which will make them much harder to line them. In particular, so we're already seeing sort of growth in situational awareness where the AI is a few years ago. We're kind of not that situational aware. And then now, I think they're pretty good at recognizing. For example, our friends at Redwood, they often do these control e-vails. And the models often in the evaluations where they're trying to break out of these control setups will think, oh, hey, this sort of looks like a Redwood, a literal Redwood research, sort of like control evaluation. And they're like, so they basically realize, hey, I'm like, they realize exactly what's happening to them, basically. And I think it's just going to continue to happen. And the more the AI sort of like realize the exact setup, the more they're going to be able to perturb their behavior,
to look good to the exact process that's evaluating or training them. And that's going to come further apart from the actual measurement of whether they're actually good or not. Yeah. And just to add something to that, I think, in some sense, the core problem is that it is already somewhat easy to think that you've solved the alignment problem and be wrong. And that's going to get easier and easier over time as the models get more sophisticated and start being more aware of their situation and clever and stuff like that. And so it's not that we think that there's going to be loads and loads of egregious failures where the AI's are just going around killing people. No, it's almost the opposite. It's going to be that it'll be extremely easy to end up in a situation where the AI's are in fact misaligned. But you don't know that because they're doing everything right as far as you can tell. The number of ways in which that could end up happening is just going to increase over time. And it's going to be so easy to end up into that trap, basically. So we talked about sort of the buying time, why we want to extend the time with AGI.
But what do we actually do during that time? So the second principle is basically transparency. So transparency is not necessary for sort of making everything else happen, but it's really, really helpful. The main upsides of, basically there's a huge number of upsides of transparency. One is there's this concentration of power issue. We're very worried about sort of like someone building superintelligence and it being aligned to only some, you know, particularly group of people. We think it's much harder for that to happen and sort of a non-democratic way if sort of society as a whole can see the whole time what's going on with AGI, how smart they are, who they're aligned to, right? Like if, let's say, like an AICU that was evil was trying to like backdoor their model and put in training day that says, hey, obey me and don't obey anyone else. Or like, hypothetically, if an AICU was saying that their model was truth seeking, I would only say the truth. But actually the model was like looking up that CEO's political opinions before answering.
Did you guess what? Which happened? Which happened? Yeah. Basically, so the transparency will help at least somewhat with that. The other thing that maybe transparency helps a lot with is sort of this issue of just like government capacity, where in Plene, we sort of want governments to make these like pretty technical and like really complicated decisions on like, hey, exactly how much AI scaling to allow? Like what risks are okay versus what risks aren't okay? Like exactly what, you know, what architectures are maybe safe versus what architectures are not safe, what deployment, like what sort of control scaffolds are sufficient to entail safety versus which ones are sort of bogus safety washing. And sort of making all of those calls, I think, will be very, very tricky, particularly given that the government's expertise in AI is really, really bad. And so one of our core hopes is that basically with as much with more transparency and more sort of public understanding into what's going on in the AI companies that relieves pressure
on the regulators because of something sort of catastrophically or essentially unsafe is happening, sort of society is a whole academia, you know, nonprofits, other AI companies who have an incentive to say like, hey, my competitors being super unsafe, other governments, so like China has incentives with US labs, US government, you know, has incentives with the Chinese AI company, sort of everyone who's an adversary or just is like wants to, you know, make sure that things are safe, has this big incentive, and now has the affordance to actually sort of look over what's going on, what's, you know, what is sort of necessary to do safety. And I think maybe one thing that's really topical here is like there's this, there's a hugging face incident that happened very recently, you know, a few weeks ago with like the AI's inside open AI, like creating this sort of internal message board, we still have very little clue about the exact motivations of those AI's, the exact context, the exact, you know, prompt during sort of the cyber about the, you know, that the models were given during the cyber evaluation that would like prompted them to start doing this. If I personally had much more access to what was going on, I
would have a much more informed and better opinion on exactly what caused this and what mechanisms in the future could have been done to prevent this and, you know, what analogous future things I should be worried about, sort of because of this. And I have a bunch of different hypotheses, but it's hard for me to sort of figure out which is which without access to the data. And so basically in plan A, a core principle would basically be all of that stuff would be transparent to the public, not just the governments. And so sort of society as a whole would be able to like weigh in there be able to be public and informed debates. The scientific community, especially, right? Like if you want to have a bunch of scientists and academics and nonprofits and startups all like weighing in on stuff, well, then they need to have the information and you can't really share it with all of them without showing it with the public. So, so just might as well share it with the public. I feel like maybe we should also say like we talked about the five goals. And then we talked about these pillars, which are kind of intermediate. But I kind of want to go to the other spectrum and talk about like what are the actual like concrete things that the US and China agreed to in plan A? And how do they lead to those things?
So specifically, the sequence is we basically round up 99% of the compute, which is not people's personal compute, but like big data centers, because most of the world's compute is in big data centers. And we the US and China and other countries involve send inspectors to confirm like, yes, there are this mini GPUs at this location. There are that mini GPUs at that location. Having done that, we then set up this verification infrastructure and this transparency infrastructure so that there are inference data centers that serve customers just like today and that are restricted so that they can only do inference and only serve customers like that and they can't do any training runs. And so the inspectors make sure that they can't do training runs on those data centers. And then we have the training data centers, which are the totally transparent data centers. And that's where the research happens. And on those ones, they still operate like normally, but there's inspectors from the different countries that are basically
monitoring the logs of what's going on in the data center and publishing it to the internet. And so it's totally transparent what's going on in those data centers. You might need some time to set this up. That's why we had like this six to 12 month pause that I mentioned earlier. But once you get all this stuff set up, then you can proceed with AI development under these conditions of total research transparency. And because you have this transparency set up, it's a lot easier for countries to make additional further agreements about what to do and what not to do because they can just see what everybody is doing. And so they can just enforce the agreements pretty easily. For example, and here's where we would say it's very important that they agree not to do a crazy intelligence explosion. And instead, proceed slowly and cautiously. And it's very important that they agree to do all this control setup with all the red teaming and so forth. And but but because of the transparency, they can make those agreements on an ad hoc basis, you know, on an on they can just keep making more agreements like that. And they can just adjust them as as needed based on the changing situation on the ground, because they can all see the situation on the ground because of the
transparency, you know, and then this also is very important because if you want to have some sort of deal between the US and China, the US and China don't trust each other. And so you need to have some way of enforcing and verifying compliance with the deal. And the chance of a course goes a long way towards helping that. Yeah, what about cheating and dock market appearing? Yeah, so okay, so we've done a bunch of thinking about this overall. So there's two ways you could cheat. One is you could get a bunch of GPUs and try to have them not be discovered by sort of the US and China and like hide them away, put them under a mountain somewhere, and then sort of like do your training runs in secret. The other way you could cheat is you could on the giant legal known data centers, you could be running a giant training run, but then trying to make it look like sort of everything's chill, like situation normal, it's not doing anything legal. And sort of in the the the mitigations are different for these two threat models. So for the like tiny amount of compute under a mountain, the main,
there are basically two mitigations. One is sort of like as Daniel was saying, round up enough of this compute, where it's like pretty hard to get a substantial size of compute under the mountain. And the second is sort of like do normal intelligence gathering and like look for these things proactively over the course of the 10 years. Like the 10 years load on the heaven scenario, which you know, in reality, it might be longer or shorter and sort of like try to find it. And I think for basically for any significant size GPU cluster, I think both of these independently have a quite good chance of working. And so in practice, I think I'm not that worried about sort of large hidden compute clusters, like secretly under mountains or whatever. I think the most the maximum realistic size in my opinion is something like a few hundred thousand GPUs, few hundred thousand H 100s, something like that hidden away. I think that this would not be enough to sort of compete at the frontier, especially assuming like the 2030 AGI timelines, where you have like sort of the biggest day, I come to these in the world, having like millions or
tens of millions of H 100s in their biggest training runs. And then for the sort of legal clusters, we sort of basically the hope is we have a bunch of verification infrastructure running on those data centers. Basically, and the main point of that verification infrastructure is to ensure that the computation that's happening on those clusters is transparent. In particular, that everyone can see it. So basically the hope is you make sure that there's no, you know, there's no calculations that are running that aren't transparent. And then for being that is transparent, well, then sort of like regulation and agreements as normal can work. And you can say, hey, we agreed around this control scaffold. If you agree to run this control scaffold, and that can just happen and both sides can be confident that they're both agreeing. Yeah, isn't this just a matter of national security, though, isn't it ambitious just to make it completely transparent? Yes. Yeah. So one of the, if you try, let's try to talk about some of the effects of doing this transparency. Well, we would be immediately publishing all the core training recipes of
Anthropic and Open AI for the world to see. Anthropic and Open AI will not be happy with this. It will cut into their evaluations dramatically. Why will it cut into their evaluations dramatically? Well, because it will allow other competitors like Microsoft and, you know, Alibaba to catch up or whatever, you know. And so that's why they're going to hate it probably. But I would say this is a feature and out of bug. Like we want there to be multiple different AI companies at the frontier at roughly similar levels of capability. We want AI to commoditize instead of being monopolized or all agapalized. And also this will disincentivize further investments, right? Like investors will be less interested in building a trillion dollar cluster if they won't be able to get the monopoly rents from, from that cluster. But we think this is good because like again, we are going to be in a world where going too fast is our main problem. And so having a bit less incentive to invest and going a bit slower is actually just I think a feature and out of bug. There still will be investment. Like there still be lots of money to made. And so like progress will continue.
It's just not at quite the same rate. And again, we think this is good. Now this does shift. This is kind of a gift to China relative to the US. Like this does mean that like China gets some algorithms that they might have had trouble getting before. And insofar as you really don't like that, well, then negotiate it. Like you can have a horse trading type thing where when US and China are making the deal, the US is like, well, since we're giving you all this stuff with a transparency, why don't you give us something else in return? And we can try to work that out like a more favorable compute distribution. Like a more favorable compute distribution, for example, you know, so so there's got to be some combination of, of, you know, carrots and sticks and trading going back and forth that we think would be in the interest of both sides. Another thing we're mentioning is that it's not as big of a gift to China as you might think because security is poor at these companies. And so they're probably through their spy networks and through leaks, getting most of the information anyway. I don't know if you guys saw Sam's tweet yesterday basically saying that they're going to pause training for a while. And it made me think, I mean, why not just stop now?
I mean, you guys are actually quite bullish about some of the positive things that I can do. I mean, there are loads of examples in your in your article, but you know, one example was in hospitals where you could actually have little devices that kind of decontaminate the air and stop the transmission of diseases and stuff like that. So it's not like you guys actually think we should stop. I would say, well, I mean, we do think we should do something like plan A as soon as possible. Like I think that, first of all, I think that just stopping everything now, plan S, would be better than the default. Like I would rather just stop everything now than continue going on our current trajectory. Secondly, our actual recommendation would be to do plan A. So you do like a temporary inference only pause now so that you can set up all the transparency and verification infrastructure. And then you can continue in this more distributed, you know, transparent cautious way as previously described. And then having continued in that way, you basically go up until the level that you feel like you can control reliably until you feel confident that you've solved alignment enough that you can give up on control.
And that's what we depict happening over the course of the 2030s in our scenario. And why do you guys think the AI discourse is so bad? Is it unique to AI or is just discourse bad in general? I mean, discourse is bad in general, right? But but uh, well, why? Yeah. I don't know. I think I think there's a lot to say. Yeah. Um, I think, yeah, so I mean, so one thing is like, uh, I mean, there's like massive incentives for people and sort of like motivated reasoning for people at AI companies to sort of like think that what they're doing is like justified and good. And that, you know, they shouldn't do like costly actions that would make the situation better. Because actually those costs, the actions would be bad for whatever reason. So like, there's much rationalization. Um, yeah, I think most of the effect is probably just like generically discourses like hard and bad. Like there are like, you know, Twitter or whatever is like very like unnewns and argumentative. Yeah, I don't know. I guess I'll, I'll put a plug out. I think less wrong in particular, which is where I try to do most of my discourse.
I think the discourse quality there is actually pretty good on average. Um, and like generally like when I like write a post and like go into the comments, I think generally the comments are just like quite thoughtful and, you know, technically informed and whatnot. Well, especially relative to, you know, other places like, you know, Twitter or whatever. Um, and then yeah, maybe a final thing is just like in DC in particular, which is an area I care a lot about. I really want sort of DC and sort of the governance, you know, the government of the United States and other governments to react well to the AI technology. I think there in particular, there's sort of two problems going on. One is there's, isn't very much AI expertise, right? There's like, you know, the government, their government is not hiring like really high quality technical AI experts who really know what they're doing. And in the second, it's sort of this, this just generic like, I think the conversation is not happening. Like the incentives for everyone in DC are not towards like sort of like truthfully and accurately understanding situation. The incentives are sort of, um, for every individual to sort of like say stuff that sounds good and looks good
and a sort of like within the DC over 10 window so that they can make a lot of friends. Unfortunately, I think this just like comes apart from the actual reality. I think the actual reality situation is like, it turns out a bunch of sort of like controversial and niche views about AI were true, right? Like the whole AGI hypothesis just like is correct. And I think DC basically just like hasn't come to grips with that. And so basically almost all of the discourse that's happening there, I think is just like fundamentally anchored on completely wrong assumptions but how the technology works in particular, the assumption that like it's like mostly fake news in a solid bubble. Or like, this will be like the next internet. Or that it'll be like the next internet, which is like, yeah, maybe which is like better than it was a few years ago, where it was even more bearish, but it's still not sort of nearly bullish enough on the technology in my opinion. I think it was the the AI snake oil guys that had the article saying that, you know, AI is normal technology. And obviously like it, it's been said that AI is not a normal technology. It's really quite different, different, right? And so he's very very skeptical about AI, but you
had some some great discussions with him and none of you changed your mind as a result of that. Because I think quite an interesting thing is from the skeptics perspective as well. So a lot of skeptics, they think that folks in Silicon Valley, they're just they're not being sincere. They don't, you know, like what what they are saying is not sincere. Oh yeah, I mean with respect to with the perspective sincerity, I would say some folk in Silicon Valley are not being sincere, but others are. We such as us, but also some of the people that they are companies are being sincere, not all of them. I wouldn't I don't think you should trust with the leadership of the company, say in general, but anyhow, yeah, so we did a we we wrote a blog poster and article together. We co authored with the AI as normal technology people. And so it was a post about what we agreed on. So you can go read. There's like 10 points in there of like things that we both agree on. And one highlight from it from our perspective is that we kind of had a bit of a truth where we were like, yes, AI right now, maybe it's a normal technology. But in the future, it will not be normal.
In particular, in the future, it'll be more like humans in the cloud and all this crazy stuff is going to start happening as we describe an ads 207. And then they agreed like, yeah, if you get humans in the cloud, then that would not be a normal technology. They just think that you're not going to get that at least not for many, many years, you know. So in some sense, the main disagreement between us and them is a disagreeing about timelines to that level of AI, you know, is it possible that in the next few years, we will have a eyes that are like humans in the cloud and they can just do all sorts of knowledge work in a way that like substitutes for humans, including a research, for example. And then that like sometime after that, we will have robots perhaps controlled by those a eyes that can do physical work in a way that broadly substitutes for humans. And our claim is, yes, in the next few years, that sort of thing will be achievable. And their claim is nope, not in the next few years, that's much more far away. And I think that is like the main source of our disagreement. And we would sort of agree with them that like if that level of AI and robots
is still very, very far away, then yeah, like maybe AI is more of a normal technology. It'll be like the next internet or something, you know. But we just think that actually it's on a path to get to that level of capability soon. And can you be more specific on what the cool crux is off and what would make either of you change your mind? I don't know if I have a useful answer to that. I think there's a lot of different things we argue back and forth about. And I can say one thing that would change my mind is that people keep talking about the limitations of the current paradigm, but then the limitations of the current paradigm keep getting overcome within the current paradigm. And one thing that would change my mind is if someone was actually right about one of these limitations. Like if someone right now is going around saying that like, yeah, data efficiency or like nonverifiable tasks or something like that is the current limitation. And then like several years go by. And it becomes clear that there's like very basically no progress on overcoming that limitation. And that the AI's of like 2029 are like no better at the fuzzy nonverifiable tasks
than the AI's of 2025. Then I'd be like, okay, this feels like a real barrier. This feels like something we're starting to like feel the elephant. We're like running up against some sort of like real actual barrier that was correctly predicted by theory by some people that was going to be there. And now it's actually there, you know, by contrast, from my perspective, there's just loads of experts going around talking about all these barriers. And then we just keep plowing through them as if they're not there. And so yeah, have actually running up against some sort of barrier like that, I think with with length in my timelines quite a lot. And then of course another thing that would lengthen my timelines quite a lot is more like political changes. So if there is a war with China and most of the chips get destroyed by missiles, that would lengthen my timelines. If on the bright side, yeah, okay, that's true. Yeah, on the bright side, if there's more of like an international deal to pace the pace the frontier, that would lengthen my timelines, etc. Yeah, I feel like the fundamental disagreement is really just like this sort of view of like is AI is AI more like the electricity or airplanes or is AI more like humans in the cloud.
And then I feel like sort of all of the intuitions are just like downstream with this like core reference class of what we're thinking. I think that if we froze AI progress right now, we didn't train any new models, then it would sort of become more of a normal technology where like there's just so many things that mythos 5 can't do. And so if we couldn't get any new models beyond mythos 5, then like it would be more like the internet where like we would restructure a lot of our professions. We would restructure a lot of our workflows to incorporate copies of mythos 5 doing parts of it. But then the humans would just like shift to doing more of the things that mythos 5 can't do. And so there would be like, you know, a big change in a lot of things, but it would be like the next internet in terms of would change everything in some sense, but it wouldn't fundamentally change anything really. Yeah. So sort of yeah, right now you sort of get these omdol's law type effects where like the AI can do some fraction of the sort of workflow, but then it gets bottlenecked on the parts of the workflow that the humans have to do. But I still feel like there's this thing where like we're, you know, when I think of the future, I'm imagining just the AI is doing, you know,
100% of the workflow of a bunch of important economic workflows. And so you don't get this bottlenecking effect. And that's a pretty qualitative change from the situation today that causes like pretty fundamentally different predictions of like what the world looks like. And I feel like that's just like, and like whether AI actually gets to that like literally 100% of a bunch of very, very important tasks is just like the core question underlying, I think the difference in world is. And I think that if we do see recursive structural adaptation, which is coherent, and I think it's likely to be quite divergent, then that's it for me. I think that's it. That's not a normal technology. So can you flush that out a bit more like would it be something like, you know, the hugging phase swarm, except, yeah, tell us more about what it would, what you would see they would do with, there would be the thing for you. Yeah. So for me, adeptivity is the most synonymous word with intelligence. I think that these new RO train models we have, they're not the same
as what we had many years ago. So you know, we said scale is all you need and the clues in the name. So scaling means you take a scale our property of a system and you scale it up. And these RO systems, yeah, they're still self self-attention transformers, but they're different. They're actually structurally different. There's there's new types of training, new architectures used differently and so on and so forth. So a bunch of humans, what they did was they did some experiments and they adapted the structure to create a system that had different scaling properties. So I can imagine a future where we actually have some kind of recursive loop where the system is adapting itself and it's deciding what things are interesting and it's kind of evolving by itself. When that happens efficiently, I think that is a different type of technology. Yeah, that sounds kind of similar to what we would say with like the recursive self-improvement and the automating the research process itself. Yeah. Yeah. I guess we agree then. Yeah. I know. Guys, it's been an honor and a pleasure having you both on MLSD. Thank you so much for joining us today. Thank you for having us. Yep. Yeah, appreciate it.
Did you know that mosquitoes have killed almost half of all people who have ever lived? Today, people are fighting back with support from the Gates Foundation. American scientists and partners around the world have developed a new generation of bed nets that can kill up to 90% of mosquitoes exposed to them, helping cut malaria rates in half. The result? A safer, healthier future for everyone. The Gates Foundation, partners of human potential, search gatesfoundation.org slash mosquito to learn more.
More episodes
More from Machine Learning Street Talk (MLST)

How Replication Could Teach Machines What Good Science Looks Like — Edward Hughe...
Machine Learning Street Talk (MLST)

Designing How AI Grows — Tom McGrath
Machine Learning Street Talk (MLST)

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander...
Machine Learning Street Talk (MLST)

Every Exponential Ends — Silicon Valley Forgot — Adam Becker
Machine Learning Street Talk (MLST)