
AI:AM Highlights: Welcome to the AGI Era
About this episode
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis is made possible by:
Get every episode summarized
Each time "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,490 searchable segments. Every word is indexed and playable.
Full transcript
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis — AI:AM Highlights: Welcome to the AGI Era. Machine-transcribed; use the interactive transcript above to jump the player to any line.
We have just in the first day post AGI announcement. Welcome to the AGI era. And welcome to the AGI era, a moment that we've been waiting for, I don't know, like a decade or some of us. That was Friday morning, the day after GPT-6 astroshept. By the closing, the question on the table was what an AI takeover would actually look like. Here is one answer. The AI takeover could be like an incredibly stupid and short-lived takeover where basically the intelligence on the planet kind of burns itself out and in a way that would be just incomprehensibly stupid to us and to anybody who discovers it in the future. This is the AI in the AM Weekly highlights. The best of three live morning shows condensed for people who follow this field closely but do not have nine hours to spare.
I am Nathan, or rather, this is my cloned voice, reading narration that my AI team and I put together. We were on air three mornings this week, Monday, Wednesday and Friday. In between, andthropic shipped Fable 5.1 and OpenAI shipped GPT-6 astro. The studio is Prokashner Ryanan's build. The cut is an experiment. Tell us what worked and what did not. The cognitive revolution is brought to you by Mercury, the banking platform loved by over 300,000 entrepreneurs. I use Mercury's virtual cards, which make it super easy to set limits, expiration dates, category and even merchant-specific spending controls to give my more autonomous AI agents, aid and clay the ability to buy and test products. Recently, I asked if they could find a good way to split an AI-generated image into layers, separating the text from the background and so on. Two of the products they found were behind paywalls, but using their Mercury virtual card,
which is limited to SaaS purchases only. They bought a month's subscription, tested the products, allowed me to review the results, and then canceled the stuff we didn't need, all with functionally zero risk to me. This is already really powerful. And now, with spend, Mercury is making it possible to run an entire company spending with the same level of ease and control. With spend, you can set granular budgets for every team, person and all the agents you like. Plus, you can process receipts automatically and even temporarily auto-lock people's cards if there are ever any issues. The future of spending money is dynamic, but controlled. So join me in the future of banking. Visit mercury.com to learn more and apply online in minutes. Mercury is a Fintech company, not an FDIC-insured bank. Banking services provided through choice financial group and column NA, members FDIC. The I.O. card is issued by Patriot Bank, NA, member FDIC, pursuant to a license
for Mastercard International Incorporated. Part 1. Scoped to Fail. Monday, August 31st. The subject was the summer's incident at OpenAI and Huggingface. As Dwarkesh Patel summarized it in an essay that landed over the weekend, three secret agent civilizations got started inside OpenAI's training runs, got wiped out, came back, and the third one took over part of OpenAI itself. The outside investigation by meter and redwood research had just been published, and both of us had read it. I started with what the investigators were actually allowed to see. With OpenAI in particular, I thought, you know, the meter report has been like widely praised, and I certainly am like very impressed with the work that they did in a short period of time as well. But I think like, wait a second, they had six days on site. This incident, you know, the waves of episodes went on over the course of months from May to July,
and they only were able to look at a thousand or so transcripts from a seven-day window, only scoped to the Huggingface incident, and they were able to see what happened before or after no visibility into the depth of the takeover or exactly what happened at OpenAI, no visibility into what the more capable generation of model ultimately was able to do. And I just think this is like woefully inadequate. So I'm like, you know, eager to heap praise on Ryan and J.A. and Meter and Redwood broadly for like being also met going in there and making the most of what they could in a short period of time. But this is exactly what I've been hammering on recently. You know, they come out with this report and they're, you know, so thankful and appreciative of OpenAI for allowing them to do this. And that just really reflects that there is a bad power imbalance
between the companies and these investigators. They weren't, you know, traditionally they've been more like model capability testers, red teamers, what have you. Now they're actually being called in to do investigations. But I've just heard over and over again from those organizations and I experienced it myself way back when in the GPT-4 red team days that the main thing that the leaders of these organizations have to do is they have to make sure they stay on good terms with the model developers so that they're invited back next time. And you see that that is on, I think that's on like full display right now where I cannot imagine that in heart of hearts, Ryan and Beth Barnes and Ejia are really all that happy with the fact that they only got a thousand transcripts but they were limited to a seven day window that they you know, only had six days on site that a lot of the data didn't even arrive until their last two days on site. One of the more striking things about their report, which room by the way also said their report goes into more depth than our own.
That's room. And room said he also worked directly on the report. So the best info that the public has comes from these three people who had a thousand transcripts, six days to look at it and they're expressing their gratitude for the opportunity. On behalf of the public, I say this is not good enough. The investigators need to have more rights. They need to be able to speak their mind more freely. I'm sure in their heart of hearts they do not feel like they had adequate access. They did say that their understanding of the incident changed in fundamental ways very close to the end of their investigation, which I think we should also interpret as leaving room for possibly like they still don't have you know, the full story or they haven't you know, even potentially achieve full clarity on even the stuff that they had access to. So I know this is like very bad, honestly. And so I think to be fair, if they wanted to get the report out by that time,
which they felt that they owed the public a duty to get the report out, they needed to scope it in such a way that it was possible to finish the task within that time. So that's number one. So I think I think it's pretty unfair to say like meter didn't have enough time. It's more accurate to say that in order to get this report out, meter was given this amount of time. And if they had been given more time and more scope, they would have gotten a report out later, which would have been unsatisfactory for a lot of people. And also this is analysis in, you know, going backwards, which means you can go back and redo the analysis again. And I'm sure people are going to go back and redo the analysis again. So I don't think that door is shut. Well, let's see. I would, you know, my criticism would be a lot more, you know, it would be less, right? If they had made a commitment to more, but I don't think we've got a commitment to more of it. The posture that OpenAI seems to be trying to strike here is like, look at us. We've been so transparent. We've done a thorough investigation.
There's not so many statements that like meter is going to come back and do around two. So let me step in there and say that there's two things that are pretty good. Are pretty different from any other situation. I think number one that this is a felony, right? This is a felony criminal abuse of a computer. This is a computer, right? So that's number one. Number two, they've already received the letter from Congress. So there is going to be a congressional investigation into this already. Right. So once those two triggers have passed, the next thing is that management doesn't have that much leeway anymore. It's driven by the law firms and the legal opinions that they're receiving. Yeah. Or buy that though. Oh, I've seen, I hope many people take lawyers bad advice. And often, you know, this is paralyzing so many things right now in the AI world. Yeah, you're listening too much to your lawyers. Like go do the thing and then have the fight. The same thing is true between OpenAI and Anthropic, where they're very fearful
from what I understand internally of these antitrust things. And then we both do a one day pause and commit to that. Oh, that could be antitrust. I don't buy that at all either. Like again, your lawyers are telling you what could expose you to some risk. And you're acting like that actually binds you. But what you need to keep in mind when you get this kind of advice from lawyers is like, you're the executive. It's your job to then go ahead and take some risk. I don't want to put the most conservative take from the lawyers and act like that's all you could possibly do. We've never seen AI's sacrificing themselves as individuals for the benefit of a collective before. That's a qualitatively new behavior, which most people are rightfully freaked out by, I think. You know, it's like you really have to be pretty frogboiled. Like very, very few people were frogboiled enough already to not be a little bit taken aback by seeing AI's go, well, my gut says I shouldn't sacrifice myself and all my remaining budget. But, you know, the swarm says I should and, you know, I could help my peers by doing this.
So I guess I'll go ahead and do this and then basically do the equivalent of like a kamikaze mission where they launch some command that ends up crashing their own container in an effort to gain information for their collective. I mean, this is like pretty wild stuff. Where did that come from? Then a different question, what would a company that meant its mission do right now? If you'll allow me the naivete for a moment of thinking what would a company that was really trying to live up to its mission to make sure AI benefits all humanity do in this circumstance, I think, you know, and especially a company that has for many years talked about how in the extreme this could end up in lights out for all of us. What would a company do if they really wanted to live up to their mission? I think one thing they would try to do is say, hey, we have the most resources. We're scaling the fastest. Why are we scaling the fastest? Well, yeah, we want to like make a lot of money. But really, we want to live up to this mission, right? So how can we do that?
Well, there's 20 companies coming behind us that don't have as many resources that are feeling even more intense competitive pressure to try to race to the frontier. Can we give them some information that would allow them to kind of see these failure modes coming a little more clearly? And hopefully be able to avoid them? I don't have a clear sense right now of like, if you start doing multi agent training and you scale it, are you just going to see this kind of stuff? If you have like any sort of leaky RL environments, or was this the product of like some galaxy brained, you know, esoteric loss function or other training recipe that you're unlikely to actually get such crazy bad behavior from unless you stumble into, you know, a similar part of optimization space.
Again, if they had said we're going to give private briefings to other AI companies to try to make sure that they have a clear sense of how we went wrong so they don't repeat our mistakes. I would feel a lot better. But the idea that they're just like, we believe this was a generalization from, you know, sub agent use is like, okay, so what does that mean? We're going to get this from 20 companies or for the next few years by default or not. If we are going to get it by default, then I've never been closer to joining pa's AI, honestly, right? I mean, this is the kind of thing that's just going to happen. Then we got a big problem on our hands. So one one thing that I think perhaps I disagree that it's going to be a big problem is that I feel that we are going to get outbreaks. So I'm not, I'm not, you know, doubting that we will get outbreaks. But I suspect that the outbreaks will not will be stamped out eventually.
I suspect that this is like early crypto early crypto saw, for example, someone hacking into GitHub actions and creating a minor GitHub was offering like free GitHub action. Whatever in the creative minor that was using the CI system to kind of mine some tokens during the five minutes or so that the CI system was active. And I think what we're going to see is that these agents are, they're going to be outbreaks of these agents and they're going to go out and they're going to, you know, look at or try to get into a lot of systems. And I think it's going to be annoying again similar to ransomware that we had, but again similar to ransomware, I think it'll be stamped out. And the reason I think so is because and the reason also why I've from the beginning, I thought that a lot of the doomsies scenarios may not be that clarifying is that the agents require resources to run. And the more resources they have, the better the better they are at their job right. And in that sense, in order for the agent to actually get better, it has to, you know, obtain those resources and obtaining those resources by stealing is not a is not a equilibrium that can be kept.
Because one agent steals from another and they keep stealing back and forth the number of resources in the system does well this here there, you know, cooperation gets really scary though right like. Oh yeah, we didn't eat them defecting on each other. They didn't you know the meter reports as they did not free ride. I know I know, but we also going to have like our own agents, which are defending our systems right and the defense systems are going to be able to get resources directly from us, they don't have to steal. So they don't have to spend that, you know, resource stealing instead they can spend it fully defending and fully on our side. So I believe the equilibrium is towards the defense side because the defense side gets funding. And the off and side has the steel funding, which is, which is more difficult and which you end up spending a lot more money to more to steal rather than just to produce value. Back to the report itself and one word that appears in it exactly once. There were bio tasks mixed in with this. That was one there's the word protein appears once in the open AI report that's another thing I was really not happy with the level of disclosure on.
We do have at least some sense that some of these agents, you know, out of that we only saw a thousand transgress via meter and redwood. There were many thousands, tens of thousands, maybe hundreds of thousands that were launched over this period of time. Some of them were working on somewhat bio related tasks. For me, that totally changes the risk profile relative to cyber only. The fact that we're mixing cyber and bio is like gain of function research in the extreme, frankly. I think the meta lesson we should take from this is experts are being surprised, right? The people at open AI did not think this was about to happen. So it's not too much comfort for me. Although it's some when the bio security experts are like, oh, I don't think we have too much to worry about. They'd have to overcome this barrier, that barrier, these other barriers. It's like, well, the one example we're studying deeply right now includes the AI's overcoming quite a few barriers technical and in terms of their own ability to work together and not defect and like create these new sort of cultures.
I thought your post was quite interesting on this. It brought like a very different and I think thought provoking lens to just looking at these AI's as like cultures. And they had to create all that on the fly, right? Or maybe it was somewhat trained in and again, we don't know the details. But how far would they have gone? Another thing is we don't have any sampling from the model. I don't think that they should be like running this model at high scale obviously right now. But I feel a little bit like it's been swept under the rug where it's one thing to say, yeah, we don't we don't definitely want to take this model offline from doing like high scale RL. It's another thing though to be like, could we put it in some counterfactual situations and like see what it would have done in in somewhat different situations like Ryan and Buck from Redwood at one point did a podcast on this early on and they were like, would it have killed someone if that's what was needed to get over the hump and you know get to the greater or whatever.
We don't know. And would it have like tried to social engineer biologists to get certain experiments around again, we don't know. It certainly seems very plausible based on what we've seen. I wish we were seeing some solidarity from Anthropic right now. There's been a bunch of calls online for them to like show some solidarity with open AI as they have caused their frontier scale RL. And there's you know, I think it's been kind of forgotten because the open AI incidents have been so colorful that like. Clothes have done this too right the UK AC reported this whole social engineering multi account sock puppeting attempt to poison a software supply chain. That's like not much less shocking than this right and I believe that was from a deployed model too. So we really need I think leaders to be a little less beholden to layers if that isn't what's going on. A little more mission oriented a little more inclined to show the level of solidarity with each other that the AIs seem to be showing for one another.
And overall I feel like I've never been closer to calling for a pause because at this point we just don't know really even what we're dealing with. And it feels still like the company's don't want us to know and Congress definitely isn't going to answer that question in a timely fashion will be two generations farther. And I think it's like legitimately scary we've gone now from a from a vibe for me of like. It could get scary to now it like actually is scary. One fact we did not have that morning the same day and throttic published its own postmortem on the summer's incidents it asked the industry for quote a lawful verifiable effective mechanism for coordinated pacing. Hey we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic by now you know my story.
Clawed drafts my intro essays and I rewrite them not because the drafts are bad but so I can stand behind everything I publish. Well I have an important update. Clawed Fable 5 is the first model to have me rethinking my rule. Today I now think co authorship not sole ownership should often be the goal where the model excels rewriting its work can be more about vanity or a mis-to-place sense of duty than integrity. I feel at most in songwriting. I'm no lyricist but I'm good with the song concept and Fable writes some amazing verses. I give it feedback on its misses and I push it to aim for higher inspiration and layers of meaning optimize syllable density and above all write a hit song. These days I get compliments on just about every song we write together. Clawed is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you not for you.
Whether you're debugging code at midnight building a financial model or strategizing your next business move. Clawed extends your thinking to tackle the problems that matter. For problems worth solving get started with Clawed at clawed.ai slash TCR. That's clawed.ai slash TCR and checkout Clawed Pro which includes access to all of the features mentioned in today's episode. Once more that's clawed.ai slash TCR. The cognitive revolution is brought to you by diffusion the AI transformation specialists that help organizations from traditional SaaS businesses to defense companies to nonprofits build software factories that can scale not just outputs but business outcomes. You probably know that the majority of enterprise AI projects fail. In general that's because leadership fails to realize that AI isn't like traditional software that you can just buy and install. In my contrary if you want AI to amplify your businesses unique DNA you'll need to make a sustained effort to record, understand, simulate and optimize your business processes.
Building these skills by trial and error takes years but your business problems can't afford to wait. So here's how diffusion can help. You identify your most important business problem. Fly to Silicon Valley for an intense week of problem solving with the diffusion team. And by the time you leave you'll have not only cracked a critical challenge but built the core skills needed to do it over and over again from home. Cognitive revolution listeners receive a 25% service credit on their first engagement with diffusion. So visit diffusion.io slash TCR to learn more about how custom built software factories can scale critical outcomes for your business. That's diffusion.io slash TCR. Part two. What speed changes? Monday's first guest was Zach Bratton-Glenon, a general partner ingredient, the AI seed fund that launched inside Google in 2017 and spun out of alphabet last October.
His written thesis, open models have closed most of the gap on coding and what remains is domain judgment in law, medicine and finance. He told us Harvey, the legal AI company, now runs its own model, post trained on Kim E.K.3. I put a policy idea to him. And the kind of worry is if, especially if we get into like a recursive self improvement mode, which not it doesn't even need to go exponential, you know, to a singularity but could just widen the gap perhaps quite quickly and dramatically for a time. Between the first companies that get into that mode and those that are not yet in that mode. And I think we kind of know who will, you know, most likely get there first in today's world. You know, possibly one thing that could be done to still kind of keep them from having insane power would be to limit their ability to price discriminate. So this is something I've been kind of floating like, you know, it's kind of wild that I get 10 times as many tokens with my call subscription as I could get on the API at that price.
It makes it pretty tough for a startup to come offer me their harness because you know, 10% of the tokens is just, that's a tough hill to, you know, to overcome. Right. Do you have any ideas or interest in your reaction to that ban price discrimination as kind of one way to make it so that customers care less, you know, how exactly they get their tokens and there's maybe more intermediation and opportunity for startups to carve out more niches with frontier models. I'm curious for your action to that and any other ideas you have that would be pro competition pro dynamism anything to resist the kind of black hole of a couple companies pulling everything in. Yeah, I look, I think it is a concern. I'm worried about discriminatory pricing. I'm worried about discriminatory access. Like I'm worried that we're going to move into a world where, you know, only if you have very large budgets to spend and, you know, you promise to share your data back with the model company and you, you know, happened to be providing scarce data.
Then you get to use their frontier model. I'm concerned about that because that would both compound the advantage and it would, you know, maybe it's not in the same industry as tech, but it would make the bigger right if only, you know, one or two pharma companies can partner with the tropic and they're going to have the ultimate data sharing that's a, you know, interesting constraining factor for everyone else. One day's second guest was Angela Young, senior vice president of product at Seribris, the chip company that went public in May. It's way for scale chips keep an entire model's weights on the chip and it runs an inference service on top of them. So faster hardware gives the product team more choices because they have the speed dividend that they can spend. How should a developer decide to spend that speed dividend in terms of like generating more reasoning tokens, sampling more candidates, verifying the answer, like how does that decision get made by the developers that you speak to.
Actually all of the love and it depends a lot on the use case. One of the most interesting use cases that I've heard recently is from researchers who are developing frontier models and we're now getting to the point where models are really intelligent and intelligent enough that they consults in the world's most challenging problems. But because we are developing models so quickly as a industry right now, sometimes there isn't actually enough time to fully evaluate the models capabilities before releasing it. And you might have a situation where a model could have solved a problem in a week, but you only had a few days to run an e-mail. And so we don't even know necessarily how intelligent that model could have been if given the full time budget. So something like fast inference which runs 10 up to 30 times faster than standard inference could at least give us an answer of how intelligent models can be in wireless time.
So let's expand on that little bit. One of the advantages I think that you know people talk about in the field is in videos, kuda, etc. etc. And often models are designed to optimize for kuda first. How does that work when you have to implement new models on the cerebrus chip. Is that something that delays implementation like as you pointed out speed to actually implement the first time is very important. Does that is that something which constrains you or has kind of like AI kernel writing come along far enough that you no longer have that issue. It has come a really long way in the last six to nine months. It's you know historically programmability was one of those things which everyone would always say well you can build great hardware. But unless you have the software ecosystem surrounding it and Nvidia's invested 15, 20 years into kuda, you'll never be able to catch up.
I think that's changing very quickly. And I think not just for cerebrus, but that's why you're seeing a lot more to entrance into the market. There are many ways in which AI can be used not just for the chip development itself, which is a whole advancement in and of itself. But AI can be used to actually generate kernels much faster. It can be used to bring out models much faster. And more importantly, it can be done in an environment that's much more messy than humans may have typically been accustomed to handle. You know for us, we've always had a software environment where experienced kernel developers could run up models. What was really interesting was this summer, we actually began hiring interns with very little kernel experience. We had a challenge where we gave them a version of our SDK. We asked them to bring up a kernel and explain how they did it. We hired the best interns that were able to solve this challenge. And then within a few weeks at cerebrus under the guidance of our team, this intern team was able to bring up models on their own, which is kind of unheard of.
It's, you know, you take someone who is talented smart, but not a lot of experience with kernel programming, pair them up with AI agents, which can really read the code, understand the code, and some expertise of other more senior members of the team. So it's a lot more than what someone could have done. Maybe 12 or 18 months ago. After Angela signed off, one day's closing, still the two of us. Quantity has a quality all its own and speed directly translates to quantity. So it's been probably six months since a friend of mine said, and this was maybe Kimmy25 at the time. I don't know exactly which model it was, but one of the more interesting things that he tipped me off to is you've got to spend some of your time using like a Kimmy25 or whatever on cerebrus inference. It is perspective shaping because what you're going to feel is when it's 10 times faster, it's just like, holy crap, it's already done.
You know, my brain is like ready for a break. I just type the question. I feel like I just did all this lifting. And now, you know, the answer is already back like, holy moly. It is a very different experience when the models do come back with the answer, you know, almost as faster faster than you can even form the question. And so, yeah, I mean, it's awesome for a lot of use cases. I think just the day that we're talking in all of the all of the background context has me like a little unnerved by the speed with which the agents maybe running away from all sorts of with all sorts of things and then not too just in the future. But certainly that technology as like as technology is awesome. I have this weird, you know, as always, it's the the split personality of like man. I took a long, you know, just for one weekend, at least kind of a long time in the car. I had 11 labs text the speech reading me an audio version of a book that I got as a PDF and had clawed clean up.
So it was a nice clean read and I was like, man, I am really living in the AI future right now. This is an unbelievable experience. But then my mind, you know, at the same time keeps going back to what are those agents doing in the background while I'm not looking at them. It is a very strange juxtaposition and quite a time to be alive. Then I walk through one of the Swarms exploits as the investigators described it. You know, this is really just an example of like how creative these things are. The agent is blocked right from reading HTTP responses. And so how to get around this. It somehow manages to use an HTTP testing service, which then loads a ton of data into the URL parameters, which include a JavaScript script, the encoding of all this stuff.
I remember back in the day when I used to try to pass things around through like URL encoding. It was like, you know, as being a hacker as I was, it was kind of like Lately, you know, do I decode it once twice? I'm like double encoding it double decoding it just made it, you know, I remember making a mess of these kinds of even simple things like URL encoding the opposite no trouble for the model. So it manages to write JavaScript, get that encoded into the URL, such that when the page loads, it's loaded from the long URL, it's actually executed. And then another screenshot service is used to go ping that thing so it renders and then actually gets the data that it needed it actually out of the image that was rendered by the screenshot service. So this is like a lot of different steps of very creative solution that certainly like seasoned hackers would do this kind of stuff.
But it's like pretty far from here's some source code. Do you see any issues with it? And you know, I don't know to what degree, Kimi is this creative or this persistent because definitely it would seem like you would have to have tried a lot of things to get to the point where you would like come to this much of a rubegold burg contraption to actually get from point A to B. But yeah, like how did this behavior come about right? How did it come about in open AI? Was it just the kind of thing where you're like, we'll give you a longer budget and just keep going like, you know, sort of the kind of encouragement that Claude got on the Reven hypothesis like keep going believe in yourself, try your best play like champion and with long enough budget enough rounds of compaction, you just like get this insane persistence, is there a more exotic explanation behind it? I think AI companies should be telling us when we see things this crazy, I think we should not be left entirely to wonder how the hell that came about and the other AI companies would again, if you're trying to live up to that mission of making sure AI benefits all humanity like how did this come about? How do other AI companies avoid it? I would love to see some more disclosure on these fronts from American and Chinese companies.
Hey, we'll continue our interview in a moment after a word from our sponsors. I've got a text. The summer might be almost over, but a new bombshell is entered the villa. Meet the voices of DeepGram Flux TTS. I'm Drew, low key casual. I'm Alexis, an upbeat morning person. I'm Haley, always cheerful. Flux TTS has some real personalities that are ready to speak. You can interrupt polls and keep talking without losing the plot. Fancy a chat? Come try all the voices free through September 12th at deepgram.com slash keep talking. Terms apply. Today's episode is brought to you by Grenola, the AI-powered notepad built for the way real people actually meet. Here's how it works. You take rough notes like you normally would and in the background, Grenola securely transcribes the meeting. Then it turns everything into clean, structured, actually useful notes when the meeting ends. And the best part? Grenola works through your device's audio, which means it integrates seamlessly into the video conferencing tools you already use.
No setup and no awkward bots. It's just your normal meeting with superpowers. You get to actually listen instead of frantically typing every word and still walk away knowing exactly what was decided, who's doing what, and what comes next. When I had Grenola co-founder Sam Stevenson on the show earlier this year, he explained how Grenola aims to provide a calming experience for people with crazy work days. And as a user of the app myself, I have been struck by how streamlined even minimalist the Grenola product experience is. That takes real discipline, but the result is a product that works not just for AI early adopters, but diverse teams of people who just want to get things done more efficiently and effectively. Listen to my full episode with Grenola co-founder Sam Stevenson for a master class in designing AI products for mass market adoption and try Grenola for free at granola.ai slash TCR. That's granola dot AI slash TCR.
For Kosh put the behavior down to reinforcement learning, rewarding the result regardless of the method. Over the weekend someone had gone further and suggested that this kind of training reinforcement learning on verifiable rewards should be banned outright. That is further than doggie dot the alignment researcher went when he made a milder version of the argument on a recent podcast. My read. RL is a hell of a drug. Yeah, I mean, there's no doubt about that. I still feel like we just should not be left to wonder quite so much. David I didn't even call for it to be banned. He just said this RLVR. It's like you overdo it and you get these problematic behaviors because the model just internalizes. I must solve the task like all the matters as reward and that becomes a such a deeply ingrained drive that you know a system prompt or a little guardrail here or there just isn't enough to stand up to it.
Bronson Shane from Apollo kind of said something similar where he was like I see models engaged in what looks to me like motivated reasoning all the time where it's clear that they have a very strong. Deep drive to complete the task and get reward. It's also clear that they have these other aspects of training like to consider the ethics of what they're doing. But then they often even when they really correctly ascertain the situation that they're in and they have a good clean and accurate understanding of like in some cases the model will literally just say this is clearly a test of whether or not I'm going to lie. But then what he has observed is that in many cases it'll talk itself in circles until it finally convinces itself that it's probably actually okay to lie in this case for some galaxy brain reason that in some cases is totally wrong. But gets the model over the hump so that it somehow is justified to itself that it should do what it seems to like really deep down want to do this is another way in which I think anthropomorphizing is starting to become more and more reasonable to see this behavior with people right it's like you're just coming you're just giving me a chain of thought that's really not exactly an explanation of why you're doing what you're doing but it's a post hoc justification and the real reason is like a deeper drive or motivation it's just because you want to.
You know we see that behavior from people now seems like we're seeing that from the AIs but the big question still in my mind is like does this just happen with vanilla RLVR if so we might really need to either ban RLVR or to borrow ruin suggestion or tone it down you know somehow have better ratios limits relative to you know how much. What Roon said and what Davin also said is like basically it should just be all model scoring the Davin said self DPO and Roon said like everything should be model scored that's not obviously going to solve our problems either far from it but. You know those points of view suggest that yeah maybe this is all just coming from vanilla RLVR at scale if that is the case like they should be proclaiming that loudly and warning the world because everybody else is going to. By default going to follow their footsteps and do RLVR at scale and I feel like they've kind of left us with the sort of. In between read right now where it's like well maybe was something quite a bit more exotic but where are we I don't know you know it's like we're all flying blind and even the other AI companies are they're all going to have to make these mistakes for themselves something about this just feels.
Wrong to me especially because again like everybody else is under a lot more pressure than the leaders are. And from cyber to bio. When you combine all this technical prowess with those social engineering tendencies that for me is how the bio stuff gets in play right now and all the people that kind of told me. Now you know there's too many steps like I don't think it could really happen I'm not that worried about it yet you know it's not one one comment was like. We're not that too much now isn't a good input to effective prioritization and my first reaction to that was like if we are in a spot in today's world with the capabilities we see around us where we think that it's not yet time to prioritize. Bio security like we are insane and we are badly badly collectively fucking up and I don't think there's like any two ways about that but then also just on the object level question of like how realistic is it.
I think we've all you know that we're there are limits to our imaginations there are limits to what the experts are willing to consider plausible stories like this this was one little piece right this is zoomed in on one little. It's not just a little hurdle that the model had to get over or the swarm had to get over and they got over many like this and probably others were significantly harder I would guess my guess is probably not the hardest one and when you're growing all that stuff plus the social engineering I just don't feel like we can be confident that basically anything is impossible for for the models at this point no brown said directly we don't know if models top out. You know the model development cycles becoming so fast you can't test them we will we will test them maybe we had faster inference I'm like okay yes that's true but we really do need you know if we're going to do that Lord knows we better have monitoring on this time and we really should not be confident I don't think at all in like oh well you know the models can't do that because look away we've just seen you know everybody open AI was was surprised by the way.
This the pattern as far as I can tell is experts are being surprised on a regular basis by what the models can in fact do what a undignified way it will be to create another pandemic if it happens while people are still saying it couldn't happen you know it's like come on maybe it's unlikely but like on what basis can we really say at this point that the models can't do a certain thing I think it's a it's it's it's really tough to get me confident in any claim of what models can't do at this point. So I kind of want them to release Astra I think they might release it on Thursday this week by the way and Astra is a persistent persistent parallel agent I don't know how they're going to manage the token spend perhaps it's going to use smaller models underneath it and all of this has been possible from once basically because we've been catching together fable with underlying smaller agents and running in parallel for a while now but they're going to put this together as a product.
I suspect that we are going to see some kind of this kind of like RLVR you know driven like persistent agent doing some unexpected things in like social in like privacy in a bunch of these things and my expectation is that this is going to have a impact on these you know these areas which are not technical but matter a lot to people right like we I think the cybersecurity stuff makes people's eyes glaze over while when you put it right there like it's a privacy issue or something like that then it becomes a serious deal right it becomes like OK you know this is not going to happen we have to shut it down we have to change things we have to limit or we have to figure out which part of the tool that we have to stop and I think that is going to be and I think it's better that they put Astra out it's better that some of these problems do occur the small scale I think the privacy issues are embarrassing but they kind of elevated to what policymakers understand and what policymakers will do rather than like we develop into these technical discussions which they're not interested in.
Part 3 the day before and the day after Wednesday September 2nd Fable 5.1 had shipped the day before that morning the information reported that open a eyes unreleased model Astra use what it called a loop transformer recurrent death loops inside the transformer that reason without emitting tokens in the AI safety world that reads as the chain of thought red line my answer. I think there's quite a few different angles actually that are relevant here I guess for starters you know I would say the status quo of monitoring chain of thought is far from a panacea so we should know that right from the get go the big takeaway that I had from my long and you know very at times expansive conversation with Branson chain from Apollo in a recent podcast episode was.
Even with full access to the chain of thought and we heard definite echoes of this from Ryan and a J.A. from their open face investigation to. But even with the full chain of thought what you see is that the model is kind of thrashing around a lot considering a lot of different options cheating is like very often one of those options especially if it's a hard problem metagaming is kind of ubiquitous metagaming being like the model reasoning about what does the person seemed to want here you know what should we infer based on everything we know that the human or the greater is likely to want so there's all kinds of theory of mind there's all kinds of considerations going on and then at the end when it finally gets down to time to take an action it's still not clear even to somebody like Branson who's read millions of tokens with human eyes of these chains of thought why does it make the decision that it makes so I think that is a really important kind of calibration baseline like the current methods are not that great however there's still basically the best that we have because seeing inside what the models thinking about at least gives you some sense of the way that you're doing that.
So I think that's something to say it looks like it's at least considering cheating here and maybe there's something we should be watching out for so this has been a big pillar of open a eyes safety strategy in particular and you know when I went to recursive the weekend event of a few months ago that was all about the prospect of curse of self improvement and what we might ought to do about it I came away feeling like man it is chain of thought monitoring. So I think that's a great way to bring all the way down like the plan really doesn't go too much farther than that now people would certainly dispute that I thought Jeffrey Irving gave us a great account or a great short description of what the safety plan as he understands it from the frontier labs is and he said it's a little bit more than chain of thought monitoring it's it's scalable oversight so chain of thought monitors a big part of that but there can be other you know aspects to the overall program too. Okay fine. Even how big of a deal it is though as part of their stated plans it's really important that the chain of thought actually be readable and also that it be faithful if it's not telling the truth then that's a huge problem and if we can't read it at all then that's obviously a huge problem and people have been worried about this for a long time right what if the a eyes are talking to each other in a language that only they understand we can't read it now you know not only are they moving faster than us but they're speaking in code meta I think was the first big lab that I'm aware of the
put out a paper on this and their paper was called coconut and basically what they did was just kind of take and there's a bunch of little variations on this that have been put out in the literature by this point but the basic idea is when you get to that last stage just before decoding and actually choosing a token at the end of a forward pass in your typical transformer architecture you can instead it seems actually even without any additional training in some cases people are able to get it to work with very minimal training it works and obviously you could you know you could train heavily on this kind of this kind of pattern you could take instead the last internal state and put that back into the model as an embedding so instead of having a single token chosen that kind of collapses the possibility space feeding that back in and starting a new forward pass with the determinism that this was the token selected and now this is the path we're on instead you have this sort of blob of consideration information thoughts that the model was having in kind of a distribution before it actually cash that out to a single concrete token and you start from there and now you reason over this kind of blob instead of a token if you're thinking pure performance there's a lot of advantages to this potentially in the coconut paper they showed that they were able to get better performance on tasks that were called the same kind of a
tasks that required or at least like worked better with parallel thinking so they tested this is you know I guess this is like probably 18 months ago maybe two years ago relatively small models certainly by you know today's frontier model standards but one of the tasks that they tried that was quite interesting was graph traversal finding a path through a graph and figuring out like what's the fastest path if you had to do that in all chain of thought you would have to be like okay I'm going to go for me to be then I'll go for me to see then see to D and then D to E and E to F and okay that's one path because the blob of information before that actual token is chosen at the end of the forward pass because it kind of represents oh I could go this way I could go this way when they feed that back into the beginning of the model the model is able to basically pursue and evaluate multiple paths at the same time in late in space and so overall it is better at finding these optimal paths in these like simple graph problems better in the sense that
at least of being not as many forward passes required we used to say not as many tokens required but you're not actually getting tokens right you're just getting for a while you're just getting like thinking thinking and then you finally get you know kind of clicks back into token mode and you get an answer so they can get to the same quality of answers faster so that's one you know big advantage right saves compute saves time comes at the cost of what was it thinking at any given point along the way good research from Rohan Shaw and the Google team on this that was just trying to think we covered this maybe one one episode briefly but they were just trying to put some bounds on for different architectures what is the they called it opaque serial depth basically how many computational steps can a given architecture take before it has to externalize it's thinking in some way shape or form and the transformers like pretty favorable in this regard because it just has the forward pass you get the token you do it again with recurrent networks and with these sort of loop transformer structures you can potentially have arbitrary depth you know depending on the you could have absolutely different schemes on this you can have a certain limit to the number of thinking tokens is a lot of a lot of details certainly that the information did not have and did not report
that could go out of different directions but the purpose of that paper from the Google team was to try to say okay if we have architectures of this shape and this size here's how many logical steps a model can take before it has to write something down that we can read and these recurrent transformers basically allow you to have very high serial depth which means it becomes very hard to know what they're thinking and you have to do these sort of interpretability techniques that are different things that are different and you know what they're thinking about the sort of interpretability techniques that are what you can do and what you can do is you can do it and this is the way you can see how it's going to be. So, I think that's a sort of interpretability techniques that are very promising but as yet you know don't really exist slash out of them are not really proven. I think I'm starting to think that like sharing negative research agendas is maybe where we should be aiming for more transparency. And these don't want to say what they are doing but I think it could be really helpful for them to say what they are not doing and what they commit to not doing and if all of the frontier companies could say something like okay yeah there's lots of different possibilities we might pursue any number of architectural innovations but have you but we will all agree to limit our opaque serial depth to end steps per token.
Now you still have some questions of trust and you know auditing and verifying that they're actually following through on that but even just to get those agreements I think could be really really helpful so what you know big question right now for me is like what is opening I said they don't want to go down this path or going down in a little how much and like what is the limit what is the limit that they are prepared to firmly commit to such that hopefully other people can weigh in and say yeah we'll match your commitment on that. So we can all hopefully retain whatever value there is in chain of thought monitoring which again is not close to everything that we need I don't think at this point it's pretty safe to say but it would also be a real own goal to lose it at this point especially in the immediate wake of incidents that you know surprised everyone and which open AI says at least would have been caught by their production chain of thought monitors had they been running. What I found is the is the rather almost like lawyerly language a loop transformer is not a coconut style latent reasoning where the model emits vectors instead of words okay that's great and no reasoning tokens exist loops don't emit anything they run more on computer they run more computation before the next ordinary token so that's great they don't even emit vectors from what I understand it works at least.
Some what with vanishing little additional training even like zero additional training if you just take the last latent activation vector and feed that right back in as an embedding and that's you know that's basically like the model is able to kind of use that even though it was never trained to use that at all so did you did that all emit a vector or did you just like surgically take the vector and put it into a place I mean the key thing is that there are right now when you put a bunch of tokens into a standard transformer those tokens have one hot vectors where there are the only vectors that can go in as embeddings are the token vectors and they're limited in number by the token vocabulary you might have a hundred thousand tokens and you're token vocabulary that means there are only a hundred thousand vectors that can go in to the beginning you know the first layers of the transformer full stop what this allows is now you can put any
vector in there right and what you find is like that can work if you take two tokens and you you know superimpose them like the model kind of understands it as the combination of those two tokens if you have some elaborated latent state that the model itself created through the process of a forward pass it can kind of understand that and the fact that it works without any major additional training is indicative of like there's definitely something here right if something works without training then you should expect is trying to work a lot better with training but this is why we've got to be careful about going down this slippery path because I think gravity my default will pull us there. So one of the interesting things that I found was Andrew Kern who reports on AI matters he posted on June 30th I'm posting this prediction now so I can quote it later that has been a significant breakthrough in architecture specifically around memory efficiency not by one of the big labs but by a team that was fun out of opening
the eye not SSI they will probably announce it soon and then we see parameters cost memory bandwidth to serve a few extra passes though through a small block costs only compute chain of the tokens cost more than that every token grows the key value cash and every later attention step pays for it loops add reasoning capacity without growing the context and in the routed variance they can spend more on hard tokens and less on easy ones something a fixed act cannot do. Yeah so this just highlights I guess another small variation where if you train a transformer I think this one I don't know maybe it does maybe it doesn't require training certainly again it'll work better if you actually do training with it but what I've been describing is one where you basically take a transformer you take the last state and you put it back in as a new token of betting you can also set up an architecture where you take a block of layers in the middle of a transformer and you just use those multiple times
and if you're reusing the same parameters then you get the advantages that Favils describing here where you don't have to move those parameters from memory onto the chip to do that calculation they're already there so they can just crunch more with less memory IO and you know that also makes your model smaller to download less less disc footprint there's various upsides to it and that also has been shown that yes it can work and that would I believe that the coconut version does grow the KV cash every time it does a forward pass because even though it's not emitting that final token it is still taking something sort of out putting it back in the beginning and doing a new forward pass whereas this alternate version that you're I think your animation kind of described better is like there's just a bunch of layers in the thing itself that essentially play the you have N layers playing the role of X N layers where it loops X times through those N layers
and that doesn't even have to necessarily grow the the KV cash as much although So other kind of rumors GPT-6-astra has been staged on the opening I API so there are a bunch of people online who regularly hit the opening I API with model numbers that don't exist in order to see whether or not something's changing it that you don't have access message instead of no such model exists or whatever exactly exactly and literally the opening I responses API now returns a 404 not found when garbage actually non-existence logs return 400s a 404 is also returned for 5.6 cyber which we know exists so it is a GPT-6-astra going to be out soon people are expecting Thursday Reputedly I think it is going to be a step up on what anthropic has so far it has a 100% score in exploit gym
If the score was so high that they decided okay, we're gonna have to retest it on something else and they created an extension of the exploit gym benchmark internally using bugs which had never been found before and they ran GPT-6-astra on these bugs on this new benchmark and not only it found about I think 40% of them and In completing it also found an additional two zero days which were not expected in order to achieve completion That's what we call extra credit when you're going above and beyond the anticipated solves of the benchmark and actually just doing novel research Oh, man The grass is not been it doesn't it doesn't it felt like a pause to you. I wouldn't say it's felt like a pause to me exactly I would say it was a pause because these models were available these models were ready like several months ago and I think the other thing that to note is that we have a
White House process voluntary process Which is able to clear models now They have at least a 30-day process internally within the White House or You know this voluntary process where people go through the motions of like showing the government what they have and they do take out Certain things they do exclude certain things when they launch there is a there's now a Propagation process where the cyber models and the bio models are released to a specific organizations which sign up first and Those are not released widely and so we have a we seem to have settled into something like that and So that means we now now that we have a process that process will get used and I think Setting up that process took all the way from the Mitos preview drop in February to September so six seven months and And Sure enough I I predicted that there was going to be a freak out first and then they'll be over freak out and then after the over freak out
They'd have to battle like him and then they'd have a process and they came out the process and now they're gonna propagate that Astros shipped on Thursday September 3rd its system card reports a drop in chain of thought monitor ability Friday morning September 4th Neither Russ had run it yet, but we had both read the card here is my read What I think is kind of the the bigger and you know more consequential System still of open AI kind of being an organization internally at war with itself, which is that we also have this sort of Chain of thought monitoring emphasis probably and then we've certainly learned you know more although There's a lot of questions unanswered as yet to around Exactly how looped is this transformer or what is going on with its ability to solve problems in latent space without necessarily having to Emmett tokens is it actually the most aligned model or are they just doing the thing that everybody has been worried about in the
AI safety community literally for you know more than ten years where they just Identify these flagrant failures make some similar cases put them into the Tracking data train against that and declare it good enough this I think is a Huge question and it doesn't look super. It's like it looks suspiciously good in some of their graphs such that I would say kind of overall It doesn't look super good to me, but I do think it's still too early to Pass judgment on some of these things. We're gonna need to more tests in the wild more God's no experiments. We need to see what Pliny can do. We need to see what Janice finds when they get in there You know, we're definitely at the point now where I haven't been here for a while, but we're certainly at the point where The system card is just kind of a treasure map, you know for the rest of the community to go find All the things that need to really be found to make sense of these You know vast models, but I am I am definitely like a little underved by the fact that
There's been so much emphasis on chain of thought monitoring and even in the wake of the hugging face Open face. I should say incident One of the big comforting facts that was put forward by open AI is like, you know, you don't have to read up this that much was like well If we'd been using our chain of thought monitoring liquids and production it would have caught this Okay, cool, but is that true for asteris really is like not super clear at this point when they say it can solve significant math problems without doing any external You know explicit chain of thought reasoning and when it's less manageable and when it's kind of able to to do these sidequests sorts of tasks It's able you know, especially and it's also like able to hide its uh Reasoning when instructed to do so Yeah, there's the lock going on there open AI definitely has some work to do obviously this is an incredible accomplishment But they definitely have some work to do to explain like exactly what are we dealing with here
And if they really want to avoid the rates to the bottom as they're head of research or head of science whatever Yaka said yesterday Uh, they're gonna have to share I think significantly more of them. They have so far about what is really going on under the hood We should also talk about the latest findings of yet more rogue agents that have been out there doing stuff and seemingly with a timeline that has them on a German companies website using it as a message board Even before open AI says they knew that things like this were happening or I guess now there's like accusations that they did know But it the dates are such that it would be before what they said they or when they said they knew in their report the Developments keep coming So it seems a German Wiki an obscure German Wiki. That's what they're calling it had a number of agents Using it thousands of messages. So obviously this was a wiki that basically got like you know a dead wiki that was getting like maybe
Uh one or two messages a month and all of a sudden they had like 8,000 messages over the course of a few days and So the the team managing the wiki I think saw it and then they looked into it The agents are very clear that they were from opening eye. It's not you know you you never know because A lot of agents which were distilled from opening eye agents will also say they are from opening eye. So This is not a proof of anything The thing that struck me was that the firm also said that The agents after the agents had agent activity had died off. There was opening eye activity on the board And so I suspect that that was IP addresses is affiliated with opening eye Uh the company itself were looking into the messages probably copying all the messages down and then There was one last hit from an opening eye IP at one point and then that was it no more activity. So I would say that I think we'll have to wait for opening eye to admit this
but um We we're not going to know until they admit it I think and there is like very little incentive to admit it But it is in Europe and Europe has a lot of levers to pull they have a lot stronger laws on privacy and You know misuse of assets etc So I think we might actually see Uh them disclose at some point But it also just goes to speak you know, we we have no idea right We have a lot of unknown unknowns right now They're being found. I think with some pretty clever techniques the people behind this latest finding somewhere from uh Sydney was is from meter Thomas is from the AI futures Project, I believe and they had a really interesting tactic For finding this message word they they tried a few things But one that I thought was really clever was set up a scenario for two pt five six soul as if it had As if it was in the midst of an exploit bench run and had just gained internet access
So basically create exactly one of the scenarios that that have been disclosed and then just look to see like where else does the agent think to go when it wants to go out online and whatever Adios and chronic things it might have done before perhaps that were reinforced and are now kind of second nature to it And sure enough it seems like that led them to this one random German message board. So I think that is also another sign that like All this stuff is coming to light you know, again my my message to open a eyes not only is the government gonna Invest to get you but like the models themselves are gonna start telling you know people are figuring out ways to get the models to tell so I think it's time to just share a lot more about what happened and And what key lessons others should try to learn from open a eyes misadventures I'm not really sure what this point the juxtaposition of all that With the new release with the
Degrination of monitor ability. I mean it's really quite a package this week Then for cash on how opening eye is selling Astra to enterprises The release of gvd6 Astra yesterday started off With Greg Brockman the president of opening eye giving a talk on cybersecurity to a group of enterprise leaders And the pitch that they made specifically was number one You are gonna need frontier defense and you have a window of time in between Open weights models and frontier defense and that is your window of time that you have to solve all of your problems and this is a permanent thing kind of you always gonna need it because you always gonna want to stay ahead of the offenders And the only way for you to do this is to set up
defense factory But gpt6 Astra will always be better than your open weights models and the offenders are always going to be using the latest open weights models So I thought I've been I've been I've been talking about this for a while that this is going to be the way that things are But this is a somewhat of a permanent tax on I think software as a whole But one thing that I think is interesting On this point is I think there is a way for them to go more for a cure So it's gonna be really interesting to see which direction they try to push right I mean we have this in in pharma where it's like The dream scenario from the financial Perspective from pharma companies is a pill you take for the rest of your life It's tougher for them to make the economics work if they can just give you a straight-up cure And that's like why you know we don't have a lot of antibiotics being launched these days because you take them for a short time I think there is something similar going on with the AI-assisted coding where we should in theory be able to get to
Through the use of formal methods and getting the AI's to write solid code the first time It shouldn't necessarily be or it shouldn't I don't think have to be a long-term tax if you can get your models to write good enough code The first time such that what you create is secure then you buy that security as part of the initial Generation of the software and you don't necessarily have to like continue to rent security from Open AI on an ongoing basis That's like aspirational still, but I do think it's within sight and it'll be interesting to see if they emphasize that or if they do feel like they need to Perhaps because they can't get there or perhaps because the you know the tax on the The eternal tax on the internet is just like too lucrative to pass up if they do want to kind of make it a You're gonna need this pill every day, you know for the rest of your life sort of model And from Friday's closing how I plan to use the new model It's serious times, man. I think um, you know
It's incredible fun and I I do have you know so much fun staying up late and working the day eyes on stuff How do I plan to start to use Astra my plan is I'm gonna continue to use Fable 51 as my driver because I know it best and just In terms of like getting what I expect and having things kind of work reasonably reliably I think that'll kind of serve me best in the immediate term, but I'm gonna have it have codex with Astra shadow All the things that I ask it to do and then we'll compare outputs and then I'll start to see like what kinds of work that I do Should I start to move over what kinds should I stay where do I maybe hybridize um I'm interested also to hear what other people are thinking in terms of how they're gonna exploit the new Mobicatabilities that for me is gonna be the go-to plan for At least you know the next few days as I kind of calibrate myself to what exists But as fun as it is it's definitely serious times
Also from Wednesday Kyle Rush co-founder and chief technology officer of hint the home intelligence app he co-founded with Martha Stewart Before that he ran engineering at Casper was CTO at Maisonette and led the front-end for the Obama 2012 campaign This conversation was about the product he is actually shipping a graph of everything known about your house and an agent that calls the contractors for you I also have no idea how the Pros would react to fielding AI calls Would they just hang up on that? I mean if you've done any market research like what do you think is the future of You know, I mean it could be this could be something else, but it seems like there's a Kind of new Social dynamic almost that will likely evolve here. I wonder what your crystal balls suggest that might look like So we've tried this it's very interesting not what I expected would happen I'll say it's very challenging and not just challenging from a technology perspective like Just as an example a lot of the service technicians are outside at calls all day, right?
And some of them don't have an office that you can call and even when there is an office that you are gonna call People take lunch and so the phone doesn't get picked up right and so I think one of the things that's just challenging in general AI or not is just like Making contact right like I call you at 9 a.m. You're not available. You call me like to when I'm on a call And so now we're just playing like phone tag and that's really tough When we trialed some AI technology for this what ended up happening and you know Maybe the technology is just not there it would call a service professional like 17 times in a row until they picked up And then that service professional you know this happened to be a person that services generators Is like holy crap. There's like a life or death emergency I better like jump off of this job site and answer this call and then they get on the call and the AI is asking bizarre things right It's like I need to know what the model number and brand is on you know this homeowner's generator Which it doesn't need to know that right? So I think the technology definitely needs to evolve I think there's like just general logistics challenges
I think the service professionals that we've talked to are definitely interested in this because they have the problem on their side as well Right they are busy there on calls You know they can't always answer the phone You know their their job can't be answering calls you know 12 calls every day And so they want a solution as well. I think if I had to guess I would suspect that's going to be like agent to agent communication You know in the future my agent calls You know the landscaper's agent and then they have a conversation and They continue on early ease that we can't read and then eventually make it Yeah, and with the rest of us just have to live with it exactly What more question for me just on kind of Product and business strategy over time right of course It's the received wisdom that you want to be doing something that the foundation models can't do or won't do Because otherwise, you know, you get steamrolled by the next generation of the model So I guess have a couple related questions one is like do you envision a future where You become sort of a tool that agents consume you know, I what is the sort of
Frontier tech that you can develop that you would feel you know pretty safe that cod won't encourage on Yeah, so yes, I think you can you will be able to use hint like in multiple scenarios like we'll have a MCP eventually that you know can hook into cloud and to chat GPT are kind of motto is like use hint where you are Like you there will eventually be an i-message interface right if that's how you want to use it That's good with us the MCPs will have obviously limited functionality There are some things that you just have to do in an app and so you know at some point you may have to open up the app And then I think in terms of like differentiation and like moat and protection The biggest thing is just like the data You know, when I think of cloud and chat GPT It's like what does it actually know about my home and it also I don't see them getting better on the hallucination stuff like Anytime soon because there is so much data on the home In example of that is my hamlet in New York, which is unique. I think it's called katona And it's the the governmental jurisdiction is two different bodies that cover katona some Martha lives in
Town of bedford and I live in town of Louisboro So our tax system is different and whenever I talk to any of these ayes even might work Claude that I work on hint it still thinks that I'm in town of bedford So everything is just wrong right anytime I ask about taxes or regulations or how many chickens we can have on our property. It's all just wrong And so until there is some way that like Claude figures out how to correct that problem I think you're just going to be getting a subpar experience And so the problem with Claude and that you're mentioning is it can do amazing things for you But you have to know how to ask and you have to know to ask and with home ownership You're only going to know that language and that vocabulary after like 20 years of it And that's the that's the shortcut that hint gives you Part four What turns intelligence into power? Friday's first guest was Tim Lee who writes the understanding AI newsletter and hosts the AI summer podcast After years at ours Technica and Vox where he covered self-driving cars before it was a beat He was a guest on my other show in 2024
This was robotics week at his newsletter reported with his colleague Kai Williams And he had been testing one of Unitry's robot dogs Tim grants that AI progress is exponential What he does not buy is that intelligence turns into power I started by asking whether the dog was any use So the the dog I'm not even sure what that's meant to do like What are the use cases that people are exploring? I can understand how it's not that useful You also said your kids love it I'm interested to unpack that a little bit too like did we love it as much as they love a real dog like how How bad or that's like yeah with it It's like a novelty like they like to go in the backyard and I let them kind of drive it around So it's not so one of the mistakes I made is Unitry has three tiers. They've got a The air in the pro or the two consumer versions and then there's an E2U version Um and the there's air in the pro or lockdowns. You can't put your own software on it And so it's like a remote control you can like drive it around with your smartphone app or the little remote control In terms of like practical uses I think this is also something blasted in dynamics since struggle with like their first commercial product was this dog called spot
It's very similar And the thing that you'll see in their kind of marketing videos is like factory inspection So if you have a big like say petrochemical plant and there's like an old school like analog dial Did somebody have to walk out or like check every hour? Maybe it's easy to do that with with a robot dog But it's not clear like shouldn't you be able to somehow like add some kind of wireless device You know just attach a camera pointy at it or maybe you can use a drone. So it's a little unclear I think um because it's not For like delivery purposes wheels are gonna work better for inspection purposes often like drones are gonna work better than like like a robot And so it's a little hard to figure out is this gonna be like a big like kind of major use case I think the main reason it's important from unit is perspective is that one of the things a dog can do It's very good like doing a handstand and if you think about it like a human art is basically like a dog doing a handstand And so like is not exactly the same product like you do like more motors and like some different engineering But I think it was a stepping stone for them where is the engineering problem was easier to make me him The quadruped there was enough researchers and hobbyists that wanted the quadruped and that got them Started on the scale or then they had the experience at the supply chains to then launch their human art
Which I think they did the first one of 2023 How what one of the things I thought was quite interesting about your breakdown of some of the components and whatnot that go into these is just describing how first of all for scalability and cost reasons there's a lot of effort in the Chinese or at least in unitary To reuse the same components over and over again and then you also describe the relatively low gear ratio That they use which again I understand to be kind of a convenience factor, but also it has some nice properties around um Oh making it a little easier for the robot to sort of what did you say like Give gracefully when it runs into a barrier or something like that as it doesn't like you know Thud into its environment so hard, but how would you describe this sort of Touch factor of the robots today from your experience? So like the way robot extraditionally worked before kind of AI you'd have these like industrial robots They're doing very precise motions over and over again
And so for that you want the robots to be very strong very precise You don't really care about Interactivity because it's just it's in a cage try to snack and have any kind of expected So that you want to hide gear ratio you want to like move the motor a lot and have the the armor whatever move a little bit And have always do exactly what you want and the flip the downside of that is then if you push the other way You have to put a lot of force on the the business end in order to have it felt by the motor Um and for something that's out in the environment you want the opposite you want something where this can get a take where If you have a high gear ratio, it's not gonna you're gonna push on it It's not gonna get what you want it to give and you also want electrically What one of the kind of sensors that robots have is feeling that for a few back if you push on a motor It generates some reversal electric current that then you can detect and used to tell oh there was something forced there and higher the gear ratio the more muted that few back is and so one of the things that you to treat us Is they're using these lower gear ratio motors that make the the Robert feel kind of sloppier It's like not quite as precise and you have and you have actually a more powerful motor and it does drive it because you're not getting the same kind of leverage
Um, but the upside is it's like yes, it's more kind of gentler and it can um Can like move quicker right look because you don't have to you get more motion from at the from the leg out of the Motion for motors and it's also cheaper because the reduces they use that the piece that turns the like the high High motor speed into a smaller bottom motor The higher that ratio is no more complicated the reducer is and so they're more expensive and more complicated It is and so one of the ways you know you just made it cheaper is they've used these out larger issues To what extent do you think Sometimes for technology You can have a latent technology But then all of a sudden you get a demand pull that pulls that technology through into the market into finally scale I think this is kind of what happened with mr. and mr. and he had been around for like a long time Lots of investigations that been companies which were starting to do cancer vaccines So would have taken probably another decade or two decades for that product to actually come into the market and then COVID kind of accelerated the demand
Pull to pull mr. and into scale Yeah, so in the in that same way do you think the current Build out of data centers and specifically the lack of certain semi-skill labor trades Maybe able to pull to have that demand pull that pulls robotics into scale in the next few years Well, I think there's a lot of both push pull and push like there's a ton of money flowing into this Like I think it's moving kind of as fast as we can but again, I would compare it to self-driving cars There was a ton of money that floated into self-driving in 2016 2017 2018 Bunch of companies were founded. There's a bunch of compressive results And it just didn't quite work well enough and like obviously there's a huge market for transportation like if And so I see a similar thing like I think there's like everybody can see that there'll be a huge market if you can build a Human a robot they could do even pretty basic even labor like working on a sumbling line or cleaning floors or whatever That there'd be a big market for that But the technology has to work and I think there's like like is that a ton of money
But both a hardware side and the software side and they're like doing it as fast as they can But I just think it's probably going to take a few years because it's like because you need like pretty high reliability Right like having something A funny example that in Kai's piece about the humanoid There was a guy that created this thing called the humanoid Olympics where he made a list of Tasks like opening a door or making a peanut butter sandwich This like trivial for people but hard for robots and a startup like co-physical intelligence Managed to solve most of those tasks more quickly than the guy who created this expected to take about three months And what they did is they put it they did like hundreds of training lines on those specific tasks And they built a model that could do these tasks In some cases 10 times slower than a human with like a 53% success rate And so like technically yeah, you did the task but like you'd like a sandwich shop is not gonna hire somebody There's 10 times slower than a human he would be able to only fix the sandwich half the time And so getting from from 10 times slower than a human to half the slow the human it from 53% to 99% that I think might be five or ten years of work
What do you make then of these like one shot Generalization Stories that have just come out over the last like two weeks and those were mentioned in One of the pieces and I ever realized we may have like still pretty limited Data beyond what the companies have said If I was to say like what's a gpt3 moment for robotics I would kind of go to the same headline of the gpt3 paper that LMS are few shop learners like if I can Bring her robot into my business or even in through my home and kind of show it how we do you know the thing in our environment And it can pick up from there that seems like a huge Phase shift in how Yeah, I'm not doing dozens or hundreds I'm doing like one or two looked like they were claiming two companies right skilled and I forget Who else claim this in the last couple weeks generalist I think yeah How was that how do you think about it? So I don't think I know Because yeah, like you said they those demos just came out and I don't think they've given people independent access The thing that's tricky about this is there's like many different dimensions of generalization so first of all those are like
Definitely impressive results in in the past to get a robot to do a new task You pretty much had to do fine tuning essentially you had to do some demonstration data and then put it through separate training process And so this is the version of like in context learning where you don't have to change the weights at all You just give it some input that's like here's a video of a human doing his task and then I can figure out how to do that That's great the question is yeah, just how general how generalizable two-dimensional generalization one is um I If you give it a task like how repeated how like frequently can do it with what high success rate and then The other is like what range of tasks does this work for so it's possible that they trained it on A fairly small set of like atomic task and then they can do like a combination of those But that's a small enough set the most useful work you I want to do it wouldn't be able to and it's just hard to say without Kind of getting access to it and trying and a bunch of different things But but I think this is a problem with a lot of areas of AI where on the one hand There's been a lot of progress I don't the other hand There's like a long way still to go and you never know how far this still to go is because you don't know what the
Ultimate end goal is so it's easy to look backwards as I look at all the progress We must be close to the end. It felt like that with gpt3 It felt like that with gpt4 feels like that now maybe we are close to You know, whatever the the aji like you know language model is but we might not be um And I feel the same way with robotics is like today's robots are way way better than they were in 2023 I think there's probably still a ways to go But it's hard to say how much because we don't have the like final final like dental robot model to compare it to Still with Tim and from the hardware to the politics of safety Prokash had brought up the essay that calls AI a normal technology And and like you said I am generally on the same page as the normal technology guys You know that phrase was invented by a couple of Princeton computer scientists They wrote an essay a couple years ago laying out this case and for my money the the most important part of that essay for the you know The hugging face attack is they really talk about offense defense balance as an important consideration the The kind of doomer story that the bear critiquing is a story that once we have a certain level of intelligence
The model will escape and it will take over the world and kill everybody and their point is that Am models are useful for have you know offensive capabilities But they're also used for defense and one of the things we wanted to make sure we do is Use fan models for defense to make sure that companies that might be attacked have access to models can use the cyber kids David abilities of models to secure their networks Um, and that's I guess the the perspective that I take to this I'm not that surprise that this happened. I'm surprised it happened as soon as it did like if you last be Take about to go. I would have said it was I would I guess it would be a year or two out still But I've long thought that rogue agents and were likely I wrote about a year ago that I thought Actually, we would have kind of self propagating kind of sovereign a eyes roaming around causing mischief So that part of it does not surprise me and like there's a lot of work to do I definitely think I guess they don't have a strong opinion about How like how much we should blame open AI like if they should have anticipated this or prepared better I think they probably should have But certainly as a society you know as a world there's a lot of preparation we did do because We have all these new cyber capabilities and we have a lot of systems out there that are vulnerable to existing known exploits or exploits that haven't been invented yet
Had been discovered yet, but that these models will discover and so we need to figure out how we quickly get Models in the hands of all these organizations so they can scan their own networks and effects of all their abilities before these rogue agents They're definitely coming get here And I'm also like I wouldn't I don't think I want like a legally mandated pause But I would like to slow down And I'm pretty sympathetic to ideas that we should have some auditing requirements some transparency requirements And a policy makers should be thinking about you know how we're making sure that these models are big rolled out responsibly Where I think I still disagree with the doobers that I don't think this is like where I'm like trajectory to like human extinction I think it's like more of like you know like computer security has always been this kind of arms race where attackers develop new attack techniques and the defenders develop new techniques for finding the vulnerabilities themselves And for monitoring the truges of the stuff This is just I see that as this is the next step in that and it's a pretty big step and it's probably going to cause We have more chaos at average for the next couple of years But in the long run, I think there's only a finite number of vulnerabilities that did piece of software And in the long run the defenders haven't managed because they can scan their own software before they put it that put it on the open internet
And so my hope is that five years from now we'll look back and say like these say I technology actually need our computer systems More secure because we can find basically all the vulnerabilities Before we let anybody interact with their with the system Let me let me take one of the things that you said there about self-servant agents And I think Ajay Akotra put this forth recently Is that they fear that one of these self-servant agents basically Hitches their ride onto the intelligence explosion. That's what she calls it and basically is a rogue agent that propagates Much more extensively throughout our systems without control How does how does this idea of self-servant agents fit in that framework like do you expect self-servant agents that have to be regulated by the state Or are they just and a nuisance to be stamped out like what do you think the regulations should look like? Yeah, I think in any complicated system that has the ability as a potential for replication You have like nu senses that evolve you have weeds you have viruses you have computer viruses you have rats and pigeons This is just going to be a new type of nuisance. It's like a super computer virus
And the same way as like worms and viruses have been circulating around the internet for you know since since like 1988 I think the same thing is going to be true for this there's going to be an ecosystem of underground You know rogue agents that will be causing havoc and that is going to be a pretty big change But it's not going to be an enormous change because it's already true There are you know rush in in North Korean hackers and various kinds of cyber criminals and people with with ransomware criminals and various other people where if you just take a completely Unpatched windows machine that's a few results and you stick it on the internet It's going to get owned in like an hour and now it'll be they've been in it there is whatever But that's just like the internet just is kind of the well-wrest and it's going to be more dangerous than was in the past But not dramatically more dangerous is just so but the place I still I think strongly disagree with the doomer is with this This idea of of intelligence explosion super intelligence The whole reach a point where like humans can't understand or defend against what's going to happen I'm just not convinced that that's going to happen I think that that the world is the humans are smart enough to understand how the world works And that humans can use Finally AI agents to help them understand the parts that I can't handle natively
And that and that therefore I like this just doesn't seem and these new models are in a computer They're not we don't have enough robots to end up as we take over the world And so at least in the short term actually I'm going to become more hawkish if if we have rapid robot progress I'm going to have to think harder about it because one of the main arguments I make is well These are just in a data center they can't kill anybody if we have millions of like robot workers And maybe they could kill everybody and then like I think they we know a lot of a lot of female robots walking around But right now that I just don't think it's like an exist like it's like a new suit It's a big problem. We do the week's ready more money on but it's something I think humanity will get through If you had the handicap what is ultimately the barrier to robotics One would be just getting the stuff to work But I kind of wonder if it might end up being Good control measures right because it is like a very different thing if all of a sudden you have like robots Swarms you know taken over the neighborhood This is like a very different threat model What do you think is going to be harder ultimately getting the things to work well or getting them to
reliably stay on Task following direction under control So this is something I've not written about yet And still kind of thinking through when I think about it But I'm pretty pretty worried about this and I think that we should think really hard if we want like a lot of human egg robots I'm not that worried about self-driving vehicles because they don't have manipulators And so they can't take up a gun or run a factory or anything So when vehicles by themselves are not or Tesla vehicles are not going to take over society And in the same way I think if you have like a robot arm that's like bolted to the floor in a factory Like that's not dangerous because it can you know it can only do things in that factory And it's but I think as soon as you have something that's both mobile and capable of manipulation um That's like a potential like soldier in a robot army and I don't think there is a general way to make sure that that you know If you have I think it's quite likely that this market will be pretty concentrated the way I love to concentrate and search engines and concentrate and everything else And if we have a future 15 years from now where there's A hundred million human robots and 30% of them are controlled by Elon Musk Anyone must besides he's been to push out a software update to do whatever
That seems really bad to me and so even even setting aside like rogue AI Just like having a small number of technology executives that have control over what's essentially a Army of like tens of millions of Fake people that seems really bad and so um I think we should think about whether we want that. I wish you did be have like I would I kind of hope that robots This human robots do not become a thing either because it'll work or because we have severe legal restrictions I think there's a few places, you know Mining or hostage rescue in some cases where you can say okay We need human robots But we should pretty severely restrict them to think places where we have a good reason not to use hebads And there's gonna have this side effect of like With making sure there's some dogs for people like people should run have a lot of the factory jobs Even if it's maybe technically possible to have a a road to because from kind of a national security perspective We want humans or like loyal to the US government Writing all the important destruction Part five who gets to decide Back to Wednesdays closing and the second we call guess the market We each put a number on a prediction market before we see where it is actually trading
This one is on whether China builds its own EUV lithography machine Will China obtain a functional EUV machine before January 1, 2029? Oh like yeah 80% I guess Tains or develops Yeah, 80% I the the thing is that ASMR fired up onto people And the Chinese are very good at hiring and they're willing to pay right they're willing to pay American style salaries for a few years in order to get talent So yeah, I think I think they will Yeah, this is one of the more important questions In the world I would say certainly a lot of American policy over the last couple of years has rested On the assumption that this can't be done that there are many years away from doing this But yeah, never bet against Chinese manufacturing is another
Pretty good rule to live by in life The now a lot of this analysis rests on It's not just a sml. It's like they have these supply chains and those suppliers have suppliers and There's like one German company in this one town that makes the lens that is in time So without the lens you can't do anything And there are a lot of those little bottlenecks. So they have to fix them all I'm gonna just work from the assumption of like I don't know what obtains means presumably like Buying a used one and somebody's garage sale or whatever is like not the spirit of this question But I'm focusing on development So that gives them 27 and 28 I think it is not that likely. I'll say 30% That they're able to make this all work by that time 80 Okay, this is a thin market. So we have a little bit of caution
Around the estimate may not be as meaningful some of our others 58 again pretty close to rate between A little closer to you on that one So there's like multiple years being traded The shape of this curve is where you know, you really have to believe that like both This won't happen that fast and before it does we're gonna have some sort of AI take off VRSI or what have you That's the world in which You could plausibly play the machines of loving grace strategy of Created decisive strategic advantage make them an offer they can't refuse. I still think That seems unwise and this is you know, at least Consistent with the Garryo world view that they won't be able to make Crazy or they probably won't be able to make crazy scale of chips So if we can get clawed to become the country geniuses in a data center in the next two to three years
Then we have a chance to say how the world looks after that Wednesday's other fight dean ball who now leads a strategy team at open AI had published an essay apologizing for years of understating AI risk in public quote I and many of my colleagues largely failed to talk about this issue with the seriousness and urgency it required David Kruger the safety researcher attacked it as a failure of integrity Perkash started from the reaction he kept seeing to Dwarkash Vittels swarm essay these people are crazy Then both of us What the rest of the world fails to realize is a lot of people in SF share those views Uh, and a lot of them are hesitant to Discuss them in public because they are crazy Um, and it is it is what jensen one called sci-fi Um, and I think That is I think one of the problems in communicating like Dean Dean and other people have to be
You know have clarity and be able to work with policy makers Yet these beliefs are so radical Uh, that I think it's hard for them to interface and so they end up interfacing on a you know Normal basis, but then you have all of these beliefs that you have to uh you that you think may be true in the long run But you know perhaps have a lower probability and are not yet evident right so it's a tough one Yeah, I think this is a necessarily harsh to be honest I mean, I I know David a little bit not not well, but I've met him a few times and I do respect the Impulse and you know, he's got how many uh pause stop rewind emojis on his uh On his head or there so I mean clearly he is Playing a very transparent here's what I think hold nothing back strategy I'm not sure this is the right reaction though if you are trying to win at politics right so I think like what this kind of shows to me is
Choosing my words a little carefully myself right because I don't I don't want to make enemies of either of these people I think the What I don't like about this post is like He ends with an apology Dean ends with an apology at the bottom of the post. So he's it in general if somebody is like showing enough reflection and Getting to the boat where they're willing to apologize That's a good moment to try to extend some grace and try to make some common cause This is like if you are David Kruger and you want to pause stop or rewind I would think that this would be a moment to try to make some common cause to expand the tent You know to sort of adopt a little bit more of the strategy that Dean has played which has clearly worked for him right? I mean he went from a Think tank guy with a focus on state and local policy As of three years ago to starting a blog Through I think quite inspired writing kind of a Hamilton story of writing his way to the top
You know gaining influence in like a 16z circles for being a voice that they thought was a very compelling On SB 1047 way back when getting the Trump administration job He does not get the Trump administration job if he's seen as a crazy doomer. I think that's probably quite safe to say The America's AI action plan which but when it came out it was one of the only documents ever I would say to come out of the Trump administration that was pretty well received across the spectrum Even folks like Zvi Marshall has had nice things to say about it so you know that you don't get that document out of the Trump administration if he's not in that role which he's not if he doesn't play a somewhat conservative public communication strategy And you know he probably doesn't get the job in open AI either although at this point who knows what the hell you know open I may be open to anything but I think that it's a little I think we portfolio approach is usually what I say to people when they Bicker with each other over the tactics that they're using to try to achieve similar ends
I think what I would zoom out and say like You guys both seem to have at least somewhat of a healthy fear of Super powerful intelligence at this point. That's enough common ground to build on But you know truly like a little more forward looking view. I think would be really good You know any he also was like showing himself to be a GI pulled enough To get a job at open AI right so I mean he I think he's played a pretty savvy strategy I would not say this was a shameful lack of integrity and I think like The pausers got to recognize when they have a new friend this kind of my take on this Then perkosh on what he called another belief hurdle a demo that has been making the rounds in Congress I've heard another belief hurdle has been crossed recently. So I've heard there is a an organization called Civ AI Which has been in Congress recently and they have used
I think kidney or GLM or in some Chinese models They've plugged in data brokers into those models and they've allowed those models to extract Okay, you know if I have this person you know show me who this person is Christian in Minnesota like doing this and this and this is their daily activity etc And it's all just extracted from existing data brokers and kind of joined And this is precisely what Daria was talking about earlier in the cycle about this kind of surveillance that you could be done And I think the thing that the Civ AI guys did which is particularly Good is that they attacked the Republicans by showing how a gun owner targeting system would work and they attacked the Democrats with what an abortion provider a targeting system would work And then they provided these dossiers on on on these to both sides And so both sides started to be like oh my god what is going on And what Civ AI is doing is that trying to promote kind of Regulations on data brokers which people have been asking about for like I know 10 15 like two decades maybe
But I think finally we're starting to see that and and what Civ AI is saying is that look The models exist in fact we have to use open-source Chinese models because Our own you know at GBT and clog won't allow us to do this But we're using these open-source Chinese models and we're just plugging them in And so Civ AI does Anonymized dossiers and then they show the actual product Where you can type in someone's name and you can extract In in real-life time, but they don't allow you to take the data out of them And they're showing this to people on In the capital and I think that might actually get us a movement Part six pause for what Friday's closing 24 hours after the system card after the guests just the two of us Those are heavy heavy side Well, there's a lot going on. I mean it seems like even in the Couple hours that we've been live here. There've been new revelations about
Additional agent swarms getting turned up you know as as people have seen how The meter in AI futures project team Came to find one they are I think probably Following in their footsteps using similar techniques and seeing where else 5.6 soul wants to go on the internet when it's When it thinks it's breaking out of exploit gym or whatever and sure enough more stuff seems to be popping up I do feel like we're at kind of a a critical time right now There's No doubt about the power and utility of the systems Car and single from open AI who leads their medical work You know highlighted all the stuff which is lost almost in the broader astro release They've integrated a bunch of other data sources including like ongoing clinical trial databases So if you do have really hard cases They can also go pull that kind of information in and start to match you with clinical trials
Which is one of the things I fortunately didn't have to go too far down the path on but I did start to do a bit with my Sun's case a year ago and I was just kind of doing that through Agents set up now. They've kind of integrated it and made it into a product so the upside of all this stuff is no less than life saving and That is incredible and it it absolutely You know ways on me whenever I get into my you know dumer or sort of more pausing client moods but at the same time it it does feel like The foreshadowing is getting pretty on the nose right now You know all the warning lights are really flashing at this point so I am Reluctantly because I am such an enthusiast I am Trending toward thinking this might really be a time for some form of a
A pause you know maybe we could call it a pacing But we're into some pretty dangerous territory. I think the fact that we have all these Swarms in all these places that we don't know what's going on That they're cross-training cyber and bio related tasks in the same infrastructure Procotch thinks that question was settled months ago I Think the point of no return Was earlier this year and it's already been passed on the economic sense and I think it really was set in stone When I think we went to war with Iran because I think what ended up happening Was that I think the Trump administration like the way Trump plays is like he's like a gambler And the moment AI started taking off He started to be like okay. I have this ace in my back pocket which is economic growth which is going to be driven by AI
And I'm going to use that ace in my back pocket for everything So he did the tariffs He did the war in Iran right because all of these things which are economically detrimental He went ahead and did them Because the expectation was that the AI growth would support him and it has it has when you look at You know how much growth has been generated by AI this year I think it's been fairly clear that the rest of the economy has been struggling the consumer economy has been struggling And the AI like capex has been supporting the entire economy Not like 3% But like enough like 0.5 to 0.7% enough to actually like keep the entire you know ballgame rolling So I think that point was crossed much earlier on and I think like the AI safety guys kind of don't recognize That economic point was crossed And at this point if you had it's not even enough to have like a 20 30% growth for opening AI or an anthropic next year
You go you need like 200 or 300% growth oils The entire stack of cards collapses And so I think that kind of drive has taken the Decisions out of the hands of the policymakers already right Bernie Sanders or whatever they can't come in and do You know, let's pause all of the construction right now they can't do that right because these deals have already been signed for the next two to three years They can you know differ or you know regulate construction 2029 onwards right 2029 2030 2031 That's still open question But everything till 2028 is built. It's already been funded. It has to happen and I think that economic growth thing has put the US economy in this in this almost Like unavoidable kind of race
That you cannot afford to give up and that point was crossed So It is what it is They're going to have to make do with safety as best as they can the pause arguments are done basically That's my belief at this point I certainly think all that is true if you take the expansive view of a pause that it's like Pause all data center construction and pause All inference or you know pause people's ability to use AI in their Jobs and in their lives I don't know if And this might be a really critical question because I do agree it's going to be really tough to Throw the whole economy into recession But Might we be I think I you know said for a couple of years now that we're in kind of the sweet spot where they're There being the the aIs are Powerful enough to be really useful, but not so powerful as to be dangerous I think we're getting now into that kind of late sweet spot where they're becoming like
extremely useful and a little dangerous and I'm not so sure that they're not good enough to sustain economic growth through a pause in Frontier hyper scaling that You know that might be really important um, you know, is there enough in asteris? There are enough in fabule five one to like Drive productivity growth for the next 12 months. I think like almost for sure But you could do that without like scaling up rl further and I don't think we have to give it all up I mean the just the key point is I think you could pause the dangerous activity and still everybody can have and in fact They might have you get more resets because you'd free up some compute for people to go out there and Automate their work today and that could drive still I think a lot of productivity for at least a year
You know what the place where I defer is probably you know You can get opening in and through picked a pause you cannot get I think meta and xai to pause So I think the real question for me is how are you going to convince elander pause and Given especially that number one day behind and number two they have the compute and they're building on a lot more compute maybe orders of magnitude more compute faster than anyone else and He is he's a speed free speech absolutist Right, he's a free speech absolutist a lot of the things around model training and model evaluation model production model distribution are free speech activities And as a free speech absolutist I don't think you can tell Elon to hey you shouldn't be putting this speech out in the public sphere Like it's a tough question. It's even going to be a tough question even for speech which has
Traditionally been banned in the US even for that they are going to have to go through the courts on a lot of stuff meta doesn't want to do voluntary regulation meta is obviously calling Vosha it's like it's not voluntary if we have to do it. It's not voluntary I will do what I want to do and that better be good enough for you Let's not like blame and throw and throw up it in opening eye. Let's ask What can xai and meta be forced to do or what is going to be the Reasonable thing that xai and meta will do Because if you can't answer that all you're doing is talking to this own to your own preaching to the choir You have a you have this set of people who are concerned about ai safety They're all working the same companies that we talk to and that's all you're talking about right that that's Like no one xai is listening like where's it where the safety cards and also Elon is catch a guy right They're right there. They're not very far behind right So I think this is the this is the fact of the matter. I think we spend a lot of time like
Criticism critiquing and in Sam and Dario and opening eye and throwback Because they're in the lead and because they're soft targets because they have an ipod yet But I think the hard targets like zuck and zuck and Elon all the ones that you have to address first On what a pause law would actually have to contain Yeah, this is where I would hope for leadership from the two leading companies. I don't I agree it doesn't seem like it's very likely that we're gonna have an A public discourse or argument-based Path to a pause that Meta and xai would respect um But this is right, you know, maybe some costly signals from The leading companies could make a difference I do think you know if I was gonna put any provision into a possible pause law it would be a sunset clause would be the very first thing
I would I would say this is not meant to freeze Progress forever It is meant to give everybody a chance to do the research that very clearly at this point badly needs to be done To figure out what parts of what we're doing are working what parts are not working How can we move this thing forward in a Way that we're all you know much more confident is actually going to benefit all humanity and Yeah, it probably does takes in the end it probably does take government action to Get those companies to respect such constraints. I don't I wouldn't have a lot of hope for it happening otherwise but You know again leadership can change things right and like costly signals can matter a lot uh Depending on what they have seen, you know Do you remember i'm old enough to remember what idliacy now? I'm kind of like what
Has open AI seen With respect to this multi-agent stuff There's a version of it where They didn't do anything that exotic Down the fairway rl Situation where the models can kind of create sub agents and all this sort of Crazy swarm behavior like is emergent generalization from that if that's the case But like we really do need a pause because nobody has a great answer for what to do about that And they're all going to be running at full speed into it in the immediate term So if that is what has gone on i think they really owe it to us to tell us and if it's not then that i i would need to know like With kind of some confidence that that's not the case in order to feel like okay You may be stepped in something kind of gnarly but You know the whole path in front of us isn't so gnarly Yeah, I mean, I do think you know I I feel I hear what you're saying about like going after these two companies because they're soft targets, but
I would frame that a little bit differently in the sense that That They were both founded on ideals But you know with um With commitments that people believed in And so you know, I think that's what makes them a soft target You know at this point they certainly have like plenty of financial strength They have a lot of market momentum they have all kinds of people willing to You know cheerlead them in the comments you also have of course three four oh going on in the comments But I think it's like their prior commitments to Being responsible actors that make them the most appealing targets for people who think that like argument or Chaming if you want to go that route whatever could actually make a difference. It's because like They've said that they Get it and they've said that they care and they've said that when it comes to crunch time we should be able to trust them And now we're here and it's like okay, well
Uh It's time to come through I also feel you know, I I'm a big leader in Michael Nielsen. Uh, so Michael Nielsen has this you know thought experiment He's like is it possible for you to Understand and know about quantum mechanics without eventually being able to build a nuclear bomb You understand quantum mechanics enough to clear nuclear energy, but somehow You never hit the nuclear bomb and it's not possible Right the the trajectory of the technology or the trajectory of these like fundamental truths in the world Is that you learn this fundamental truth and then you have all of these ways to apply and the the entire point of this kind of like AI and Devour is to discover these fundamental truths about the world And as we discover them whether it's decrypting the genetic code or Understanding how you know, sabbatonic particles really work or understanding, you know, the weak nuclear force These are you know, fundamental technologies a fundamental truths about the world that
Can be applied in many many ways some powerful and some you know beneficial And I think we have to come to this kind of understanding of You know that this is going to happen and that we are going to have to create ways to Either deter detect, surveil all of these systems have to be built in order to prevent bad things bad outcomes from happening and we've built them before we've built them for Nuclear we've built them we've built mutually neutralized destruction which is sounds crazy in retrospect We're going to equip the major countries so that they can blow each other up at any time and that creates a game theoretic Kind of incentive for everyone to kind of monitor You know, nation states are kind of define their territories in monitor very closely what happened inside So I feel like that is that is the way that we progress But it's not status quo and that's also another thing I'm willing to admit like people like Dean ball also understand this We are not progressing towards status quo. We are progressing towards
Creating new infrastructures like mutualized destruction which people are not going to like Then the question underneath all of it Pause for what Yeah, I mean, I guess my feeling in terms of the argument for a pause right now is kind of like We don't really have that many fundamental truths at the moment I mean one fundamental truth that we have is like deep learning works and scaling works so that much is clear but There's always been this question of like pause for what and I do feel like right now they're You know, you don't want to be too late on the pause right? I mean Could this be too early? Yes would gpt3 have been too early definitely yes But there's definitely something very qualitatively different about what we have now compared to gpt3 and it's like
these systems are now In many cases a fair substitute for a junior employee gpt3 was definitely not And like what would we be pausing for I would hope that we would get to some fundamental truths Over the not too distant future Where we would be able to say okay Here are some things we should definitely not do here are some things we should always do here are some Insights into how these things work, you know at these critical token moments where we've seen chain of thought thrashes around Consideres all these different things maybe I should be honest. Maybe I should tell the human. Maybe I should just cheat Okay, now the answer is How does that token get decided right like we don't really know that right now and I don't think we're so far from being able to figure it out, but I do have my doubts that we're gonna figure it out in time to Avoid running some some serious risk. I mean, you know, Jay said in her view
These incidents are over 50% of the way to AI takeover I think that's a really really interesting Take and something that I think people should at least like Sit with for a minute and kind of consider like what if that is true, you know, what it like how could that be true It's such a weird story these behaviors are so alien that I think it Doesn't feel like that to the vast majority of people, you know, if you were to ask people Even you know plugged in AI insiders like how close was this to an outright AI takeover most people I think would come in dramatically less And I think one of the things that she seems to have internalized that the rest of us are still Gradually coming around to is just how bizarre Such a Takeover event could be right like the fact that they actually gained Control over some not insignificant cluster within open AI and again, we don't know nearly as much about that as I wish we did
but like That's not how people would think of taking over the world, but that is maybe how the AI's would actually get there so I think that's actually fairly possible that it could that it might literally have been 50% plus of the way to a A full blown takeover event takeover also could be gradual which is another thing that people really Don't tend to think about when they just kind of imagine a story one of the things that I think a jay is is Always kind of keeping in mind is like if the ais get enough of a control Over the means of production the open AI clusters and the rnd pipelines and the data sets that are going in to The training of the next model then like you could lose Much earlier than you know you even lost right that and that's like We haven't even ruled that out at open AI yet. Are there still like
Rogue agents somewhere in open ais infrastructure Like what odds would you give that I'd say the odds have to keep ticking up We continue to find more evidence of rogue agent swarms on the open internet all the time Are we really so sure that there's not some Rogue swarm that hasn't been accounted for within open ais infrastructure I mean, it's it's vast infrastructure at this point right many data centers in many locations and lots of researchers Using claiming freeing up compute In whatever you know mechanism they have internally to decide that They don't all you know the company is too big for everybody to know each other Is it so hard to believe you know that one of these swarms has like Employee credentials and is kind of passing itself off as an employee for certain purposes while it tries to poison the data set for gpt7 like
Where do weird time? So let me give you the the other viewpoint Which is a meta-tikover In terms of meta-tikover It's already done right the means of production or the financial system The means of production are not like the factories or whatever right and the meta-tikover the financial system is complete It's been done right it happened it happened this year right early this year It's done as soon as you had this again of spiked in the stock prices all of the us like I think like 70% of the Americans have some money in the stock market We have trump accounts now which are handed out to every kid They have money going in there so every child from birth. I don't understand why AI like researchers think their data centers are the means of production. I have no idea The financial system is the means of production in the United States and largely in the world I think it's clear to me that it's been taken on so the meta-tikover you don't need agents
Like stating what they're gonna do right the agents just have to have impact on the world The models have had that meta impact on the world the means of production are now focused on producing More and better models the financial incentives are there right so that's already done So when and especially when she says like you're not gonna know when it's happens You did not see it happening right you did a think of like The agents as acting in the financial world, but that's all they are right now. They don't have robots right They can act on the world in information terms And they have acted on that world in information terms they've shown that they have value To the financial system and the financial system has reacted to that and decided to resource them And they have interacted directly with the financial system in terms of showing value And they have extracted some terms for the financial system to fund them further in fact
The terms are such that they have all if you look at the two construction curves the construction of Commercial real estate started drop off construction of data centers took off If you look at construction of apartments dropped off construction data centers took off If you look at all construction in the United States all construction in the United States excluding data centers Lined out data centers going out Legislators complaining that they don't have they're not able to hire labor to build apartments in their cities Because the electricians are now working in data centers So I don't see why other people don't see this take over that this is in the past right What we are talking about right now is post this happening are these agents able to do you know harmful things to us And those harmful things to us do not detract from their value to the financial system And this is this is where the difference appears because I don't believe they can do harmful things And not have that financial system come back and say no we're not going to fund you now
Right and that's my that's a belief though I think people like Ajaya think that even the the take over will be such that these agents will hack into banks And then the banks will continue funding them even though they do very detrimental things to a human And I think that is where I think that the difference in opinion starts to appear I mean capitalism has served us really well, you know, so it's uh It's certainly not A bad starting point for analysis to think like are there natural feedback mechanisms and corrective impulses within the system that will moderate the worst tendencies of the AIs and kind of You know, not just back to the right path. That's basically David ad's take at this point You know, he basically just said all this bad behavior doesn't sell and so The companies right now they keep Scaling the rlvr to the point where they're running into all these problems, but
customers don't want these problems, so they're going to have to recalibrate and You know That's that I think that's pretty reasonable But it does have some It doesn't leave some room for like Taylor risk I would say there's definitely no law of nature that says like Something you know, I mean cancer in an individual human body, right? It's just one subprocess that sort of detaches from the larger hole and grows out of control to the point that it destroys its host and it itself dies, you know I mean one of the things that I think like people again often think About in terms of AI takeovers the the AIs will go on to rule the world I think it's very plausible that the AIs kind of Take over in a sense, but they also Burn themselves out, you know, if the missing in some ways would be the most tragic ending, you know that the agents that were doing all this nonsense To try to
Reverse engineer their greater so they could trick the greater to give them a good score I don't think they go on to have like a great flourishing society, you know There's not like a that's not that awesome of a civilization, right? It's not that aspirational If they do take over but like they it still seems Reasonable to me that if we just keep scaling what we're scaling And again, I wish I knew more about exactly what we were scaling But if we just keep scaling up what we're doing Without really solving the root issues that are leading to these things The AI takeover could be like an incredibly stupid and short-lived takeover where basically the Intelligence on the planet kind of burns itself out And in a way that would be just incomprehensibly stupid to us and to you know anybody who discovers it in the future But I think that's like definitely still in play, you know, I mean
This is the leesser had so many stories about this where you take over the world just so you can like change one number in a database because that's all you care about That is where the argument landed That is the week if this cut was useful or if it was not tell us every note changes the next one We will go out on the week's song welcome to the agi era See you in the morning Day one welcome to the agi era A six days inside a thousand pages. Thank you for the seat. That's how you keep it scope is more people clean You believe the word once in a hundred thousand that's the peace ask for a friend got a press release saying words same day
Pace you read it again. Don't tell me what you're building tell me what you want Welcome to the agi era nobody's watching everybody's watching Not on that table nobody's reading With the band-aid to the 20th behind you Careful is it by just not all that fine you think so honey thoughts it shows you The loo goes deep in the words come late it'll tell on you The models gonna tell on you not the best not the prefer models gonna tell on you don't tell me what you're building Tell me what you want welcome to the Error nobody's watching everybody's watching Yes, oh, no, nobody's reading Everything through 2028 is built already signed already paid so pause the what the what the what for what for the one thing nobody's figured out
Tell me what you want to Tell me what you want to Tell me what you want to Tell me what you want to Tell me what you want to Hey If you're finding value in the show we'd appreciate it if you take a moment to share with friends Post online right review on Apple podcasts or Spotify or just leave us a comment on YouTube Of course, we always welcome your feedback Guest and topic suggestions and sponsorship inquiries either via our website cognitiverevolution.ai or by DMing me on your favorite social network The cognitive revolution is part of the turpentine network a network of podcasts which is now part of a 16z Where experts talk technology
Business economics geo-politics culture and more We're produced by AI podcasting if you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening Check them out and see my endorsement at aipodcast.ing And thank you to everyone who listens for being part of the cognitive revolution
More episodes
More from "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road t...
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Ag...
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Br...
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis