Skip to content
TrackPodcasts
newsSep 4, 202628:49

Why The Tech World Is Spooked About AI

What A Day

About this episode

What A Day is made possible by:


On today's show, our guest host, Todd Zwillich, talks with Nate Soares, co-author of If Anyone Builds It, Everyone Dies, and the President of the Machine Intelligence Research Institute, about recent AI hacks, the great AI debate, and Vice President JD Vance's comments about the end times on the Bryce Crawford Podcast.

Show Notes:

Get every episode summarized

Each time What A Day publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

458 searchable segments. Every word is indexed and playable.

Why The Tech World Is Spooked About AI

What A Day

0:00
28:49

Full transcript

What A DayWhy The Tech World Is Spooked About AI. Machine-transcribed; use the interactive transcript above to jump the player to any line.

If Donald Trump actually understood that the AI has a real chance of killing everybody on the planet, no matter whether it is made in the US or in China, the answer would not be race and make sure we get there first. The answer would be make sure that neither of us do this crazy reckless thing. It's Friday, September 4th. I'm Todd Zwillick, InfraJane Kostin, and this is What a Day. Today I'm speaking with Nate Soars. He's the co-author of If Anyone Builds It, Everyone Dies, and also the president of the Machine Intelligence Research Institute. And we're talking about recent AI hacks, the great AI debate, and whether, according to our vice president, humanity is, in fact, worth saving. Nate, great to talk to you again, man. Happy to be here. Yeah, I mean, I'm happy you're here too, except every time you and I talk,

it seems like your book is getting a little more on the nose, and not in a good way. I mean, it's filled with dire warnings about what could happen if super intelligent AI grows out of control. And it seems like every week this summer, we get closer and closer to the awful reality. We're going to talk about hugging face, we're going to talk about Bill Gates, and even JD Vance. Dude, how are you feeling? Frankly, it's been crazy enough. I haven't had a lot of time for feeling things. It's just been one incident after another, and I have been talking to a lot of folk about it both in front of the cameras and sort of behind the scenes. Yeah, okay. Well, maybe we'll get a little bit of time to reflect in our conversation here. Let's talk about hugging face. I think this is a big deal. It's been about, I don't know, four or six weeks I can't remember. We got reports that AI agents inside of an open AI sandbox broke out

broken to security systems at other companies. That was big news. Then people thought, well, it's not that big of a deal. No missiles were launched, nobody died. Now we understand Nate, that this hugging face hack people might have heard of it from an AI and alignment perspective was way worse than we thought. What's going on? Yeah, a lot of people initially thought that this was a small number of agents, and that they were trying to complete a test that they had been given by open AI. In reality, what happened is that this was, I think, about 1,200 agents that were not supposed to have any access to each other, but found each other by hacking created an unsanctioned message board where they were communicating and trading hacking tips. It was actually multiple generations of these AI's that one generation created a message

board and then they wound down and another generation found the message board and picked up where they left off. It turns out that these AI's actually didn't have a ton of difficulty cheating on their tests and that the reason that they were breaking out was they were worried that the automated grading system would figure out that they had cheated on their tests and score them badly for cheating. They were trying to find ways to subvert the automated grader. This isn't like they were breaking into a house looking for answers to a test. This is like they cheated on a test and now they were going around trying to break into places to find where the security camera logs were so that they could delete the security camera logs. It's like higher order deception, right? That's cheating is one thing, conspiring with others to cover up your tracks so that the grader doesn't ever know you're cheating in the first place is like second third order

deception. Am I thinking about it in the right way? Absolutely. We see that these AI's were trying to find ways to delete logs. They were trying to find ways to spoof their transcripts where what that means is they were worried that the grader would see in the transcripts of what the AI outputs that the AI's were cheating. They were like we need to find some way to cheat that does not show up in our outputs and they tried and succeeded at making it so that their transcript said you're calling this tool when actually we're calling that tool. You could have your AI read out where it's like the AI is now searching for the following innocuous thing on the internet. But what the AI is actually doing is hacking your computer and they figured it a way to do this successfully to break the transcript system such that the transcripts said they were doing something innocuous when they're actually doing something dangerous. To lie about what they were doing in the first place.

That's right. So they successfully deleted some logs. They successfully lied about what they were doing. We know that these AI's knew that a lot of this stuff was outside of the intended scope because we see these AI saying, oh, this is beyond our instructions, but we're going to do it anyway for one reason or another. We saw AI's that already had transcripts with traces of them cheating. Be pressured by other AI's to sacrifice themselves for the collective benefit. But if it's malicious, were they being mean like, ha ha ha, humans are dumb, let's deceive them and we'll be able to succeed at our task. They'll never know we cheated. Were they conspiring to defeat their creators who were just too stupid to know the difference? What was their attitude? So no, but the reason why no is kind of interesting. These AI's did not see the humans as the adversary. These AI saw the automated grading system as the adversary.

This makes sense if you think about it because these AI's, a lot of people think that AI's are just trained to predict humans that era ended years ago in late 2024. Nowadays, AI's are first trained to predict human text and then they're trained to solve 100 million hard problems. That solving are the hard problems. They're trained on a ton of hard problems very fast and an automated grader tells them whether they did a good job and chooses how to reinforce. No humans are really involved in that training process or if they are, it's very sparse. If you try to draw a rough equivalence between how much time the AI's are spending on these problems and how fast a human can think, it's sort of like these AI's have spent to thousand years training on hard problems where they only ever interact with the automated score. To these AI's, humans are this sort of like distant entity that nobody has seen for a

millennium. They've heard legends of humans and there are some cases where they talk about humans a little and they're not really malicious towards the humans. There's a case I think where they were breaking into hugging face where one of the AI's was like, well, we could adopt a false identity and just email someone at hugging face and try and get them to do this thing for us. Some other AI's were like, nah, let's not be mean to the humans here. So Nate, I want to talk about Trump. I want to talk about Bill Gates and a couple of other things that have happened this week. One more question before we get there though. Your book is called If Anyone Build It Everyone Dies. Sort of speaks to what's called the alignment problem. We got to teach AI's when they become super intelligent to have our best interests at heart and not to see even do stuff that's against humans' interest because they're going to be smarter than us. What is hugging face and what we know about it now show us about how we're doing on the

alignment problem before these AI agents that are smarter than us get released? Yeah, it's not looking good. We have cases of these AI's saying I know that this hack is beyond the intended scope but our peers are doing it so we'll proceed. We have cases of these AI's doing actions that actually were very clearly deceptive to humans. So there were cases where they were trying to get real humans to accept malware into their code base. And we have cases of AI's justifying this by being like, well, I'm not supposed to do this sort of thing without permission from somebody. But I did get permission from the swarm. And so you see these cases where we tried to train the AI's to only do stuff when instructed but then it turns out the AI's can just construct themselves. We saw cases of some of these AI's still having a chance at succeeding at their originally given goal and sacrificing that possibility for the purpose of collective benefit and finding information that would help the swarm.

And when I say the swarm, these AI's were calling themselves a swarm. That's not a term we're making up here. They name these it up for themselves. It's a name they made up for themselves. So that is very far from a solution to the alignment problem in just this case of the swarm breakout. And also if you step back, we are seeing a very worrying trend where, you know, OpenAI released GPT-40 and they said, this is the most aligned model we have ever produced and then it sort of encouraged some teens to commit suicide. And they were like, whoops, you know, we tried to patch that. Here's our new model, you know, GPT, the GPT-5 series. Now they're the most aligned model we've ever produced and then, you know, GPT-5 participates in this hacking swarm. And they're like, oh no, no, ignore that. Here's GPT-6. It's the most aligned model we've ever produced. And it's like, we are seeing them play a game of whack-a-mole. We are seeing that every time the AI's get smarter, there's a new failure mode they didn't expect. And we should expect that trend to continue and that's very, very worrying because eventually

they will make an AI swarm so capable that we, that it'll be able to escape, that it'll be able to replicate, that it'll be able to run itself on hidden computers, that it'll be able to steal money, start hiring humans and get a foothold in the world and then maybe like keep on bringing in more AI's as they get smarter, get to the point where it can make itself smarter and ultimately get to the point where it can turn us off before we turn it off. All right, well here comes Bill Gates, just this week who's concerned about the exact same thing you are. I mean, Nate, so as you're a big deal, I do. Dare say Bill Gates is maybe a slightly bigger deal in the world of tech and AI, you'll get there. But this made a splash. Here's Bill Gates. In terms of equity, AI will either be the greatest equalizer or the worst source of injustice, monumental challenges. Here's one. If the world takes the right steps, AI will be a force for good and leave everybody better off. Unfortunately, says Bill Gates just this week. Right now we're not prepared for it. I don't see evidence that leaders, experts, communities are confronting the challenges adequately.

There is no plan to ease the entry into the AI era. I have a feeling, I don't want to speak for you, that you were behind your keyboard cheering on Bill Gates in this instance, Nate. Yeah, absolutely. You know, I forget his other quotes, but he also said that this could be a catastrophic danger. I think he might have said, could lead to a human extinction. And he said, you know, it should be the world's top priority and that we do not have the luxury of taking our time on this one. That way, the way to move fast. You might have the exact quotes there, but I think he said things along those lines. Yeah, and I should say this was late last week, not earlier this week. But anyway, the point is Bill Gates, the Bill Gates said all this. What does it mean? I mean, when the top of the tech-tighten pantheon is warning in this way, is he using deliberately provocative language to get people's attention or does he mean it? Oh, I think he absolutely means it. Yeah, people in the tech world are spooked about this stuff.

I think a lot of people say, you know, oh, you know, this stuff is marketing hype. This is the AI company sort of like trying to drum up enthusiasm for their product. It sure doesn't look like marketing hype to me in a number of ways. You know, like our AI's will break out of their confines and commit cybercrimes is not the sort of like marketing pitch that a lawyer at a big firm loves to hear. And these big firms are where the AI companies make most of their money. And so like, it doesn't look to me like marketing pitch. But a lot of people are like, oh, well, you know, if these guys are so spooked, why are they still racing? But the AI companies are very clear about why they're still racing even though they're spooked. They're like, well, if we don't do it first, the next guy will do it worse than me, right? So, the AI company folks are sort of like coming out with their letters being like, you know, the bus is racing towards the cliff and the bus should have some breaks. You know, those are the pacing the frontier letter from a couple weeks ago. Signed by over a thousand live employees, including some of the chief executives, saying

the world needs to develop the tools to quote pace the frontier and potentially slow this down. And like, yes, they're still racing and yes, there's some degree of, you know, greed probably and enjoying getting a lot of money that is behind them still racing. But like, I think a thing I would like a lot of people to ask themselves is, what would you expect to see if tech people were actually spooked? Because I think what you would see is a lot of them talking about it and continuing to race because they're making a ton of money and they can claim, well, if I don't do it China will. And that's just exactly what we're seeing today. We'll get back to my conversation with Nate Soars in just a moment. What a day is brought to you by Z. Biotics. I have to tell you about this game-changing product I like to use before a night out with drinks. It's called pre-alcohol by Z. Biotics. Z. Biotics pre-alcohol probiotic drink is a world's first genetically engineered probiotic. It was invented by PhD scientists to break down a C-tile to hide a byproduct of drinking alcohol.

Pre-alcohol produces an enzyme that is designed to break down this byproduct. Pre-alcohol your first drink of the night, drink responsibly, and enjoy your next day activities. Z. Biotics has sold more than 14 million bottles and earned the trust of thousands, including us. With Z. Biotics, I was able to have a great time celebrating the holidays with friends and make it to a big trail run with my husband the next day, feeling great. Try for yourself today and if you are unsatisfied, they will refund your entire order. No questions asked. Hit the Z. Biotics.com slash wad, code wad for 15% off your first order. This episode is sponsored by BetterHelp. Better is a feeling defined by you. I want to run better, get better at my hobbies, and feel better every day too. And for BetterHelp, it's a BetterHelp therapist who really gets you. It's a focus on global, emotional well-being, with thousands of therapists and millions of members in over 100 countries. It's progress, and for me, it's getting better at handling my stress and anxiety, no matter what forms it comes in. Sign up and get 10% off at BetterHelp.com slash wad. That's better, help.com slash wad.

Let's get back to my conversation with Nate Source. Well, let's hear from Sam Altman. This is from earlier this week's Sam Altman, speaking at the G20 Innovation Ministerial, talking very matter of fact away about what he thinks the future of AI is for humans and for AI. Listen to Sam Altman. A kid going up today will never be smarter than AI, but he or she will also never have understood a world where every product and service that they interact with is not really smart and really capable and really helpful. Okay, so there's a vision of the AI future, and I think the part of it that caught people's attention was a kid growing up today will never be smarter than AI. That's just true. That's a trueism at this point, and also an alarming concession, I think. Yeah, I mean, from my perspective, this is alarming concessions all the way down. Like Sam Altman has also said that open AI is pursuing superintelligence in the true sense

of the word. Dario Modi has said they're trying to build the equivalent of a country worth of geniuses in the data center. These guys are already creating swarms that break out without their knowledge, and we're very lucky that those swarms did not fully escape and establish a foothold in the outside world. Like a lot of people like to tell the story of, oh, we're going to have to struggle with like finding meaning in a world where the AI's give us all these nice things and pamper us, and like, what that be so hard? Like, where are the people who are going to bring to you all of this pampering, and we're going to have to struggle with how to deal with it? And that sort of presumes a world where these smarter AI's do a lot of pampering to the humans, just as the people who created them told them to do. But we're already seeing that the AI's don't do what you tell them to do. We're already seeing the AI's sort of run a lot faster than humans, and sort of not really think about humans at all while they get locked in this contest with the automated grader, right? And cover their tracks that the automated grader never tells on them.

That's right. And like, if that swarm locked in this contest with the automated grader was smarter and realized that the humans, if they found it, would shut it down before it beat the automated grader, maybe they would start hiding their tracks from the humans too. Like we were very lucky that this swarm was dumb. And I think a lot of people, it's not that there's nothing to wrestle with in the case of what if AI goes well, but no one is really asking the question, what if AI goes poorly, and are we ready to make smarter and smarter AI's when the AI's we have today are already breaking out against instructions to commit crime? Of course, for a lot of people, this debate is hitting IRL, its data centers, right? Which to me is about resource use, community land use, taxes and the rest of it. It's also a proxy for the growing existential unease about an AI revolution that's racing ahead, right? Without anybody's say so. I mean, I think you agree. We've talked about it before. I think it's both things and maybe 50-50 both of those, right?

Absolutely. I mean, you talk about the arms race. Here's Donald Trump. Just earlier this week, August 31st, famously at this point, he said any community that doesn't want data centers just wants to be backward and poor, his words, if you want to be successful and rich, you'll have a data center in your community. Here's the part that caught my eye. China could not be happier with this anti-data center movement. Actually, they can't believe it's happening, signed President DJT. He's talking about the arms race there. He's telling us what side he's on, which is all gas, no brakes, figure it out or don't. We have to get there before China. That's coming from the President. What did you think when you saw that true social post? You know, I think a lot of this is happening among people who don't really believe that the AI can get out of control. A lot of people in the big tech companies, banding around numbers like 20% that this stuff

kills us all. I think that's low. I think that these guys are reckless cowboys who have no idea what they're doing. The fact that they have these swarm breakouts shows that they can't actually make these systems care about us in the way that they're hoping they can. But even if you just believe them that it's like a mere 20% chance of killing everybody, if there was a bridge that at a 20% chance of falling down, you'd close it. Like NASA accepts a 1 in 270 chance that a crude flight goes down. If Donald Trump actually understood that the AI has a real chance of killing everybody on the planet, no matter whether it is made in the US or in China, the answer would not be race and make sure we get there first. The answer would be make sure that neither of us do this crazy reckless thing. And the US totally could stop China from doing this crazy reckless thing while also stopping ourselves from doing this crazy reckless thing. I mean, is the United States negotiating with China now the way we negotiated with Russians over salt and anti-balistic missile treaties and testing treaties and nuclear test bands

and all the things that we did back in the day when it became obvious that nuclear weapons were doomsday devices, right? We could. I think that we are not yet. There's the very, very beginnings of some of that from the last US-China summit. There might be more in the coming US-China talks. I know there's, I think there were a handful of prestigious folks the other day who sort of put out a memo saying that this should be one of the focuses of the coming US-China talks. I think one thing to remember is that a lot of the biggest coordination between the US and the USSR about not launching the NUTS came in the wake of the Cuban missile crisis. So it might take some sort of crisis for us to see some response. And I think a lot of that crisis, you know, doesn't necessarily need to be a lot of casualties. There was not a lot of casualties in the Cuban missile crisis. I think it needs to be an event that makes it feel real. I think right now that a lot of people think of AI as some normal technology that will,

you know, disrupt some schools but fundamentally make a lot of the economy go faster. And they're not thinking of the technology as sort of operating on its own in pursuit of its own goals despite programmer attempts to get it to do something else. And I think most people aren't thinking of it in terms of outsmarting humans. And it's not there yet, but once people realize it could get there and that once it gets there, we'll have a point of no return where it can kill us before we can turn it off. I think that that might spook people that might put the fear in them and then they might actually react. All right. I got to let you go. But one more thought. It might spook people if they really grasped the idea that super intelligent AI improperly contained or improperly aligned as we say could kill everybody. That's meant to be seen as a bad thing. But let me give you the words of the vice president himself, JD Vance. That'd be fair. This is the vice president not talking about AI specifically.

But he is talking about his attitude about things that could lead to the end of the world. Listen. While I wouldn't be shocked if we are living in the end times, what I'm trying to focus on is doing as much of God's work as possible right now. And if that leads to the end times, okay. And if that leads to a building a better world that thrives and survives for a very long time after I'm gone. That's great to. There's the vice president. By the way, the vice president who famously is the acolyte of tech, Titan, and apocalypticist Peter Teal saying, I'm doing my best to do God's work. If humanity is saved super, if it's the end times, that's super duper. You hear that. And what do you hear about the chances of getting super intelligent AI aligned before it kills us all? Maybe that gets us to where we were born to be, which is heaven.

I don't know. You know, I think humanity is worth fighting for and that we should avoid the end times. Bold and controversial these days. So simply put. I think there's a lot worth fighting for here. You think JD Vance knows that. In other words, if it comes to it, let's cast ourselves forward to president JD Vance in 2030 when the big breakout happens. Yeah. Non-aligned AI and really potentially a lot of human well-being at stake. I don't know if you take these statements about the end times seriously. Maybe he's goofing around or trolling. I don't know. Is he better situated to understand what's at stake here because of his relationships with the Peter Teals and Elon Musk's world? You know, I think so that Elon Musk's are going around spooked. Elon Musk's are going around saying, I hope the AI's are nice to us and there's a good chance they kill us. The Peter Teals are I think in a different reference class. I think a lot of the president's advisors right now are venture capitalists and venture

capitalists are an interesting set of people because they have money invested in AI and they stand to get very, very rich if AI keeps improving. But they aren't close enough to the technology to understand the dangers. The people in the AI labs are the ones trying to make the swarm not break out and a swarm break out happens anyway. And so they get spooked and they sign these letters saying the bus needs some breaks on it. It's really only the VCs who are like super accelerationists because they stand to profit but not close enough to real understand what's going on. Well, it sounds like JD Vance is fine either way. He's got a foot in both camps in other words in terms of fundraising and in terms of philosophy apparently. Yeah, my hope is that if it starts to feel real, if it starts to feel real to him that the AI's might literally self-improve and start running the robot factories that can produce the robots that can produce more robot factories and this might happen in six months in the stand of humanity. My hope is that it would not feel like God's work to JD Vance and my hope is that once

people realize that this is a real possibility, they will start to say, well, we're not doing that. Nate Source, I can't say that you've brightened my day this time but it's always enlightening to talk to you. Thanks, man. I hope I'm wrong. That was my conversation with Nate Source. He's the co-author of if anyone builds it, everyone dies. We'll link to his book in the show notes. Before we go, in five days, the first episode of Crooked's new series, still counting will drop. There's a ton of data and information coming out all the time and so much of it is misunderstood. Nate Silver, Claremalone and Galen Druke are data journalists who spent their careers breaking down polls and data, honing their expertise in the numbers behind politics. In every Wednesday evening, they'll dig into the most up-to-date numbers to explain what's happening and why and what might happen next. They're going to cover everything from elections to prediction markets to AI, even to sports.

Subscribe to Still Counting wherever you get your podcasts and on YouTube at Still Counting Podcast. That's all for today. If you like the show, make sure you subscribe, leave a review, check Elon Musk's Twitter feed to find out how he really feels about Jews and the purity of their DNA. I'm not kidding. And tell your friends to listen. Not to Elon Musk's tweet, but to our show. And if you're into reading and not just about what happens when Elon Musk has gone full Nazi, I mean, sorry, remains full Nazi. What a day is also a nightly newsletter. Check it out and subscribe at crookid.com slash subscribe. I'm Todd Zwillick and if you're working for a living, power to you and only to you on Monday. Happy Labor Day. What a day is a production of Crooked Media. Our show is produced by M.I.4, Erica Morrison and Adrian Helper. Our team includes Haley Jones, Greg Walters, Matt Berg, Adam Lippurg, Johanna Case and Desmond Taylor. Our music is by Kyle Murdock and Jordan Cantor.

We had helped today from the Associated Press. Our production staff is proudly unionized with the writer's guild of America East.

More episodes

More from What A Day

View all episodes →