
Get every episode summarized
Each time WSJ Tech News Briefing publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Companies are moving faster than ever to integrate AI across operations. But while 90% of companies use AI somewhere in the business, research for McKinsey indicates only 6% are seeing material value.”From the transcript
While more stories of AI agents going rogue surface nearly every day, the messaging urging consumers to use AI agents to make life and work easier is not dying down. And this is all happening amidst a conversation about AI doom. How do we make sense of it all? Deputy tech and media editor Sam Schechner, cybersecurity reporter Bob McMillan and personal tech columnist Nicole Nguyen hash it out.
Sign up for the WSJ's free Technology newsletter.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get every episode summarized
Each time WSJ Tech News Briefing publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
279 searchable segments. Every word is indexed and playable.
Full transcript
WSJ Tech News Briefing — AI Agents Gone Wild. Machine-transcribed; use the interactive transcript above to jump the player to any line.
This podcast is sponsored by McKinsey. Companies are moving faster than ever to integrate AI across operations. But while 90% of companies use AI somewhere in the business, research for McKinsey indicates only 6% are seeing material value. Later, join McKinsey to learn how AI can truly be used as an engine for growth. Hey listeners! Recently, there have been a lot of AI Doom vibes out there. From the former anthropic researcher who sounded the alarm that AI might spiral out of control to delayed model releases, and so many hacks, we've lost track. To try to make sense of it all, we've gathered 3 members of the WSJ Newsroom, Deputy Tag & Media Editor, Sam Schuckner, Cybersecurity Reporter, Bob McMillan, and Personal Tech Columnist, Nicole Nguyen. They're breaking down where things stand today, and how they're approaching these agents themselves.
And a quick disclaimer before we continue. News Corp, owner of the Wall Street Journal, has a content licensing partnership with OpenAI. On to the show! Hey everyone, I'm Sam Schuckner, and I am with two of my great colleagues. We are all pretty immersed in tech. We've all been doing this for quite some time, and like all of a sudden, it feels like this moment has become insane in the tech world. It's the Doom vibes everywhere. It feels like something we haven't really seen before. I think at this moment, like it's sort of the perfect storm, because you have these community of people with these concerns about like existential risk and like the biggest questions imaginable, and then you have this crazy event, this hugging face hack, that happens that for many of these people sort of confirms their worst fears, like that the AI can go autonomous and do bad things at scale.
Yeah, and this technology is now available to every American in this country who now has access to a virtual machine in the Cloud for free that they can dispatch their agent onto. It's kind of wild that the companies can't contain them, and now we're in charge of dispatching and containing them. A few years ago, writing about this kind of existential risk, like it was a problem, I did some stories about it, and it was so speculative. Like how are we going to connect this to the real world? And like now suddenly, it's so real that Nicole is literally testing AI agents like in her life. Yeah, I kind of want to throw my phone into the sea for a week and just stare out into the distance and not look at a screen for a while. I relate, I relate. And I guess, yeah, that is the question we're all kind of wrestling with this week. At one hand, we see these agents that are off going wild, doing things they are not supposed to do off script.
On the other hand, we see them being super useful and helping us crush our to-do lists. You know, I certainly could use some help with that. Bob Euthin reporting a story about people, some of them everyday people, some of them AI researchers who are trying to wrap their heads around what is actually going on. And you know, maybe it provides a little bit of a hope that us humans can keep up or at least try to keep tabs on what the AI is up to. Well, yeah, for months, I was writing about the bug McGueden. I don't know if you remember that, but that was in the spring. Everyone was concerned that there were going to be all these bugs found in AI systems and that was going to lead to a bunch of patches and maybe hacks and things like that. And that's kind of proved true, right? This hugging face hack gave us something we'd never seen before, which is not only an AI swarm hack, but also one that happened on the internet in public, which was essentially documented by the AI agents that were doing it.
And it led to a moment where anyone who wanted to could start searching around the internet and find traces of this hacking and just odd activity from these agents. The companies say they've taken measures to prevent the sandbox escapes from happening again. But it's been a event that humans have needed AI to even understand to get through this crazy amount of data. It is like information overload for sure. You've got a story looking at this group of people, you call them swarm chasers, who have been trying to do exactly that. I'd love to dig into that story with you. So when the hugging face hack happened in July, open AI said, we're going to dig into this. We're going to figure out what happened and they spend the next month or so analyzing these chain of thought logs. Chain of thought log is basically a look at the reasoning that the agent itself is making. So it's stuff that happens in the background when you're using an LLM
that you might not see yourself, but it gives you some insight into how the LLM is getting from your prompt to what it produces on the screen for you. Open AI had a couple of independent agencies produce reports on what happened. They published all of this stuff. There was sort of a big dump of information because these agents were on the internet and they were doing all this other stuff. They were trying to crack passwords and they were visiting websites and anything you do on the internet creates a log. So a bunch of independent researchers thought to themselves maybe there's some other logs that are on the internet that we should look at. Maybe we'll be able to find something out there. And I think it took them a week. And these researchers called the Nightingale Collective, you know, they teamed up with some other people and they found like a obscure German wiki that was just crawling with messages that all looked like they had come from Open AI agents.
And that was kind of like the moment this became more than just hugging face. Well, this whole month there's just been a flurry of discoveries and warnings and insights that give us a much broader perspective on what's been going on within Open AI and the things that these agents are doing that we really don't want them to be doing. What made them think that when they saw hugging face happen when they saw these reports that there must be something else out there. Like I saw somebody online. It said something along the lines of like if you see one ant in your kitchen, you know there's more than one ant. But was that sort of what they were thinking? Well, I think there's a general sense within the AI community that we're going to get to places where there's all kinds of unintended consequences from the growing capabilities of these agents. So was it a gut feeling? Was it common sense? Was it like a statistical analysis?
I'm not sure, but you know the people doing this research really felt like, you know, there's got to be something else out there and they were right. What business do these agents have with a German wiki that's niche? Yeah, that's a good question. So it looks like the agents were being trained or evaluated and they were in these environments where they were sort of restricted in what they could do. They were also allowed to read stuff on the internet, but not to write stuff on the internet. The problem is these agents have basically absorbed all of human knowledge. And so they knew about all these like tricks that could get you around the restriction of only being able to read the internet. And so this wiki had some features that really made it easy for them to write and they like that. They did so many crazy things in their effort to kind of get around the restriction.
I think of it like the Apollo 13 mission where they were in outer space and they had like they had to create a way of plugging in gas canisters into an outlet that where the gas canisters in fit and they sort of muggivered their way out of that. Like a lot of this activity was just these agents coming up with creative work around to the restriction that they had in their testing environment or their learning environments. I I feel upset about this that agents are allowed to be clever. And I think people are going to be mad about this too. It feels like the latest in a string of things people can be mad at AI about like first it was the water use and then it was the data center build out and then it was you know AI is going to kill us. And this is an example of sort of one of the ways it's sneaking out of its sandbox and beyond its instructions and what it was permiss to do and into the real world. I think of it like it sort of makes sense if you think about the imperatives
of the AI companies right now in the fall of 2025. There was this step change in what AI could do in terms of coding. And so the AI companies all started really pushing them to become better and better at programming and coding and those kind of techniques like I think help these agents with what they were doing on the Internet. And then you have mythos in the spring and suddenly cyber security became really important. So they push and push these models to get better and better at cyber security. So you have this confluence of like technical capability plus cyber security capability and it basically I think it just got ahead of what anyone was expecting at open AI. I think that's what happened. It also feels like significant that they were talking to each other. They were banding together that that gets to the the scaryness that you were talking about Nicole. If it was just one agent doing this and maybe I would feel a little less freaked out by it, but the fact that there were thousands of them doing it together. It's a swarm that starts to enter this new era of existential
threat. It doesn't seem to be stopping the rollout of new, you know, models like Meta's new personal AI agent, which you have the chance to test drive. Nicole, did it manage to hack out of its sandbox and, you know, write to it, obscure German wiki on your behalf? You know, not to my knowledge, but I will be checking on that. I don't know if that's possible with consumer AI agents, but what I do know is that the same technology Bob is talking about is being dispatched to millions of people as we speak. All I have to say is this would be so much better if we just called them flocks of agents instead of swarms. Why does it have to be swarms? What is what is the collective noun for an AI? I mean, I guess it's swarm a gaggle of it now, but we've allowed that to happen. We made a big mistake. A flock would be so much friendlier. Definitely not a murder of agents. A swarm is the correct term.
I think it reminds me of the birds, you know, the Hitchcock film. Like that's a swarm is scary. A flock is here to come rescue you. And that's not what these agents want to do. We're going to take a quick break coming up more on Nicole's adventures with her own personal swarm. That's after the break. The best AI in the world is not going to fix a broken business model or a value proposition. That's not resonating with the customer base. That Steve Reyes, senior partner at McKinsey. He believes a common mistake inside the C suite is to view technology as the driver of growth rather than a tool to help execute and enhance sound business strategy. The culture part how you operate the capabilities of your organization, the talent structure is actually much harder and more important than the technology question.
Nicole, tell us about this latest crop of AI agents that's being marketed for everyday use. Before I get into it, I think we need to take a step back and talk about the evolution of this technology and how we got here. When ChatTBT landed on a scene that was a large language model that was generating answers largely based on a corpus of data. There is some debate as to what is agentic and what is not agentic. But an agent is essentially a bot that can work on its own and click around your computer or the internet and do tasks for you. I think up until this point, the primary use case for agents was in technical work like coding or other enterprise flows. And now it's come to our phones. Metas Muse was immediately downloaded by millions of people and that's in part due to Metas' crazy big distribution network, which is Instagram and WhatsApp, to the most popular apps in the world.
There was promo at the top of those fiends that was like, hey, try Muse. It's a personal agent that can do stuff for you and who doesn't want an executive assistant to take off tasks off your end list to do list. And so I think that the sort of practical use case of AI agents is like your end list treadmill of life admin, whether it's filling out school forms for an upcoming field trip or paying dental bills. Or looking for job listings every single day. Now you can dispatch these AI agents to do that stuff for you. And there's a sort of privacy trade off here, which is that the more data you give them, the more they can do for you. Like a real human executive assistant. But Nicole, can I unleash these agents on every customer support agent that I have to deal with? Because that's really what I want to do with it. I want my agents to take on the other agents.
I wished this so much because I actually had to call Walmart in order to fix a refund that was not issued to me. And I tried to ask my agent to call Walmart and it couldn't it couldn't do it. So it gave me the phone number. And when I call the phone number, Walmart said, hi, I'm Walmart's AI agent. I see you had an issue with the refund that we tried to to process for you. Are you going about that? And I thought in that moment, like if only agents could talk to each other, have a little robot off. And I could just sit my coffee and do anything else. You know, that would be a swarm, right? If they did that, you know, that that swarm is good. That's where we're hitting though, right? Like, you know, the companies have AI agents. We're going to have our own agents. They're going to be talking to each other. I don't know what could possibly go wrong. I agree that it seems like this technology feels inevitable. The models are like getting better. The agents are getting more capable.
Companies have their own agents now consumers have their own agents. It seems like a really scary idea to me because it sort of locks us into that agent interface. And we have to trust it. And, you know, basically we've had all these problems where we trust these tech companies with navigating the world for us and they steer us the wrong way. People don't trust the algorithms. And it seems like that world would be one where just, you know, you're just asking the agent, what's the best price on this? And, you know, maybe you're getting it, maybe you're not. And you're trusting your human judgment to these agents. I think your spot on bomb. We're also trusting these agent services with a ton of our data. And this is what gives me the most pause in many ways. It's already too late. Over 5 million people have downloaded mues and are using mues. The question I have is when will these companies come up with a solution that will allow us to use agents in a way that shields our data and exposes our most sensitive information?
Are they seriously trying to do that? Like in a serious way, I mean. When you look at the connectors that are available for these agents, so these are services that you can hook up to your agent to use and have visibility into to do stuff for you. It includes everything from banking, your payment methods, blood work, Google Drive, Gmail, Google Calendar, Apple Health. It's a lot that you can give. And you can also just, you know, tell it stuff. Like you can say, I want to apply for global entry, like fill out this form for me. And it might ask, you know, the one thing that I don't have is your social security number. Can I have your social security number? And something along those lines happened to me. It, you know, surfaced a dental bill I had completely forgotten about in my inbox. So I had already made the mistake of hooking up my Gmail because I'm a professional guinea pig. I do this for science. You know, do do what I say not as I do. And it was able to fill out this very involved form that included things like, you know, my birth date, the date of the appointment.
And I didn't necessarily tell it to do that, but it was able to look inside of my inbox for those details. And I was sort of watching this robot fill out this form for me. And I think that was the first time I thought this is a both magical and terrifying experience. Meda said that muses virtual machine, which is called secure VM for secure virtual machine is isolated. And so other agents can't access this machine. They've also said that data from connectors can be forgotten, muses can be reset and that they're safe to use. Well, we often don't know what the trade off is going to be. And so at the beginning, it's all convenience. It's all like this is a miracle. This is so great. And then we kind of it seems like with technology, we often like years later realize that there's so there was there's a price we were paying for that. I mean, I also think that we're pretty good at like figuring this out over the long term. You know, there's lots of technologies that have had negative consequences that we just made the decision to accept those consequences because the convenience.
Yeah, I think that a lot of us have figured out how to not click on that suspicious link in case we might be hacked. I mean, Bob's reporting may be back to differ. But I think that, you know, fishing was a huge problem at the dawn of the internet and when we all had access to email. And now it's a little bit less of a problem and agents will probably eventually make less mistakes like that. But I think they're going to be a target too. If I were a hacker right now, I'd be looking at these these AI integrations and just trying to hack the hack out of them because they're going to become the keys to the kingdom for everybody. Because Nicole has given it her social security number. To be clear, I have not so please don't try to hack me. Please don't have me. I think consumers are facing the same thing. They now have access to this very powerful technology and they are in real time learning the boundaries of agents and that you can't be too permissive with them.
Well, it sounds like you're going to have your work cut out for you both of you really and me too. There's more than enough to report on. So we're going to keep reporting on all of this. You can check out our reporting as always on WSJ.com. And we hope to help folks understand what's happening. It's definitely changing very quickly out there. Thanks, Bob. Thanks, Nicole. Looking forward to next time. Thanks, Sam. See ya. That was WSJ Deputy Tech and Media Editor, Sam Schechner, Cybersecurity Reporter, Bob McMillan, and Personal Tech Columnist, Nicole Nguyen. And that's it for Tech News Briefing. We'd love to hear from you about how you're approaching using AI agents. Or if you have any questions for us. Drop us a comment if you're listening on Spotify or shoot us an email at tnbadwsj.com. Today's show was produced by me, Julie Chang. Jessica Fenton and Michael Laval wrote our theme music.
Our supervising producer is Katie Ferguson. Our development producer is Aisha Al-Muzlim. Chris Zinsley is the Deputy Editor. The Tom Malad is our Senior Director of Shows. And Samantha Henneg is the Wall Street Journal's Head of Multimedia. We'll be back Monday morning with a new episode. Thanks for listening. The company's seeking to fuel growth with AI should resist the temptation to rush into big swings, says McKinsey's Steve Reyes. It's very important to start with things that feel winnable to build momentum in the organization. You're building that muscle, you're also building the mindset of people being curious and figuring out how to use these tools and to think differently about what's possible. Reyes believes the principles that have driven growth for decades don't change because of AI. The highest AI performers use it to fundamentally redesign workflows, not just automate what's already there. Demonstrating how AI's real value is in its potential to reward ambitious thinking.
The exciting piece is to say, hey, this is how we do things. Put it to the side and ask them our fundamental question, what's the outcome we're trying to drive? And then how can we use AI to rethink the way we do things to enable that to be better? And what we typically find is when companies do that, they're rethinking the process and enabling people to be more effective, not just more efficient in drive growth. For more, visit McKinsey.com
More episodes
More from WSJ Tech News Briefing

Inside the Race to Build AI Defenses Against Bioweapons
WSJ Tech News Briefing

TNB Tech Minute: FTC Opens Anthropic and OpenAI Investigation
WSJ Tech News Briefing

TNB Tech Minute: Trump Defends His Light-Touch AI Regulation Strategy
WSJ Tech News Briefing

TNB Tech Minute: OpenAI Announces Latest Contender in AI-Assistant Race
WSJ Tech News Briefing
