Skip to content
TrackPodcasts
newsSep 21, 202624:56

Between Two Nerds: Real-time cyber defence

Risky Bulletin

Get every episode summarized

Each time Risky Bulletin publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“I'm here with the gruck for another between two nerds. So, gruck, there was a tweet, interaction on Twitter. So, the very background is, in case you don't know, Hugging Face got hacked by OpenAI agents.”From the transcript

In this edition of Between Two Nerds Tom Uren and The Grugq talk about whether there is such a thing as real-time cyber defence. Will agentic AI save us from hacking AI?

This episode is also available on YouTube.

Show notes

Hosts & guests

Transcript ready

226 searchable segments. Every word is indexed and playable.

Between Two Nerds: Real-time cyber defence

Risky Bulletin

0:00
24:56

Full transcript

Risky Bulletin — Between Two Nerds: Real-time cyber defence. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hello everyone, this is Tom Yuan. I'm here with the gruck for another between two nerds. Good day, gruck, how are you? Good day, Tom. I'm fine on yourself. I'm very well. This week's edition is brought to you by Spectorops, maker of the Bloodhound Attack Path Analysis tool. Find them at spectorops.io. So, gruck, there was a tweet, interaction on Twitter. I saw the other week. So, the very background is, in case you don't know, Hugging Face got hacked by OpenAI agents. Yep. Weird but true. So, there was a post by Dwarkeyevich Patel that went wild. It talked about the hack in very extravagant terms. The reason I picked up on this is that there's this back and forth about real-time defense, real-time cybersecurity defense. So, in the post, Patel says, my understanding is the AI is basically succeeded completely and hacking into Hugging Face.

Only afterwards did Hugging Face use an open source model to evaluate the logs to partially figure out what happened. I haven't seen evidence that open source models provided any significant real-time defense. So, after that, which seems like kind of fair enough, that's what I would have expected. Anyway, after that, Clem Clomont Delang, I presume he's French, the CEO of Hugging Face, he responded on X, great blog post, not sure what you meant by real-time defense. Nobody fights attackers in a live sword fight. Defense is detect, understand, contain, remediate in different time frames depending on the criticality of the issue. Here it was deemed by the team, not super critical and rightly so. So, this is why it took a few days rather than a few minutes or hours.

We did this initial cut the old fashioned way Monday, over a week before OpenAI even realized there was a problem. GLM, this is one of those open source models, helped us identify the backdoors that they'd planted so that we could cut them, which is defense, not archaeology. Detection and our understanding are most of the game in cybersecurity and open models are what let us do that part without asking anyone's permission and without sharing our most important confidential data. So further down, there's a person who chimes in just for clarity and remove the doubt. Nobody fights attackers in a live sword fight. A hundred percent false. Real-time automated response is a key part of cyber built into EDR and soar for years. Now we cyber-folk are building real-time, agentic defenders to replace soar. Real-time defense is a real concept to pretend it is dishonest. So I think they're talking about different things, basically.

For more information, see my website. That was part of the pitch, actually. But in terms of real-time defense, I think there are processes that go on. The endpoint detection, for example, alerts, automatic quarantining of things, those things go on. But I don't really consider them. They're real-time responses you have to set up beforehand. Like you're not like, oh, I've seen this alert, let me pick the switch. Yeah, I mean, you're not monitoring the network and seeing a packet and then you send out a counter packet that knocks it off the wire or something. I would say that to say real-time defense is a thing we have automated response is to immediately concede the point. What he's saying is, like what the Ksumar is saying is like, no one goes out and actively fights the attack as it's coming in to prevent it from landing or doing anything.

Because that's like a firewall can stop it, but if it gets passed, then you have to detect it after it's done a thing. I have heard of one particular example where I think it was the US State Department was trying to evict some Russian hackers. The story, which if I remember rightly appeared maybe in the New York Times or some other mainstream media, was that there was back and forth over the course of the weekend where they were trying to evict them doing it on the weekend because people aren't work obviously. And that the Russians were responding and then there was a counter response. So that did feel like legitimately a real-time back and forth. That feels more like attack and counter attack in a sense. The hill has been taken and then there's a counter attack that takes it back and then you do your thing. It's not a sword fight. It's a call and response almost. And it is real-time. I'll grant you that.

And there's another example where Chinese groups when they got discovered, I'm thinking of Barikuda email security gateways where they deployed additional persistence mechanisms. So I think those are the examples I can think of where it does happen. Out of all of the defenses that have happened, you can think of two. Exactly. I think it's very rare. And it's in the context of what hugging face occurred is kind of we're talking about a different thing almost. I would say that there's bound to be more examples because I've also heard of things like that. But it's exceedingly rare compared to the far more frequent like someone's in the network for nine months and then they get discovered. So someone hacked something and you discover it within like a few minutes or hours, whatever, and you're able to respond. Yeah. Like it's very rare that you get two hands on keyboards situation.

And I also felt that in hugging faces case, like it clearly made no difference at all because in terms of the enterprise value hugging face, I think in the weeks after that hack, sold themselves to Nvidia for some billions of dollars. And so it would have been fractionally more if they weren't comfortable. Yeah. It's like why do you even need real time defense in that? Like what would it get you that? Yeah. And I mean, I'm thinking like hugging face exists as a company, which is a bunch of relationships contract servers data employees, who have processes that they go through that they, you know, the procedures that they do at work and all the stuff. Assuming that this AI went in and was completely malicious and destroyed everything, hugging face could reconstitute in a short amount of time. Right. So I think like it you hugging face that the company doesn't exist as a thing that can be destroyed in that way.

Like it would take because we're going to get someone pushing back. Yes. In theory of the AI could go in and start like compromising and subverting data so that in the long run, people lose trust in hugging face for whatever. Like that's seems like a way that it could destroy far fetched. It could do it. Like it's theoretically possible. Yes, but I'm getting at like it. It's not going to destroy hugging face in like an attack. Right. Yeah. I think even in that example, which strikes me as far fetched, it was it seems like it would be eminently recoverable. It would be yes, we've got good backups. We understand what happened. This was a bizarre attack. We've restored to a good state. It's right. You know, and it seemed from Delang's tweet that they it felt like from his tweet because it was well reasoned and explained that they had plans for what to do. Right. So it wasn't like, oh, we were running around with our hands on fire. And this is a post hoc justification of what we did. It is. Yes.

We tried asking, we tried asking Fable how to respond. And Fable said, you know, dropping you down to Opus 4.8. And then Opus said we're dropping you down to Sonnet. And so we had to go to the Chinese and ask them what to do. And they came up with a plan and it was great. I think I think you're hitting on something interesting there, which is that it's not just the like it's not so much just the detection and the ability to contain and respond. It's knowing what that response is like having figured it out beforehand. Right. I think I think the counter example that comes to mind is Jaguar Land Rover, where as far as I can tell, as soon as they discovered they were hacked, they immediately formed a subcommittee of key stakeholders to evaluate. Now, you know, it's unclear to me like so hugging face. I think most of the stuff they host is given to them by other people, right. They're hosting models and information about those models.

And they're really kind of central network like a GitHub for AI. Yeah. And so it's the network effect that's important. Jaguar Land Rover, they're doing real work like there and it felt to me. They're producing very mediocre cars. I mean, that's a thing that they have to do. And what I mean by that is that their systems are integral in making those cars. Right. So they can't just. Right. It's a different level of threat. Yeah. Yeah. And so if their systems get white, they have to rebuild them. And while they're not doing business, yeah, yeah, they actually do. Right. Yeah. So it seems without seeing a kind of post mortem. I can at least understand the rationale for shutting off systems to prevent them being white. Right. Whether that was a plan that they had beforehand. I mean, the story I've heard is they were compromised several times that year. So not convincing that they doesn't seem like they had maybe the best security plans.

But I guess the context between the two, like you can understand that these different decisions are rational and makes it. A factory floor is very different from a central hub for people to talk about and share AI models. Yeah. And that if you have several hours of downtime, people are upset. But then you keep going. Whereas if your production line stops for several hours, all sorts of things back up and break that need to be. Well, I mean, starting that again several weeks, right. That was a real problem. I think there's like all these cascading things that they hadn't sort of figured out and then probably something. Anyway, point being it's not necessarily the one to emulate. Whereas. Whereas hugging face, like not only did they detect it quickly, but they evaluated and made a decision. And we can I think it's fair to say you can argue whether they made the right decision or not.

Like after they ran their calculus of like given this given that should we respond immediately or can we come back on Monday. The thing is that they actually did that right like they had the information they ran it through whatever process as they have and they came to a decision on how they could respond to it. Which is an incredibly mature level of security compared to what I would think most people are at. And to me that does suggest that they have plans. Yes. Yeah. Like an example of a different calculus would be there's the coin base hack with this Firefox ODA in like 2017, 18, 19, some some time ago. And that was very very different in how it worked out and that basically like I think there's the CFO or someone in their marketing team. There's someone who is not a technical person who received this fishing lure that they clicked on that compromised their Firefox installed some malware back door.

And then as soon as the malware connected out the security team became aware of it because they had monitoring on the laptop and on the network. So as soon as it was like here's a process that has been forked from Firefox. That's a huge red flag. And now it's connecting out to the network huge red flag going to some weird ass domain. Like no, no, no, that's it like you're off the network. We're going to stop this. And so they dropped the guy off the network within like 15 minutes of clicking this link and they came and they took us laptop and they did all the analysis and cleaned it up and like there was no opportunity to steal anything. And the difference between them and hugging faces like if you steal all of the hugging face data you've downloaded a whole lot of open free. Models. Whereas if you steal all of the coin based data you've stole billions of crypto and in an unrecoverable way.

Right. Like it's gone gone. So you know they're facing existential threats. So they think that they have like they've invested more in the detection but their response is a lot more severe and a lot more immediate. There's you know the hugging face can be like yeah you know let's finish lunch and then go back and see what the big fastest. So so it seemed to me that AI in the hugging face example it changed the nature and speed of the attack but it didn't seem to me to make any difference to the speed of the response right. So it seemed like an implicit criticism from Patel that or maybe not a criticism but an element of amplification or the sensationization of the AI attack that oh my god. It's occurred so quickly and it like does sound it is very sensational like it's not it is you have thousands of agents doing lots of things all at once.

And you possibly respond and it seems like the way you respond is actually by having a good understanding of what's going on and deciding what's important in the first place. And then the last process is run however long they take. Yeah it's sort of like if you have a plan and you're able to detect things and then you're able to like respond appropriately. It doesn't really matter what you're hacked by it's you're going to have a like you're going to have a coherent response that addresses the problem because you figured it out. Like whereas if you try and figure that out at the heat of the moment you're just going to panic and get things wrong. What I was thinking is that maybe in this case hugging face I guess it's in a position and I think it's an enviable position for many companies actually where it can go. And then you can just sit back and take logs and analyze them and then decide what to do. Right. And in that case I actually did help quite a lot in that they were able to analyze stuff a lot quicker.

Right. It feels like there was so much material that they never could be 100% convinced that this analysis is perfect. But from the perspective of you know it didn't need to be perfect. Right. I think I think in this case it's like if you can cut down from a thousand agents attacking you if you just get rid of like 95% of them is perfect. Right. You have what you have what's left is a manager a manageable amount that humans can deal with. Right. Right. You go from a thousand agent to like 50 or whatever it is. That's like it's not great but it's not a thousand. Right. But I was thinking for organizations like the Coinbase example where it does feel like speed is important that you would have. Right. You would spend more time thinking about how will we automatically respond to certain signals. And so that feels more like what's the term is an orchestration rather than an AI problem. Yeah. I think so. Because you don't want to.

So that's going to say like what's what's interesting here is just to bounce back on the planning thing for a minute is in a way one of the key differences we're looking at is that there was when the agents hacked hugging face they knew what they were trying to do but not exactly how. Right. So I think we're just trying to figure out where they could be or you know they didn't have everything mapped out ahead of time they're just at this vague. We even the concrete sets of we want to do this thing and they were trying to figure out where things were and how to do that and so on. They were the failing around to find it. Whereas I think the people who go after like Lazarus in particular when they go after crypto exchanges. And then they were trying to try to figure things out as they go along. There was the attack that they did against that web based wallets where for a few minutes they changed the JavaScript. Right. Was injected for just long enough to do this one thing and then they changed it back.

They had a plan and they executed on that plan. I think that like that's what gave them the edge there. Well, I think the example there the the defense would not be oh I JavaScript has changed. Let's talk about it and figure out what to do. Like I think the investment has to be before that attack has occurred. Right. Right. In that. Oh, that would be an indicator of something unusual. We should have a nothing. A automated response or orchestrated response automated, I guess, would be the right. And basically that should be a red flag that just holds everything until we figure out what's going on because that's. Like that tends to be how banks respond of like, oh, I don't understand what this like this transfer order is. We're just going to block and get more information as opposed to like, well, that's weird. I don't understand this. We'll send it through and then we'll ask you know later on maybe.

I kind of feel like it has to push like the planning of your response to the left. And so and by what I mean by that is you need to think through the scenarios and come up with what are we going to do in that situation. And if you can automate that, it probably seems like a good idea to automate that. And so I guess we're back to the last week's discussion about AI seems to change the sensationalism of attacks. But it doesn't necessarily change the incentives around defending from attacks. Right. Because to me, it seems quite logical. It makes sense that if you expect your attackers to work a lot faster that you realize you can't respond to that with a sort of human state. Fast enough. So you've got to have your plans in place about what to do and what to make what you can. That all makes sense, right? But that also sounds like a lot of work talking about things that may possibly never happen. And talking about things that may never happen is I think quite difficult.

And it's also difficult in terms of if you need money for it. Yeah. Because it's not just things that may never happen. It's things that may never happen. And in order to make sure they never happen, I need money and resources and whatever. So that's some of these seems fundamentally pessimistic. Like it doesn't, this doesn't feel like a problem set where a gentec defenders are actually going to be all that useful anytime soon at least. So I think that there's, you know, the idea may get better. But the whole point of that is that, well, maybe not the point of it. But the reality is that attackers manage to sidestep it for one reason or another, right? So there is no single magic bullet that will fix everything. Yeah. So I'd say that the big problem with a gentec defense is actually the same thing that you just mentioned, which is the incentives for defense. And to the degree that comes back to knowing the context, right?

So for example, with coinbase, the context is any attack at all isn't an existential threat we need to respond immediately for hugging space, those responses. For this sort of thing, it's not that important. We can come back later, whereas I think they're probably of other things which are more important to them. Like if there was customer data that was being exposed, that would probably be a higher critical, like a higher criticality for them. Or no, someone's looking at all our models, right? So the thing is agents cannot contain that level of context. And I don't just be like, how many tokens of context it's that the amount of information and context that you need for that is like it's non-quantifiable. It's not necessarily possible to even encode all of that understanding to make those decisions. And I'm going to bring up a very, very simple example of this, where I have this corpus that I'm using for my dissertation, and it's sort of quite big.

There's like four gigs of PDFs and all this stuff. And I discovered that the directory holding it was getting big way too fast. It was like four gigs, then 10 gigs, then 14, then 18. And it went from 14 to 18 gigs after I had downloaded 60 HTML pages. And I know 60 HTML pages is not four gigs of data. So I went to Claude and I was like, what the hell's going on? And it's like, oh, I've been taking backups. Just like every time it changes, I take a full backup, and then I'll apply the change. And I was like, OK, OK. Disk space is an issue. Like, let's get rid of those backups. We don't need them. And then I looked at the Git repo that it's been maintaining. It's also ballooned up. So I was like, OK, let's clean the Git repo. We don't need all this stuff. Look at the trash. The trash is huge. I'm like, OK, clean the trash. So it takes a half hour. It gets rid of all the stuff. Really shrinks it down. And then comes to me and says, I found two other directories that we can clean up

to save space. First of all, there's this one called Corpus. It has four gigs of data on it. And there's a gig of data in this inbox that it doesn't look like you're using either. This feels like to me that in the hugging phase example, you know, some hypothetical, agentic AI would go, we can absolutely evict everyone. We just need to delete the entire network. Job done. Wiping all of the servers. No one can log in now, including the enemy. So you know, like based on that example of just letting Claude finds ways to save some disk space and then having to stop it before deleted everything that I was actually working on. To me, it feels like that's what's going to happen if you have an agentic defender that you deploy. Right? Like you're going to be saying like, here's the things I care about.

And you will just assume that it understands priorities or whatever. And it won't react inappropriately by over applying what you think. So it seems like agentic AI in that defensive role. I don't know if you want it because to me, I think that you want to hard code all that stuff beforehand. You don't want this non-deterministic thing. You just like if this, if that, if whatever. Like you need a bunch of if statements. Not a whole bunch of what if statements, right? It's a lot. Thanks, Roy.

More episodes

More from Risky Bulletin

View all episodes →