
D2DO315: Running Agents Securely at Scale
Get every episode summarized
Each time The Fat Pipe - Most Popular Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“The Gartner IT Infrastructure Operations and Cloud Strategies Conference December 8th through the 10th in Las Vegas brings together heads of infrastructure, cloud ops pros, platform engineers and security leaders.”From the transcript
Get every episode summarized
Each time The Fat Pipe - Most Popular Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
464 searchable segments. Every word is indexed and playable.
Full transcript
The Fat Pipe - Most Popular Packet Pushers Pods — D2DO315: Running Agents Securely at Scale. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Today's podcast is sponsored by Gartner. The Gartner IT Infrastructure Operations and Cloud Strategies Conference December 8th through the 10th in Las Vegas brings together heads of infrastructure, cloud ops pros, platform engineers and security leaders. Learn more at Gartner.com slash IO. Efficiency is top of mind for everybody, better results, better quality. It's all there, but it's just a matter of the industry needs to... you can't put best practices around something that changes every 18 hours. Welcome to Day 2 DevOps where the dev whoops is in the details. I'm Ned Belvance and I'm joined by my prodigal co-host Kyler Middleton. Hey Kyler. Hey Ned. Today we're discussing all the things AI. How do you observe it? Where does the data come from and how do you store it? Will we have enterprise AI on our laptops? Will we end up buried in
paper clips? Should I keep asking questions to myself? Let's talk about all the things. Guiding us through all of this AI is our guest Michael Levan, a prodigal content creator in his own right and an AI architect at solo.io. In fact, he is hands on doing this stuff all the time. So this is not just drawing stuff on the whiteboard. He's down in the trenches doing it. So let's talk to him about it. Michael Levan, welcome back to Day 2 DevOps. We're very happy that you're here. Really enjoyed our conversation last year. All I'm talking about. What do we talk about? We talked about wasm last year. Something else has come up since the end. Something else has become important in this face. Why don't you tell the good folks out there what you've been up to since we last talked to you? Yeah, last year. Okay. So one of the big things was I went from being self-employed to joining a company called solo.io, which has been awesome. And I've done everything in anything right now. I'm
an AI architect. So focusing on all things agentic across the entire stack from agents to your gateways, to your registries, to your evaluations and everything in between. And since then, yeah, I've definitely moved from like the standard, I don't know what you would call it. I don't even know what you would call it now, but the standard cloud native pieces of the puzzle to definitely more AI focused more agentic. But again, it's like so much of this still runs on the environments that we were running everything else on already. Like the majority of all things agentic are still running on Kubernetes. People still need a way to orchestrate and scale out agents. And if you're not using a managed service like a foundry or a better aqua or something, you're probably running it in Kubernetes because that's the place to scale nowadays. So there's a lot of overlap. And another thing that I would say to you is like, if you have the background in monitoring, observability,
performance optimization, call optimization, programming, architecture, like you're not just throwing away that stuff. You're bringing it with you a thousand percent, you know, like a hundred percent of what I've done throughout my career. I am now using in the agentic space. How did you so I guess I'm curious was it a natural progression for you to end up in this AI architect space from where you were before? Was it a series of consulting engagements that got you into it or you were just tinkering on your own personal curiosity? Like what drove you to dig deeper into the world of AI? Yeah, so I think it was three things. Number one, it was just tinkering, right? Like the the reason why I've always stayed up to date with my career has been because I just enjoy this stuff, right? Like when people ask me what I do on the weekends and I'm like I'm programming for fun and then during the week I'm programming at work. So like I'm just always interested in this stuff. So originally I was like building out agents from scratch. I was like looking into like fine tuning
models, seeing what went into building models. Obviously I didn't do that because I didn't have millions of dollars for any of the data centers. But I was able to dive into like light fine tuning with Laura and stuff like that L-O-R-A. It's like a memory optimized way to fine tune models. So I was diving into that and then I started diving into the frameworks, crew AI, chain, all of them. And then just kind of like one thing led to another, right? It was like it was popping up more and more, more and more customers needed to dive into it and talk about it and see how it worked. And then I just transitioned my way into it. So I think we were talking earlier about the fact that you are just speaking to a ton of different organizations about AI related things. And I'm curious like what are the top concerns they're coming to you with? Like if you could name like three top concerns that they're like these are things we're either worried about or we feel like we need to do.
Yep. Yeah. So and ironically enough, the top three concerns are across all organizations. Right? Like I'm talking small organizations to the largest financial providers, the largest telecom providers, government agencies, everybody. Everybody cares about three things. How do I secure the thing? How do I observe the thing? And how do I get cost management and optimization and check? Like those are the three things. Like you could like you could you could essentially like walk into a discovery call at this point and like hit them, hit the the speaker mute button and come back on at the end and like, okay, so you want these three things and it's like, yeah, those are the three things that we want. Yeah. And and I'm I talked to a lot of folks like I'm on anywhere from five to 20 customer calls a week. So and like this is just what I'm constantly hearing. Like a lot of this stuff, I would say is at the gateway layer because you know, people are thinking like I need to
get from point A to point B. Right? So I need to go from agent to LLM or agent MTP server or agent to other agent. And there's this line of communication in the middle. So what is that line of communication? It's going to be your gateway. So everything going out, everything coming in is traversing through your gateway. So you got to figure out how do I observe the traffic? How do I secure the traffic? What do guard rails look like? What does prompt guarding look like? What does call optimization look like? What does general performance look like? So a lot of people start there and then they start to think about like, Oh, actually, I need a runtime because if I don't have that, I can't run agents. And then I need a registry for a server serving securing storing MTP servers, agent skills prompts. And then I need to put governance around that for the whole shadow way, I think. So I need auto discoverability of my agents that are deployed once I have my runtime. So a lot of it goes into this like a genetic stack as a whole. But I would say a lot of people started the gateway layer
unless they already know, right? Like they've been playing with this stuff for a year. They're playing around with open source projects. They're like, All right, I have an idea in my head of what I need. I need x, y, and z. But the space is so new right now that it's really difficult to do that as well because we're all just figuring it out as we go. So on AI gateways, I know that concept is pretty new of an AI gate. I've seen MCP gateways for maybe a year, maybe a little longer, but AI gateways are like a pretty new concept right now. Who does it well? Who should we be looking at? Are there open source like models that I shouldn't say models because that's overused here. But open source products that can help me be in AI gateway and secure it in my company. Yeah, absolutely. So like really what you want to look at is how a gateway is architected from the ground up. Okay. So a lot of the times what ends up happening is people essentially take their standard on boy gateway and they put a third party library package on top of it to be able to
handle agent traffic. So like if we just dive into it architecturally with envoy when envoy first came out and it became more and more popular and now it's crazy popular. A lot of the concern was microservice driven stateless traffic. Like of course you had stateful traffic and such there. But really what you cared about was the headers, right? Like you didn't care too much about the body of a response. Now with agent traffic is upside down, right? Right. You really care about the body of the response, especially when you're thinking about like MCP traffic, right? It's all adjacent on our PC. You're looking at the body of a request. You're looking at the body of a response. So now you have to think, okay, I need a completely and entirely different gateway than the standard envoy model, right? So now you got to ask yourself a question. Do I want to take a gateway and just pop a third party library on top of it? You're going to have performance issues. You're going
to have integration issues or do you want to build something from the ground up for agent traffic? That's that's really what you need to care about. So and that's what agent gateway does, right? So if you can go you can look at GitHub, you can look at the open source project. That's what it is. It's a ground up built gateway for agent traffic. It's written in rust, right? So if you're thinking about it from a performance perspective, right? Like if you compare a Python based gateway with a rust based gateway, you are going to see a night and day difference in performance. And maybe, right, a year ago, people weren't really thinking about it all that much because they're like, oh, like why do I need to care about performance? I have a web chat bot for customer service and I have one or two SRE agents who cares. I don't have to think about performance right now. But as we continue down the path and especially now I'm seeing this all the time where people are like,
well, I actually need an agent for this and I'd love an agent for this and I'd love an autonomous agent. And now the agent substrate project came out from Google about three weeks ago. And that's actually a sandbox project that you use their control plane and their management plane to control agent sandboxes. And then you're using Kubernetes for its resourcing and for its orchestration and clustering. So now people are like, okay, I have a performant, efficient, and secure way to work with sandboxing agents. It's running in Kubernetes. So now there's nothing stopping me from having tens, hundreds, maybe thousands of agents. And then at that point, if then you got to start thinking about performance and a lot of companies are getting there quickly. I would be super concerned with cost management at that point too. If I'm scaling to thousands of agents, then that's thousands of requests to whatever models I'm consuming. And if those
models are like paper token, woof. Yep. Exactly. A thousand percent. And that's really even where like the gateway could actually come back into play there. So I'm thinking about like semantic routing for a second. One of the many things in semantic routing, the thing that I enjoy the most is the way that you can route to models based on what the agent is doing. So if I have a really difficult task, right? Or I'm doing rejects matching or like, however I'm doing it, the agent says, okay, I'm going to go hit opus 4.8. And then if the task gets easier or you're past the architecture stage and it's more simplistic programming, then it's like, okay, now I can automatically drop down to solid, right? Because that's where the cost comes from. The tokens in may not be hitting you all that hard, but the tokens out is what cost you. And the reason why is if anybody's wondering the reason why models, the better the model is the reason why the tokens outerm
more expensive is because it takes more processing power to run those models, which means more hardware, which means we are paying for the hardware that the models are running on. Right. And the tokens that are output also includes the intermediary tokens when it's doing reasoning. So if you're using a reasoning model, that's even more output tokens that it's generating. And that's the real cost back up to the provider. They're going to nicely pass that right along to you. That's a percent. Yeah. And that's really again, like it's funny how this all ties back to the gateway as well, right? So like, for example, taking that cost thing into play, let's say you're hitting an MCP server and this MCP server has 50 or 60 or 150 tools or however many it has. When an agent hits an MCP server, even if it's just using one tool by default, it puts all of the tools into context. Right. So that means you could hit 100,000 tokens
before you even put your first prompt in, right? And that's where things like code, mode, and progressive disclosure to just only expose a couple of tools versus all of them come into play. But that's all set at the gateway. So it always comes back to the gateway. Interesting. I haven't heard of code mode before. Can you expand on that briefly? I don't want to make this a code mode episode, but you know, absolutely. Yeah. So there's two things. There's progressive disclosure, which is it shows two tools, right? Invoke tool and get tool. So instead of the agent putting the entire schema of an MCP tool into its context, it only puts in like the name and the description, right? So it's still taking a little bit, but not as much, right? So you only have two tools with code mode. This is actually to get past the whole, what if I don't want to use MCP? Okay. So the way that the way that it started
was like in the agent space was like, hey, I have an agent and this agent's going to go hit an API, right? I'm going to have my agent go through G Cloud or QCTL or AZ Clire or whatever, right? So it's going and it's hitting an API directly. Then the question was like, well, how do I secure this? Right? The agent, the agent can just do whatever I can do pretty much. Yeah. Whatever my credentials are locally. So then everybody said, okay, we need a standard around this. We need a schema around this. So the MCP came out. So now we have a standard, but now we're back to the, what about if I just want to hit an API? Well, here's the problem. You still got to worry about the observability around that. You still got to worry about the security around that. You still got to worry about the governance around that. So what code mode does is code mode allows you to go and hit an API. Like I always use the OSS geolocation API as an example. So I can actually have my agent go through my gateway, hit this API and then use the tools within this API. The tools are just
the functions or the methods right within the code base. And then when that traffic comes back, it actually translates to MCP traffic or the MCP schema. That way it still looks like MCP. So the agent can understand it. It can read it. It can be secure. It can be governed. It can be observed. But it's not MCP traffic. You're still able to go and hit an API endpoint. Okay. So solving for two different problems with two different I got you. Yeah. Yeah. And both of them kind of come together from a call-s optimization perspective and just an overall usability perspective. I want to back up a little bit to model routing. I haven't implemented that yet. I've been reading about it. And I'm curious. There's two major use cases for AI. There's like the chatbots that talk to your customers and promise to give away cars in Canada like two years ago. That was hilarious. And then also there's your like harness, cloud code in your environment. Are you seeing model steering for like either for both the real enterprises like do
model steering for a cloud code like through an AI gateway? Is that a thing? Tell me about that. Yeah, totally. So depending on how you create your harness, right? And just to define it, because I feel like the industry is kind of defining harness isn't a different way right now. Like how I define a harness is everything around it. My skills, my MCP server, my prompt, my instructions, et cetera, right? Your harness could be what you're building in the agent. Your harness could also technically be cloud code or open code or code X, right? Because you have all that stuff built in. Is that like kind of what you're thinking about when you say harness? Yeah. Yeah. It makes a lot of sense to me. Like I have an agenteic workflow and I control it. It's inside the enterprise. But for developers, I'm curious about forcing them to use a less smart model. And how much they can get up in arms. And I'm curious if anyone's done that well. Yeah. So funny enough, if you opened up cloud code right now and you were working on it all day and you went and you looked at the output logs, I would bet money that there's something in there
that during some task, it shoots you down from high mode to low mode or from opus 4.8 to saw it, right? Every time with reading a document or a picture, it's using high cool. And I had no idea until I looked at the usage. Yeah. Yeah. So that's yeah. So by default, your harnesses that you're using are probably doing it already. Now from an enterprise perspective, yeah. Like I personally think it makes us ton of sense. There's a project called VLLM semantic router, not just VLLM, not the not the inference underlying piece of like if you want to run your own models and stuff, but the VLLM semantic router, which is exactly for all this stuff. So like let's say I want to route from one model to another based on a task that I'm doing or a language that I'm using or the VL keyword, however I want to do it, I can do it with VLLM semantic router. And yeah, I think it makes a ton of sense, right? This is the funny example that I always use to come
Italian is like you need to have cost optimization in place because the last thing you want is a $20,000 bills because somebody's trying to figure out the best check and palm recipe with their agent. Right? So if you're if somebody's asking these questions that are like not difficult questions that sonnet can handle versus like I pull down a code base and I go to my agent and I say ingest this code base and put it into context because maybe there's not a lot of docs around it, maybe there's not a lot of examples. So I needed to look at the code to understand the source of truth versus doing a web search. I want to know, okay, I'm going to use Opus for this and then maybe when I say, okay, my tasks are done. This is the type of demo that I want put it into an MD file for me. I probably don't need Opus for that. Like I can probably just dive down to sonnet because everything's already in context, right? It's already ingested the code base. It's already done the hard
work. Now I can have it do the simpler things, right? And save myself money along the way. I think it makes sense for every organization to think about that style of routing. When I asked you the top three concerns and you came back with security, observability and cost management, I was like, if you squint a little bit, that's exactly the top three concerns with cloud. We have been here before. We've solved for these problems just on a slightly different technology. So if you're out there and you're listening and you have lived through the cloud adoption migration phase of things and had to deal with those three top concerns, nothing has changed. It's just slightly new technology. Yeah. And to that point, I mean, really, the only thing that's different at this point is we're just using fancy terms, right? We're saying words like token and inference and the rest of the big fancy words when it comes to AI. But underneath the hood and reality, everything that we're doing,
like a perfect example with Agent Gateway OSS, let's say you want to observe traffic, right? Well, you need traces, you need metrics, you need logs. And Agent Gateway is a data plane proxy, right? So it's exposing a ton of data points. And then you can take that and you can put them into Prometheus, Grafana, Tempo, Loki, or DataDog, AppDynamics, New Relic, whatever observability tool you're using. And the way that you do that is with OTAL, right? So the tools and the platforms that we're using, this isn't anything new here, right? Like we're still using the same monitoring and observability stack. We're still using OTAL to send those metrics logs traces over to whatever the platform is. We're still running this on Kubernetes. It's in a Kubernetes pod and we're going to scale it up and scale it down and manage it the same way. This is all the stuff that we've been doing for however many years. It's just at a different layer at this point. Even like I wrote a white paper
a while back, I don't think it's public yet. But like my whole thesis around it was agents are nothing more than the new distributed system at this point. That's how I see it. I see an agent as just another server, as just another system that's performing a task, especially when it comes to autonomous agents and stuff. It feels very much more like, hey, this is a box that does a thing than anything else. I want to dig a little more into the observability aspect of things, because we talk to a few other people around what does observability look like for agents and AI? Because yes, we're using the same tools, but the signals we're looking for are probably a little bit different. What are the important signals that people are looking for when they're observing their AI and their agents? When I say observability, by the way, I'm talking about logs traces metrics. I know I feel like people have different definitions of what observability is,
but those are the things that I'm looking at. That's what they like to. That's like the official hotel. Beautiful. Yeah, that's it. So I would say a couple of things. Tokens in, tokens out, traffic as in where's the agent going? How many hops it took, right? Because what does an agent do? An agent goes and talks to an LLM and then that LLM says, oh, I need this MCP server tool, and then that gets sent back and then the agent calls the tool. So there's always, there's this back and forth within the context, right? So I categorize that as hops. So you're looking at that, you're looking at inference speeds. So how long it took to do the thing? That's why, for example, right? Like I have an Intel look behind me right there. And if I were to run whatever clients, open clawed, open code, whatever. And I was hitting a Gwen model that's running on that. It's going to take me three years to get a response back versus if I have a DJX Spark or a high-end
gaming machine or something with 16 or 32 gigs of VRAM. So like looking at that stuff, right? Like the performance of how long inference takes. And then I would say the other big pieces are what the agent is doing and what the agent is hitting from a security perspective. So like, should the agent have the ability to send that prompt? No agent should be able to say, delete all the Kubernetes clusters in my AWS or Azure environment, right? So it's probably something you want to be able to block. You could also look at it from like a header perspective of, okay, and this is where like agent identity and stuff comes into play. But I have an agent. This agent has an identity. It should be a read only agent. So it should only be able to use list tools and get tools and search tools, right? And you want to be able to observe that traffic to say, okay, these guard rails that I put into place, are they working as expected? And I feel like there's like a
bajillion deaf, different other data points that you'd probably care about. But like, if I had to go 50 foot view and we can dive into any of these by the way, but if I had to go high level 50 foot view, those are the categories that people are definitely looking at at the moment. A lot of stuff around cost. People just, it's funny, right? You know, organizations just don't want you to spend $20,000 on nothing for no reason. Who would have thought? I'm going to spend money back. And I guess this is outside the scope, but I'm so curious if you run into this. You see like two engineers and they have different, you know, spends that they're spending. How do you measure if they're doing a good job being efficient? Like, I think this all kind of boils down to like, is the code high quality or is the value of the engineer good? And like all of those are so squishy and human. Is there a, you know, mathematical way to look at your metrics and say, you know what, this person is doing a good job being efficient? Let's
copy them. As an IT leader, your job is to modernize your existing processes, deploy applications faster than ever and keep your entire organization abreast with the AI revolution all at the same time. Tyler, it's enough to make your head explode. It is, but you're not in this alone. Gardeners infrastructure operations and cloud strategies conference helps you and other IT leaders and senior engineers connect with experts in the community you need to keep your head intact. We know that the last year has been a decade compressed into 365 days and it shows no sign of stopping. You need to develop a strategy to keep up. Slow is fast, fast is slow, right? I remember about a rabbit and a turtle. Exactly. You need to chart your course before you venture into tumultuous AI-infested waters. Khan learned how to captain your boat and command other metaphors too at Gardeners IO conference on December 8th in Las Vegas. Learn more at gardener.com slash I-O.
And you can use our handy coupon code of DevOps to get 450 de bloons. I mean, the dollars off the ticket price. It all know over to gardener.com slash IO and check it out. That's gardener.com slash IO and code DevOps. See you there, me hearties. Yeah. So I think about that a lot. So I'll give you an example. One of the things that I do is I have a a repo in my GitHub work called the Gentick demo repos and I put everything in there. Everything that I work on I put in there. So what I wanted to start doing was internally. I wanted to start building out labs for all of our authentic products. That way we could give it to customers. But also like let's say like a new SE joins right or a new SE joins they have something to like go off of. So what I did I'm getting to your point I promise. So what I did was I have had my agent
go look at everything in one particular directory and say build me out labs on this make it look nice, etc. And then I said, okay, however you did this, build me a prompt. So if the labs change or I want to add a new agent, a new lab or something, I have a prompt that my agent will understand what exactly it needs to do. So to me, that's efficient versus every two weeks saying, hey, can you do this thing? And then it trying to figure out how to do the thing again and again and again versus saying, here's how you do a thing, right? Or you go and you create an agent skill for it, right? And you put that agent skill in your harness to be the lab creator, right? However you want to do it. So I guess to answer the question is it all depends on the tasks that you want to do. And it depends what you put around doing the task. And I want to the reason why I say it like that is because the last thing that I want is the industry to get to the point
where we're like looking at lines of code to measure who gets a bonus, right? Like we were there was a point in programming where like, hey, you wrote more lines of code, you get a bigger bonus, and then people were just putting crap over the line. So that's the last thing that we want with agentic to, right? Like we don't want people just building stuff for efficiency purposes to like get a better bonus or whatever it is. And I also think that we're still so early in the industry as a whole that it's so hard to mark like what efficiency is. I was just reading this. I think it was two weeks ago or three weeks ago where OpenAI was, was they put out a big thing essentially saying enterprises are canceling subscriptions because they can't figure out the ROI. Like what are we getting out of this? Right. So I still think we're at that stage where like we don't even know what efficiency looks like. What we do know is we need cost optimization and management. We don't want
token spend for no reason. We want quick results and we want accurate results. And I think that if you can get to those four things for a particular task, you are in the category of being efficient right now. And that's going to change probably tomorrow because everything moves fast in the world of AI. I feel like there's a couple of things we need to be able to measure. One is like developer productivity. So I'm writing code with my my agentic help in my code harness. And so you might want to monitor token usage by the developer and see are they burning through their budget really fast? And if so, maybe they're doing amazing things. And that's not like a mathematical thing. That's go talk to the developer and see what they're doing. And be like, no, no, you built a whole application that solves some real business problem. Go on. You know, keep spending those tokens. Yep. Or it could be they don't know best practices. And because best practices are like changing every single hour. But like they have that, you know, 26 different MCP servers hooked up all the time
loading 300 different tools into every prompt. So they're burning through most of their context window. And they're trying to use like 12 different skills. And it's like, okay, we need we need to help you a little bit. So that that's an opportunity right there. But on the flip side, if you've started using just autonomous agents to do things, you know, whether it's code review, or actually trying to help a customer or whatever it is, whatever that agent is doing, you also have to monitor their token usage. And see if they're being efficient with how they're going about various tasks. And maybe you go, you know what, maybe an agent isn't the best solution for this particular task. Maybe I should just write like a bash script. Right. Right. Sometimes that is. So yeah, I feel like you're absolutely right. We're so early in establishing best practices and tooling that what was correct and good two months ago. Now forget it. Like caveman was a big thing for a little bit. And I think that may have been like, like died down a lot.
Even if it does cut down your output tokens, because maybe your your your your quality response quality drops. I don't know. Like I there's not a beauty yet. Yeah, I think it's also really hard to have and maintain a project that you could literally do the same thing in your prompt just by saying max out at 500 token output for this prompt. You know, it's yeah, it's tough. But I would also say that the industry is definitely going down a direction right now where big providers primarily open AI, you know, your open AI's your anthropics, your Microsoft's like they're all trying to figure out the best way to have ROI on this because if they don't, they're not making money. And then they got to go back to the board members and the shareholders and say, well, this whole AI thing is not making us any money. So they're all working really hard on this. So one of the new things I think it was last week or the week before, uh, Boris from anthropic was talking about like the whole agent loop thing. And the idea around an agent loop is you open up
cloud code and you say slash loop. Um, and then what ends up happening is during your whole interaction, the agent is constantly going back in looping to have a list of source instructions for the best possible way to give you that outcome. Um, which on the flip side, a lot of people are saying don't do it because the tokens that go into that are astronomically high. Like everybody's saying like unless you have a unlimited token budget or you work at anthropic or open AI and you have free reign to these models, don't do it because you're going to hit your limits quick. Um, but that's the whole idea of it, right? Like if we, if we go back and we think about how conditionals work, right? If this do that, if this error message print this out or perform this action. Um, so if you're thinking about loops, like that's the way I've been thinking about them. Well, it's just like a do-while loop. Yeah. Yeah. Yeah. Exactly. Literally. Yeah. Like it's just,
hey, while this thing is going, figure out the most efficient way to do this and keep looping through the steps until you get to the optimal result, which is great. Um, expensive, obviously, but great. Um, so we are getting to the point where efficiency is top of mind for everybody, better results, better quality. It's all there, but it's just a matter of the industry needs to, you can't put best practices around something that changes every 18 hours. All right. So that's really tough. So like there needs to come a point where it's like, okay, this is what stuff looks like. Now let's start putting best practices around it. And I do think we're getting there. Like, we understand what a framework is at this point. We understand what an agent runtime is at this point. We understand what type of prompt guards and guardrails and security practices and things that we want to observe or like we know all that. So I think that we can now start putting best
practices around it. And that's why people will, uh, well, big organizations, small organizations, everybody's trying to figure out what to buy and how to buy AI. Um, because we are now getting to the point where it's like, okay, yeah, this makes sense. We got to use it for this. You mentioned agent identity, uh, a while back. And I'm assuming that's, you know, a way to assign identities to the individual agents that are running in your cluster. Yeah. Is that an official project or is that just a general terminology that's being used for associating identities with agentic workloads? Yep. Yes. I think it's both. Yeah. So I think that there is a, let me see, agent identity, OSS, there is like an open standard for this. I forget, where I forget which one it is. I'm just, I'm, I'm looking now as I'm talking. I, I know that there is an open standard for this. Um, and I also know that like if you go, if you go and you look like
author, where I like they have their pieces around agent identity, even if if you go spin up a foundry project and you go and you create an agent and foundry, you will have an agent identity. Attached this. Just like go to the agent identity tab and you'll see it there. Um, so yeah, it is an open standard. And I would say at this point, it's like a, that, that's what, that's what I think we're calling it. That's what it seems like every vendor is kind of calling it at this point. Yeah. That makes sense. I know that, um, hash group announced something with vault. The newest version of vault is supporting agent identity and being able to rapidly provision those identities effectively. Yep. And I was like, Oh, that makes sense. That's, that's one of the products that would do that. Um, so I have you, when you're talking to these various clients are most of them running their agent workloads on Kubernetes. Is that the overwhelming platform of choice? I would say so unless you're using a, has like a platform like your AWS bedrock or your
Microsoft foundry or your Gemini enterprise, um, unless you're doing that. Yeah. I would say like the majority of people at this point are running in Kubernetes. Um, because I funny enough, I was just having the same conversation with a customer prior to, to jumping on this podcast where it was like, Oh, well, like, where else would they run? Right? Because like if you think about an agent, you have, it's essentially a client server architecture. You have your agent and the, for the runtime and the framework running on, let's say, some server. And then you have a client and you're interacting with it. So the server that's running the agent has to be exposed. And then you use open coder with some chatbot or whatever you want to use to go and interact with it. Now if you do it that way, you got to think about, how am I going to scale this thing? How am I going to observe this thing? How am I going to manage this thing? And then we get to the point where we got 50 servers for one agent because they need to be highly available versus if I pop it into Kubernetes and I turn on horizontal pod auto scaling, right? It did much easier. You know, you got your
self-healing from your controllers, from the replica set controller and such, right? So you can scale it up, scale it down, make sure that it actually stays up and running and operational. So I just think it's just the easy place to run agents at this point. That's fascinating about the workload of agents where they're running. And up until this point, we've been only talking about inference that is being farmed out to the providers, existing model providers. Have you been hearing anybody asking about running local models? Because if you're trying to save money and Claude Anthropic is costing you a million dollars a year, you might be looking at some hardware to install in your data center and run some models locally to save a little bit of money. Yeah, so I hear it a lot because I work with a lot of EU customers. So like data sovereignty and stuff like they need to run models locally anyways. But even in the US, right? Like it, it's beginning to make more and more sense to run these
models yourself versus going and spending hundreds of thousands of dollars or a million to dollars with a big provider. Because really, like if you think about it, if I were to buy 410 DGX sparks there, I think 4,500 or 4,600 at this point. And I'm running a couple of models, right? I got a Qwen 80 billion parameter model. I got a Gemma model. And then I put my harnesses around that I can even do some fine tuning if I want that agent to be or that model to be specific to what I needed to do. That's a lot cheaper and a lot more cost efficient in comparison to going and spending millions of dollars with these providers. Now on the flip side, what I will say is you could make the same argument for running your own servers. But the reason that organizations don't run their own servers is because it's much easier to log into Azure and click Create VM,
then to go to Dell by a server, provision it, put it somewhere, run it, and then connect to it. I think that we live in a world of, you know, we're humans, we're lazy, right? So like we don't want to go, we don't want to go into all that stuff. That sounds cumbersome. So I'm just going to go and I'm going to pay these people to manage it and host it for me. Maybe we'll get to that point where it's like, if you look at these data centers that are purpose built, like Colossus, which is the world's largest AI data center, and you look at data centers like that and maybe those continue to pop up more and more. Like maybe you go and you run space there and you put some boxes in or maybe it's like a rack space type of deal, right? Where you say, okay, I'm going to provision and I'm going to run it there and it might be cheaper at that point. So I think that there's a bunch of options, but I do also think that probably at the end of the day for these organizations, it's still so much, it saves a lot of time to just for your credit card to anthropic or opening a high
that it is to buy your own hardware. Yeah, for sure. I've been tinkering around with it and looking at what it would take to do this at any kind of scale. The tooling around it is pretty, still pretty rough around the edges. This is not going to be a drop-in and everything just works magically kind of situation. It's not going to be like installing V-Sphere today, which is like boom, you're done. This is a significant amount of work to just get the servers up with the hardware and the correct drivers and everything before you even provision a single model. And if you're looking at that, you've got to road ahead of you. 100% and it's funny because I wonder... So the reason why right now it's still a problem where like, hey, what am I going to do? Am I going to go pay a cloud provider? Am I going to go host my own
data center or buy a rack somewhere and put a hardware in? It's because it's less like consumer centric, right? The whole world of like Kubernetes and cloud native and all these things are B2B. There's Joe Schmo is going and buying a rack in a data center. Some company is. However, in the world of AI, it's B2B and it's B2C. So when we think about the advancements that can come from that mindset, it may actually end up shifting. A really good example of this is, I don't know if you all have seen the RTX supercomputers are calling them, right? Like, a Jensen did it in the keynote or whatever we was talking about. They're coming out with the laptops and the desks, that can run 100 billion parameter models for $4,000 or $4,500. And this is somebody, this is a personal computer, right? This is like your gaming, reg, or whatever, but because of the black well chips, because of how efficient they are,
you can also run these models if you want it to. So now we're like, we are hovering around the idea of like, yeah, it's B2B to run a model, but like also anybody in the world can do it. And with that mindset of AI is hitting both consumers and businesses. Maybe there's a point where we have a more efficient way to do this versus again, the market for running a server in a data center is much smaller than the AI market. There's more people involved, there's more innovations happening. So maybe we will get to the point where it's like, hey, we should just run this on somebody's laptop in the corner somewhere if we want to hit it for specific needs. Kyle, do you have anything you want to ask about? I don't think so. I think I've learned a lot and I want to go research. How do I set up AI gateways right away and how do I start saving some money? Because AI can be so useful, but it can also be so expensive. So we've got to mitigate that cost and hopefully not mitigate all the benefits
we're getting. That's the goal. Yes. You've given me a lot of links that have been searching as we've been talking. And I'm going to have to dig into some of those projects to learn a little bit more about how they could make me a little more efficient. And also maybe think more about like how I'm using the frontier models today. Michael, this has been a great conversation. Thank you so much for coming on and talking to us. We'll have to have you back in another year and you know what you'll be working on at that point. Sure. It'll be something interesting. If folks want to hear more from you, where should they go and find you on the internet? Yeah, I would say the best places are probably linked in and exit this point. Yeah. Those are still the two spots. All right, we'll include links in the show notes. Thank you again for being a guest today on day two DevOps. Appreciate you all having me. Thank you. Thanks to Michael for appearing on day two DevOps and virtual high five to you, dear listener, for tuning in. If you have suggestions for future shows,
we would love to hear about them. Hit either of us up on LinkedIn or send some feedback via packetpouchers.net slash follow up. You can find me, Ned Bellovance at net in the cloud.com and my amazing cohost Kyler Middleton blogging over it. Let's do DevOps.com and we're both entirely too active on LinkedIn. Stop by and say hello. Did you know that we also publish day two DevOps as a video podcast on YouTube? If you want to come see our beautiful faces, stop on by. You probably do know that already if you're watching and me say that right now. But if you're just listening, you can see the faces that go with these disembodied voices on the packet pushers YouTube channel. Swing by and hit that subscribe button. I think that's what the kids say. Pay two. Until next time, just remember that doing DevOps is awesome and so are you.
More episodes
More from The Fat Pipe - Most Popular Packet Pushers Pods

TNO076: Observability and Automation for AI Networking
The Fat Pipe - Most Popular Packet Pushers Pods

HN845: Real Life Use Cases for BGP Monitoring Protocol (BMP)
The Fat Pipe - Most Popular Packet Pushers Pods

LIU024: Dennis Thankachan: Lightyear and Beyond
The Fat Pipe - Most Popular Packet Pushers Pods

PP129: Inline Segmentation – HPE Networking’s Easy Button for Zero Trust (Sponso...
The Fat Pipe - Most Popular Packet Pushers Pods