Skip to content
TrackPodcasts
technologyMar 12, 202632:21

The Hidden Challenges of Running AI at Scale in Production

About this episode

Chen Goldberg, EVP of Engineering at CoreWeave, joins the podcast to discuss the critical infrastructure shifts required as companies move AI from pilot to production. 

Subscribe to the Gradient Flow Newsletter 📩  https://gradientflow.substack.com/

Subscribe: Apple · Spotify · Overcast · Pocket Casts · AntennaPod · Podcast Addict · Amazon ·  RSS.

Detailed show notes - with links to many references - can be found on The Data Exchange web site.

Get every episode summarized

Each time The Data Exchange with Ben Lorica publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

796 searchable segments. Every word is indexed and playable.

The Hidden Challenges of Running AI at Scale in Production

The Data Exchange with Ben Lorica

0:00
32:21

Full transcript

The Data Exchange with Ben LoricaThe Hidden Challenges of Running AI at Scale in Production. Machine-transcribed; use the interactive transcript above to jump the player to any line.

All right, so today we have a great guest, Chen Goldberg, SVP of engineering at Corrieve. Tagline is the essential cloud for AI, the force multiplier for AI, trusted by the world's leading AI pioneers. And she is an industry veteran, formerly a ran big engineering teams at Google. And so with that, Chen, welcome to the podcast. Thank you so much, Ben. I'm excited to be here. And maybe one correction before we kick off. My name is pronounced Chen. All right, let's see. So for some context, I'm sure a lot of our listeners already know this, but Corrieve is, as I said, cloud platform focus on AI. It does have some strategic partners, including Nvidia, which is both an investor and obviously a supplier of chips. And then open AI, which is also one of their larger customers. So with that out of the way, I guess my first question to you

is obviously you probably have seen the same surveys I have, which is, you know, a lot of people have AI pilots, but they don't go to production. But actually, I think a lot of those surveys are a bit overblown because I do talk to a lot of people. And they do have AI in production. So as best as you can tell, how would you describe the state of reality in terms of pilots in production? I think how quickly the industry moved from experimentation to production. And even this kind of surveys is really a testament of how AI is impacting the industry as a whole. It's just yet another example, right? We can talk about how engineers, I lead engineering teams at Corrieve, how our work has changed. Just these past three months with models. So what we are hearing from customers that has definitely changed and furthermore, it's not just moving from experimentation to production. There is a lot of still with that move. There's still a lot of unknowns and a lot of changes around.

How do I differentiate? What is my unique asset? If I'm an enterprise, though, I start up. So there is just a lot of things that folks have to think about, including which cloud provider they are partnering with. So I guess for our listeners to kind of frame your responses moving forward. So you are talking to people that are just not, you're just not talking to people inside the Silicon Valley bubble, right? You're talking to people in regular companies, right? So across different industries, is that correct? Correct. And again, that change also happened pretty quickly, I would say, from a level of interest. It's pretty similar, but probably on steroids. What we have seen when we, 10 years ago, also when we talked about cloud native and move to the cloud, where the evolution of SaaS happened and how we create experiences really impacted every industry. We're seeing the same thing right now. And what I see companies doing, definitely outside of Silicon Valley, is really looking inside and thinking,

what kind of assets we have, what kind of data do I have, what kind of people, what kind of customers do I have. And then thinking about how they take that and innovate further. And as you mentioned, the cloud computing journey for a lot of companies they've had years and years of experience, right? So either they're working with one of these hyperscalers or maybe they have their own cloud setup. But for AI, a lot of companies have a lot less experience, right? And they already have relationships with existing cloud platforms. So at what point does it make sense for a company to look at something like Corbiv, one of these AI first cloud platforms? So how much AI, how am I already doing before I even look at something like Corbiv? Or is that the right question, right? It is definitely the right question. And specialized tools, I think everybody should consider that. And it's company like ours. Yes, we are of course, specializing with our infrastructure,

but we are also building developer tools with our weights and biases platform. I think that like any other area that you want to differentiate in and to have an edge, you need to make sure that you get yourself the best support you can. And with that in mind, I think that everybody should explode the best of read approach and what can really help them accelerate, right? If you just do what everyone else is doing, then you won't have that edge. Assuming you need it, or you want it for your business, or it's critical. And specifically, when we think about our customers, when we talk about our customers for our cloud services, anyone that is either doing any type of training or using inference for real production workloads, which means that they care about security, reliability, and scale that can be idolatency of performance, those are kind of customers that talk to us. And what we are doing, I think which where we are unique, we are thinking about it as like a tiger team, like the experts in this space. So when someone comes to us, one thing that our customers are telling us,

that this is beyond that, you know, just I'm a vendor, give me something I'm going to pay you. It's really a partnership. In essence, and we take pride in that, which again, it's very much similar, in my opinion, to early days cloud, where you are looking to those trusted partners that can help you exert. And there is like a, for example, we'll just give an example of a startup company, they actually took significant of the money they raised and decided to use it to partner with us. And what when talking with them, they said like, hey, we don't want to waste our resources on things that you already sold, and we know you're the best. So that's the way I think about that. By the way, I want to give a shout out to another open source project that you folks acquired, Marimo. Yes, are you a friend of Marimo? Yes, yeah, I've had them on this podcast, actually. Oh awesome. And I'm a friend with Lucas as well, of Oets and Biases, great tools that you folks have rolled into your box. I think that's really speaks about that strategy of looking for the best tools for the job. And that's been our strategy and where we're focusing.

And definitely know with folks like Lucas and Sean from the Oets and Biases or ArcShade from Marimo, it's about innovating. It's about taking a different approach to this problem and where we think it really creates the right results. So you mentioned so the two things that people tend to do, right? So training and inference. So the stereotype is training that's mostly for Silicon Valley giants and labs. And you can probably give us examples of otherwise. So are there companies that are actually doing training that are not one of these tech companies? Yes, I will give a couple of examples. Definitely anyone on the financial world when we think about risk. As an example, it's something that we see customers using training with their own data already. Really from us. So this is like a foundation more. Building their own models for and they've been doing a, by the way, ML before. Yeah, there's there's nothing new about that. So it's not necessarily doing their own labs, but just the work of training models. And on we are also seeing on the retail side companies investing in unique experiences

there is an assumption that, you know, the way we as consumers consume research data will see that definitely, you know, early on drug discovery and research around health. And we've seen some of those customers definitely running things at scale at this space. And there is some other things that are very common or very common. Maybe it's a Silicon Valley. Maybe I shouldn't say very common. Some other use cases are on call center, services, a video generation, which is again, other use cases. It's not just gaming. Did your generation, that's interesting because it's something that the foundation model builders can do, right? So Gemini, I can do it, opening, I can do it. But for some reason, there's still an opening for startups that really specialize in this, right? Yeah, I think that for everything, it's probably too early to talk about disruption in this space. Because maybe in parallel to what we are seeing around the strength of foundation models, we also see evolution of tools. Right, how easy it is to maybe train

or do an inference at scale, or use your own data. Fine tuning has become so easy now, right? And we're just at the beginning, right? If you think about like where we were six months ago, so much as evolved, actually that's another, we've acquired another company last year that's called OpenPipe. And what they're doing, they are making reinforcement learning easier. Right, the idea is that they make it simpler for you to do that trade-off between best chip and fast, right? It's always that trade-off. But can they help on some of those metrics make things more automated and like a good enough approach? And again, it depends on what task we're trying to solve for. So, too early, there's a new model coming up every day. Yeah. It's hard to keep up. Almost literally. Yeah, not off, yes, sometimes too. No, just good, yes. And I'm going to go back to reinforcement learning in a little bit here. But so on the inference side, what's the biggest, I guess, rude awakening that people have when they move to a high production kind of application? So at first, of course, they go, we can just maybe

use a traditional cloud or do this ourselves. But then at some point, what causes them to move to something like Corvieve? So maybe, I think that's a great opportunity to talk about maybe the origin story or why I joined Corvieve. Yeah. So for folks, maybe they don't, I'm sure many don't know. I spend almost a decade at Google Cloud. Kubernetes. Yes. Part of the founding team of Kubernetes. And really, what Kubernetes did in the space of cloud is create an abstraction there, making cloud resources feel the same everywhere. And actually creating the opportunity for potability of workloads and multi-cloud becoming a reality. And part of that world was on the fundamental assumption that all resources are the same. And if I'm a user, I can declare what kind of resources I don't need to go to too much details. And I can let this orchestrator deal with things on my behalf. Which for our younger listeners was a pain in the you know

where before Kubernetes. It was really hard. It was really, really hard. And there was no need for it to be hard. And then what happened with AI workloads is that some of those assumptions that we had in the past were no longer true. For example, the idea that I can easily move workloads from one node to another for one machine to another. That's actually no longer true because actually most workloads, especially when you run things at scale, are running on multiple nodes on multiple machines. And that just that makes that orchestration and how you think about highly available systems and reliability different. The cost and complexity of systems became really hard, right? Once you understand that you're moving from a single node to multi node environment and you add on top of that, the cost and the complexity of the systems, then the risk of having one point underperforming and just dragging down the entire system has huge consequences. Huge consequences. It's like it's for some customers and companies

10 of millions of dollars. And also it's an advantage from a company if I cannot move fast enough or if my service is down, it will impact my business period. Yeah, user experience is so critical. Correct. And that's where companies like Koryv came about. And I joined Koryv because I saw some of those challenges while I was working at Google and two things. One, not being part of this infrastructure replatforming felt to me like a missed opportunities and infrastructure person. It's a great time. And I realized that in order to meet the demand of those workloads, we need to do things differently. Okay, we have the opportunity to do things differently. And that's what Koryv is about. Okay, we have the privilege and luxury to really focus on a smaller set of problems. We can build simpler solutions for that. And we're doing it across the stack. And maybe that would be the other point that I would, the other thing that we've learned is that it's not one thing. Okay, it's not just like your GPU or as like it's, it's always a combination of things of how you get that sustainable differentiator

and that's a sustainable advantage. And that's what we do. We optimize different things across the stack. One of the things that I came across that I found interesting is this thing that you call arena. Yes. Which is like almost like a digital twin for people running large AI workloads, right? So it's a simulation environment, right? Actually, what's unique about it? It's not a simulation environment. Oh, it's real infrastructure, real production. It's a real cloud. What we have to say is it like a here's my workload. And then you have like an exact copy and then you simulate that. Yes, some of the decisions that people are asked to make right now, when you and you've asked Ben before about inference is making decisions around the infrastructure. Because if we said before like with the Kubernetes, if I'm running it on H100 or H200 or GP200, okay, the system has to be different, the architecture, the deployment, things will change and will perform differently between those systems. And people are asked to make big decisions

where the stakes are really high without having the data. And those benchmarks that are available, like you know, test this model, test this, are actually not a good mirror to what will happen in real time or with real production workload. And what we are offering is an experience that right now we know is funded by all of our customers is the opportunity to partner with us with your real workloads and re-bring the expertise of real infrastructure. And it gives you the opportunity to benchmark what kind of environment, what kind of storage solution should I use, what kind of networking, what kind of data I have. Let's do troubleshooting. How do I do highly available? There's a lot of questions that we are helping customers accelerate the decision making instead of guessing and not getting into those long-term commitments without having any answer, which I think is what the industry is expecting people to do. And it's not reasonable, really. So basically allows me to stress tests beyond what GPUs do I have. I'm trying to stress the system and learn. We all talk about experimentation.

So then when you, as you were describing it, I imagined like almost like many, many dimensions, almost like a multi-arm band that I have to optimize to configure it. Exactly. Exactly, I think that's exactly right. Your spot on, it's not one dimension. It's not just which model I'm using, but it's also how I'm deploying it, how the network is configured. What kind of data am I measuring? Where do I getting my data into? There's a lot of things I would like to consider. And you know, we talk a lot about performance and reliability. There is another aspect that is really critical to our customers, which is security, which also relates to of course, instrumentation and observability. So having a system where you can partner with experts on all of those dimensions and make some quick decisions is key. And there, you've also introduced this notion of efficiency, which I don't know if I've heard this term before, good put. Actually, it's helpful for my Google as well. OK, so it's the actual time GPUs

spend performing useful work. So tell us why that's something that people should start tracking. People are already tracking a lot around efficiency in the sense that because compute is so in demand and very expensive, people are looking for solutions to get the best out of their infrastructure. And maybe by the way, this is I think another big difference between general purpose cloud and AI cloud. And what we have learned and we've seen that in, you know, we are submitting our results to the different MLPF and benchmarks results. And we've seen feedback from our customers that there are different type of bottlenecks along the way. So for example, one of the things that we have done in order to improve good put is making sure that we accelerate the amount of data both from volume and throughput that gets into the GPU. Right, so anything we can do in order to make that processing more effective will be threats that we build the specific caching mechanism on top of our core with AI object storage that accelerates that.

So that would be one example. Another thing that we've been doing is giving customer transparency into the system. So that means that when a job is not maybe performing well, I can quicker know what's the root cause, right? Because maybe my job is underperforming. I want to know why I do I need to restart, I do I need to fix something is something that will be resolved on its own. Those kind of data when the stack is very complicated is really hard to learn. So what we've been doing, it's solution that we call mission control. And we are thinking about the stack as one. So no longer thinking about storage separately, network separately, workloads, orchestration and so on. But how do we bring all the data together and make informed decisions? So that's another way our customers are able to improve a good put and of course from reliability perspective. And we have our own proprietary IP that talks about how we test and validate the infrastructure both proactive and reactive. And those kind of things are all accumulating together

to improve efficiency. By the way, one of the interesting things about IT ops and DevOps is the telemetry and the amount of data that's being generated. In fact, it was always the laboratory for all these massive time series analysis tools and databases. So I imagine you folks are innovating around, it's not just collecting all this telemetry but making sense of it. So you must be using AI to diagnose AI, right? Of course. Yes. So anything to share there? I think what's definitely interesting is leveraging the amount of data that we have to accelerate troubleshooting even further, which just helps us as a team to be more effective and watching trends. But I imagine that's not unique for us. So that's just us using AI with the data that is an asset that we have as a platform running at scale. So are you, but you must have a team of people who are building AI tools for time series, massive time series analysis tools

or something like that, right? Yes. And every day that goes by, there are more and more people to do that because maybe going back to the discussion we had at the beginning when you said, like, hey, do you use AI at scale? Of course, like every other company we're also consumers of other AI technologies and we are leveraging that significantly. So obviously a lot of listeners have heard about GPUs and GPU shortages, so on. But the other thing that's been in the headlines recently is memory. So anything to report on DRAM, DRAM and memory constraints, DRAM supply, rising prices? I'm not an expert from a supply chain perspective. But even your high level is a lot more than what we know. So that's, I think fundamentally the demand for AI infrastructure continues to rise and it has implications on the entire supply chain. And we're seeing it across the industry and that's of course on one of the course a great opportunity, not just of course

for building more capacity, but also finding innovative ways to make the infrastructure more efficient and those are the kind of two things that we are optimizing for. So internally within the team, of course, we are working with our vendors and working on improving the supply chain. And if anyone is in that space and actually everybody are familiar with that challenge right now, but in addition to that, we are also looking of how we can help our customers optimize the confining solutions. I think that's like the end of the days. How do we partner with our customers? Because memory is a critical component as well, right? Because it's the model sizes, right? So. And it's not just memory, but yes, for almost everything in the supply chain is impacted at the moment. So you brought up the topic of reinforcement learning earlier. I'm a long time reinforcement learning fan, cheerleader in some ways, but I've always, it's always been one of these things that's always just around the corner beyond the grasp of regular teams, right? So always the province of advanced teams.

So you seem to be hinting that it's getting closer to being much more accessible and usable. You're seeing more people at least trying it out. We're seeing more people trying it out. And again, I think that this is where folks are moving to more of those production use cases and donate to make that trade off. So just to clarify, we're not talking Silicon Valley bubble companies here. So you're seeing RL beyond the useful tech suspects is that what you're hinting at. Tech companies, not just startups. Not just startups, okay? Not just startups, but definitely tech companies that are investing in AI right now in order to innovate, differentiate are the kind of companies that are definitely investing in this space. So for reinforcement learning, maybe you need more CPUs as well, right? Yes, of course. So then, so I said part of what you folks are trying to understand, okay, so if reinforcement learning grows, we need to add CPU capacity, is that?

Look, I think that some folks, like when they think about co-weave, they just think about like GPUs in large clusters. That's not the case. The way people... That's the stereotype. Yes, the way people should think about it is that co-weave as the AI hyper-scaler gives you all the tools and infrastructure you need in order to run your AI work, period. That means, okay, that of course, we have file systems and object storage and CPUs and we have a cloud console and API and Terraform and security guarantees and SLAs and SLOs and we're doing that is, sorry. Vector databases. We are helping customers manage their own. We don't yet have a managed service over vector database, but the way someone will consume is, is they will find a lot of things that are familiar for other clouds. You know, we've launched our serverless RL offering, so there is, of course, we have an inference service, so just think about the set of tools and services, and you yourself mentioned, like Mari Moro

and what's it's advices. So that's our approach. So it's a simpler stack, highly optimized and of course, if there is a needs for CPU and other things, then of course, it's available as part of our offering. Do you for sub-ray? I'm an advisor to any scale. We have customers using Ray on co-weave as well, and if you look at our documentation, there is some help charts and examples of how you can run Ray on CKS, co-weave Kubernetes service. Awesome. So as I mentioned, right at the beginning, right? So Nvidia and CoreView are kind of joined at the hip, but I don't know how much you can answer this question, but there's also other alternatives in video, right? So other GPUs, custom silicon, so that's not something you folks pay attention to at all, plus you're all in an Nvidia, or do you monitor the rise of alternatives? We monitor the entire industry all the time, and we partner with different vendors in the industry, as long as it's what we are hearing from our customers. Okay, and that's really for us the North Star,

and that's where the partnership with Nvidia has been so great, you know, we are working as a partner and as a customer of course, and it's important to mention that our differentiation goes beyond the GPUs of course, right? Because we talked about before the multiple dimensions, and we're investing across the stack and there is a place for solution innovation, but at the moment, we don't see demand or a need to diversify beyond working with Nvidia, and Nvidia already has a portfolio of accelerator, which we make of it. And obviously one of Nvidia's main secrets of course, is their software stack is so much better than much more mature and much bigger ecosystem than the other hardware platforms, right? Yeah, so let's see, I wanted to ask you about the other bus word, which is agents. Okay. Obviously agents, the computational patterns is slightly different than just straight up inference. So what changes for a customer that's doing a lot of agents

and what does that mean for their infrastructure? First of all, like before, like you mentioned, definitely when you're running agents at scale, there's different requirements from a computer GPU, the CPU for two, like so there's different. Yeah, and then agents also have, they might do reasoning, but they might also do these loops, they're calling tools, they're calling other agents, right? Exactly. When we are talking with customers, there are a multiple things that are very important. One, of course, is where they have the data and data gravity. By what agents also create, and we start seeing more of more people talking about it. When we talk about agents and data, there's a lot of security considerations that people are thinking about. There's a lot of innovation in this space, we see, for example, a need for sandbox, as part of that, right? Like how do I run, maybe I'm interested, a code, as part of those, part of those systems, and definitely the scale. So those are kind of things that we are seeing customers too. They're also, of course, folks that are using existing, like, of the shelf system

that they can use for agents, and I expect we'll see more of that. So we are also internally, like we are using, we are customers of other systems in some areas. Your sense is people are really using agents, it's real, and what percentage of people are using of these workloads are now agentic, you would say. It's impossible to answer, but that's impossible. Non-trivial. Only, probably, like, my best insight is from thinking about our product and engineering team, right, so where I see my own team. And I imagine that's true in an insight, but I'm not an expert into other areas. We're having conversations when you think about agentic coding, for example, the quality, how does it change? Code contribution, how do you do? Code review, the size of those that kind of conversations from quality, how much you need human in the loop, when the expertise that humans need to demonstrate in order to be effective. So those are kind of conversations that really evolves every day. So the best models are still the proprietary close models,

but the open weights models are starting to get better. They're cheaper or they're faster, they're cheaper to run, you can customize them, but the downside is they're from China, in some enterprises, some industries, that's kind of a no-fly zone. So what's your sense of adoption of these Chinese open weights models? I don't have the data, but maybe, you know, the key point that you said is, when you said open source models become better, and then you talked about those dimensions. So at the end, it's kind of a trade-off, right? And it depends on the task, and where RL, for example, can be useful, is by making those models better, and maybe cheaper and more accurate, why do we say better? And those the kind of things that we are looking into, it's probably too early to know exactly what people are, is because that just keeps changing very frequently. So in closing, we're still in the early days of AI. Early days. But people should probably start thinking ahead and worry about technical debt even at this point. So what kind of technical debt are you seeing people incur

and what should they be avoiding at this point? So first of all, yes, it's early days for AI, but it's moving really fast. And I think that that's something that if people are not experimenting heavily right now, they're probably behind. Even for your own personal productivity, for your products, for your teams. So I think that would be the first thing. The second thing is understanding the implications of that the market is still constrained. We talked about supply chain, we talked about capacity, talent will be another. So folks need to make some strategic decision with some unknowns. I know some companies will feel uncomfortable with that, but I think that's part of experimenting with a new technology. And when you think about technical debt, I think that will have been around for enough decades to know that it's never like just everything changes, right? We will always stay with some legacy. So I think the opportunity is how to take these new tools and also apply it to maybe the less sexy problems,

which can probably be very useful, like some of the things that I've mentioned is around troubleshooting and engineering excellence and productivity. I think that's definitely where we see value already today. Actually, one thing that people don't realize is that, so obviously there's a lot of talk about AI for programming and writing code, but there's also a lot of AI and agents being used in data engineering pipelines, ops, DevOps. So there's a lot of things already happening that the consumer may not be realizing is being powered by AI. And that's on the technical side. And I just, in the last few months, I'm getting more and more scared for knowledge workers. Yeah, it's just, it's really, this technology is going to be very disruptive. As she mentioned, the owners is really on you to really use these tools, to upscale, to constantly learn. Yes, one of my other passions is mentoring, again, part of me being in the industry for so long.

And if there is one thing that hasn't changed with what I tell folks, you know, whether they're early on in their career or later on is the importance of being experts. I don't see the tools as a way to avoid depth or expertise or knowing what you do. So maybe that would be the other thing that I think people should think about, especially when there is a new technology. Like, what are the things that I know best where I can apply my tools? I was just talking with a friend, for example, from she's in a marketing domain. And she said that she feels like she has those superpowers now. Yeah. She has so much experience and she can do things she couldn't do before with much less time. And that's the way I think people should think about it. But that means though that there might be fewer. You might end up needing fewer. Cause one thing that seems like the stats are bearing out is the entry-level jobs for recent college grants. That's really soft at the moment. So there's a bit of a hollowing out of that pipeline, right?

So you need those entry-level people to become mid to become senior inside your company, right? So if you kind of slow down the hiring of that entry level, you lose that job, right? I can't speak about correlation. I do believe. And that we had those kind of thoughts in the past with different technology. Changes the advice I would give to junior folks early on in their career. The technology is actually accessible. There's all the things that you can start and do and build the expertise. And that's something, maybe I will just bring it back to Kubernetes. One of the things that I loved when we just started is that there was no one with 10 years experience with Kubernetes like that or with Cloud Native. It was all new. This is the time. Okay, there is this new technology. And that's the opportunity for folks to lean in. And with that, thanks for joining us. Thank you for having me. Thank you for having me. You're going to follow the work of Hen Goldberg online at pourweave.com. Please join the thousands of people

who subscribe to our newsletter, which you can find at gradientflow.subset.com. And we are reader and listener supported. So if you can, do become a paid subscriber. Thanks for joining us. If you like the show, please subscribe and rate us through Apple podcasts or overcasts or tune in.com or Spotify and never miss an episode. The Data Exchange podcast is a property of gradientflow. And I'll be back next week. And we will do this all over again. Thank you for joining us. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.

More episodes

More from The Data Exchange with Ben Lorica

View all episodes →