
1032: Agents Need 10x More Data Than Humans, with Salesforce’s CDO Michael Andrew
Get every episode summarized
Each time Super Data Science: ML & AI Podcast with Jon Krohn publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Four years ago, Salesforce had their data fragmented over 650 different data streams, making it impossible for even Salesforce to have a unified picture of who their customers are. But my guest today came up with the solution.”From the transcript
Get every episode summarized
Each time Super Data Science: ML & AI Podcast with Jon Krohn publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
292 searchable segments. Every word is indexed and playable.
Full transcript
Super Data Science: ML & AI Podcast with Jon Krohn — 1032: Agents Need 10x More Data Than Humans, with Salesforce’s CDO Michael Andrew. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Four years ago, Salesforce had their data fragmented over 650 different data streams, making it impossible for even Salesforce to have a unified picture of who their customers are. But my guest today came up with the solution. Welcome to another episode of the Super Data Science podcast. I'm your host, John Crone. In today's episode, we are honored to have Michael Andrew, who is Chief Data Officer at Salesforce, and we're recording live from the Dream Force conference in San Francisco. In this episode, Michael provides tons of brilliant advice on how to have your organization prepare its data so that you can be effective in the AI era and particularly in the agentic AI era. Enjoy this one. Michael, welcome to the Super Data Science podcast. Great to have you here. How's it going today? Fantastic. Excited to be here with you. Yeah, we're day two of Dream Force.
It's been my first dream force. Okay. I've been blown away. All right. What is, what is stood out to you? I'm really curious as a veteran, but I don't have the eyes of a newbie. What is Dream Force like to you so far? Well, the most emotional moment for me was I'm a huge fan of no doubts tragic kingdom album. Okay. So when Gwen Stefani came out unexpectedly before Mark's keynote yesterday. Yeah. And she played, don't speak. I literally burst into tears. Oh my gosh. You can thank our events team. They do a lot of work to help mark if the right people to show up. And of course, if you're a fan, Gwen's playing tonight. I know for sure. I'll be going to that as well. Too hot Metallica is one of my favorite bands. I took out I didn't get invited. They've been a classic, but I think that some of the attendees will have to switch it up. So we got I think usher and Gwen Stefani tonight. Yeah. Yeah. It makes sense. All right. Let's get into the technical stuff here. So your career has been about listening to customers at scale through data. And you've been doing that for a long time now since the 90s. Yes. And for eight years now.
Eight years to sell. I'm going to say years of sales force. Yeah. And Chief Data Officer for how many years now? A little or two. Yeah. And so you're about the best expert around on enterprise data that there is. Now that humans are not the only people working with data to clean insights or do work with data. Now that we're joined by agents, how does that change what it means to be a CDO? I think there's two sides to it. There's we now have customers that are the agents. So for most of my career, it's been how do we get the right metrics for this different executive and their teams? How do we align the right data to help different departments? But now the workers of this department are also agents. And we have agents. We have customer support agents. Every day, they're working to help a customer. It's not like they could learn who the customers are if we don't get them data. We have to say who is this customer? Here's our help document. So more and more of
my team's job is actually preparing the data for the brains of the agents, not just the brains of the humans. So in many ways, we're seeing more demand than ever before. But also, I'm also employing agents to help me. So even though it's Salesforce, we have an amazing team. I have amazing resources of very large scale global organization. But even as big a company as ours, I could never hire enough data scientists. I could never hire enough data engineers. I could never hire enough analysts to serve every country, every market, every one of our clouds. We always had some invitation. What we're seeing now is agents are letting us do a lot more. One great data engineer with five agents, you know, our mental agent's helping is now giving the output of what you should take a team of five or 10. And that's helping us in so many ways to just do more and start to do better work to kind of spend more time on the things you want to, which is the
strategic part. So now I've got to employ the agents and I've got to, are my customers as well. I think something that seems like a little bit of a paradox that people might not have expected is, you know, if you have a data engineering agent and it performs the job extremely well and obviously it's going to get better every month that passes theoretically. That could mean, oh, a data engineer loses a job. But it seems like in most organizations, the inverse happens because, as you said, now a data engineer, a human data engineer armed with a team of data engineering agents is more effective than ever before. And so actually that human data engineer is now better value for money than they've ever been. I think so. And it's, of course, a change is a skill because you have to start to think how are you like a manager going to be managing this process, right? So maybe in the past, you wrote the great individual code, right? You wrote the SQL, you wrote the Python, you kind of did that work. Now you're overseeing the agent that's writing the Python. But again,
most businesses are not paying you to write SQL. They're paying you to, hey, I need this metric. I need this pipeline to run at a certain time and think about like tickets. Every data engineer, do they really love tickets when pipelines break? That might be a lot of your week. But now we have an agent that can monitor the tickets. In many cases, the agent can just fix it. Well, that's amazing because you actually want to spend time on that next pipeline, that next algorithm. And like I said, it sells for us. I just look at a backlog that even with these agents, I don't know how I'm going to fulfill that, you know, there's so many parts of the business that need data. And what's changing is agents need a lot more data than humans. So I'm almost at like, I need to produce 10 times the volume of trusted data than I did before agents. And so all of a sudden, our workload is going way up. And so yeah, I'm still hiring people all over the world. They may be slightly different skills. But we have more work than ever before. And there's really no way we could get
it done if we didn't have agents to help us. Why do agents need 10 times more data than humans? If you think about it, humans are very good at learning from the world. Our neural networks already figure things out. But agents are like newborn children. They haven't learned anything from the world. They learn through data, right? The models have been trained on the internet, but they don't know your business. But the only way to speak to an agent is to speak in data, to give it context. Right. So if you think about it, let's say you made a tablet dashboard and you had a metric and we have a metric like pipeline. Anyone in sales knows what pipeline is. But is an agent going to know what pipeline means for your business? Unless you've told them, here's what it means, or what the ACV metric means, or any metric. No, you have to say, well, here's what this metric is. Here's how it's used. Here's what it means. So you think that simple metric, you now have to have all these other data points. And here's how it's related to this. Before the humans understood it, you could put it in a slide, you could put it in a dashboard. Agents don't can't do that. They
don't know your business. So you have to start to create data for them to consume. Yeah, I guess a lot of the humans that you would hire would, there's an expectation in a lot of cases that even if they don't know your business that well, you probably hired them because they have the familiarity with a similar kind of business, but the agent doesn't come without saying background. And I've hired people obviously and marketing to do marketing data science, people in products or revenue. Usually they have to domain expertise, but also you train them. They learn, it takes time, but agents aren't going to learn if you don't give them the data and the information to teach them your business. So that means something is a data team. Again, we had a slide that literally showed every team pointing to the data team for dependencies. So my whole conversation with my team now, how are we going to scale up? Because now every single department in the company and Salesforce is using an agent. Every agent needs data. How are we going to keep up? And of course, the the veracity of the data is hugely important for agents because as probably
all of my listeners have experienced and new and I have experienced agents, they're a weakness for them still today. And most cases is that they don't come back and say, I'm not sure I have enough information on this. They'll typically just kind of go with what they have. Yes. And so yeah, so that's so having the the data right is something critical. And so Salesforce has described its own customer data four years ago as chaos. Multiple CRM instances plus snowflake Google and Amazon with no cohesive real-time picture. What was that situation like? I mean, the things you would expect, right? Having duplicate customer records, having lots of different data systems. And by the way, we keep buying companies. They come with our own data systems. So this is actually a big reason why we developed a product now called Data360. And we run one of the largest now Data360 deployments in the world. If we were a customer, we would be in the top five customers globally. Obviously, we don't pay ourselves exactly
to use it. And we do a lot of R&D internally where we work on our products. But that's how we solved it is we connected our snowflake. We connected our Amazon data lake. We connected our multiple different Salesforce instances. And then we were able to unify that customer picture. And now, whether you're in our sales team, our marketing team, our service team, we're able to get that information. And so we talk about this is it led us kind of untrap data. So we had data and backend systems like our licensing system. But obviously for a salesperson, they didn't know, hey, this customer is running out of licenses or they've used data. That was a backend IT thing. We were able to connect that up, merge it. And now our sales team gets an alert. Where a data cloud is listening and says, hey, this customer is about to run out. Maybe you should talk about getting the more credits. So that's really where we've made tremendous progress. But we did it really by hooking all of the data together, integrating it. And in our system, you don't have to move the data. So we can read the data from all the systems, merge it into one harmonized
view, and then supply that to all the applications and agents that need to use it. It must have been a huge amount of work with all the fragmentation. Some of the stats that I have here are that there were 266 million fragmented customer profiles from over 650 data streams. I think that sounds like the biggest headache of all. And over four years, you and your team resolve them into 141 million unique individuals. So basically having the number of customer profiles because you were like, okay, duplicate entries from across these different 650 data streams. And so yeah, I know you understand that you call this kind of your truth profile. Yes. In terms of you being a user of data 360, I understand that it's a salesforce principle to be customer zero on most of the products, right? Yeah. And customer zero to us means that we're going to be the first to put our own software in production at scale because if we can make it work, or what's about to be $50 billion a year company operating at hundreds of territories around the world, lots of businesses,
we think it'll work for others. And that means sometimes we don't get it right. But then I look at it as like, well, our problem is an opportunity for you. So if we mess up, well, that's kind of on us and we're learning, but we want to then take that and make it more resilient. And that's really our role. And so my team interestingly sits in the product organization. So we run our internal technology inside the product organization. And essentially we had to both run it for the business again. We're going to be a $50 billion a year business. Our data has to be right. It has to help all the different now 85,000 employees around the world be successful with their jobs. But we're also R&D, which means we will try new features. Sometimes we build new features if they work, then there's some features that we want to bring to all of our customers. But that kind of dual role is what customers here. I mean, so that we're always testing. And by the way, we buy other software. Sometimes we don't have the software. So just like our customers and snowflake as a partner, data verix as a partner, we have these different data systems that we also use, but we use them with
data 360, just like so many of our customers. So we're able to help our product teams see what really happens when you put this in production. Because that's really where kind of the rubber meets the road is when you run real production workloads on these systems. Yeah, most enterprises would already have a warehouse or a lake house. What does a unified real-time profile in data 360 do that a snowflake or data bricks doesn't? And how do the two co-exist? Yeah. So if you think of any of your kind of data warehouses or lakehouses, they're essentially like the repository. Here's where, you know, what do you call it a data lake, data warehouse. You're often storing a lot of data there, but you're not necessarily putting it to use for your sales teams, for your marketers, your customers, right? To do that, you need to get it into an email system. You need to get it into the call center system. You need to get it into the sale system. So that's really what data 360 does. It lets you activate your data. We call it like your untrapping data. You have all
this data, but the reality is your salespeople aren't writing SQL, right? They're not worrying about the tables and the storage, the head up jobs, the spark jobs. And so if you think in so many companies, it's actually where a data team can get in trouble because they've kind of become the back office. It's not really where the value is. The value is how do you help the business create revenue, resolve customer issues, expand partnerships. So data 360 is that way wherever you store your data, and again, we have a wide zero copy network to be able to tap into it and then make it active and available. So we use it to drive our marketing campaigns. We send hundreds of millions of emails and messages every year. All of our salespeople get it. All of our customer service. We use it inside of our product. So that's all by data 360 taking what had been. And again, we have hundreds of petabytes of data on one data lake. We have many benefits and we have data everywhere. But it was all kind of backend systems used mostly by engineers and IT. Now that data is available
in the flow of work for the rest of the employees in the company and our partners. Regular listeners will already be aware that I'm obsessed with Anthropics Fable 5 model. And it has taken over my working life. I'm writing a technical book that includes latex files, mathematical notation, Python code examples, and Fable 5 and Cloud Code handles requests I make across whole chapters with accompanying Jupyter notebooks and to end work. But a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works. Essentially like magic first time. Cloud is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at Claude.ai slash super data. That's Claude.ai slash super data. And check out Claude Pro, which includes access
to all of the features mentioned in today's episode. Claude.ai slash super data. Brilliant. Yeah, which is the key, as you say, to unlocking those data, making them usable across the organization, whatever department that they're in. Something that I found interesting is talking about this kind of coexistence between a data warehouse and data 360. Data 360 now shares files into Databricks Unity catalog. So what this means is that there's no more copying of data for listeners who aren't aware of how that Unity catalog works. And Salesforce talks about zero copy architecture generally is the era of moving data over. I think you don't want to move data unless you have a need to. So for many use cases, you're simply reading the data and then you're using at that point. So moving the data would be an expense that you don't really need to do. There's always going to be some cases where you need to have that data in a transactional system. And again,
there can be cases where you want to do that. But for, call it 90% of the use cases, you're more wanting to read the data wherever it lives and then put it to use. So this zero copy network is that way where if you've already made your investments in your warehouse, you already have your integrations, you don't need to redo all that. You just want to tap into it and put it to work. And that's why us and the industry has really rallied upon that so that companies can make that choice because it's expensive. By the way, I've done big database migrations. I have paid millions at dollars to move data from one warehouse to another. That's a lot of expense. Kind of get the same thing. So you don't want to spend that money just moving data around, re-hooking everything up. You want to spend that money on doing something valuable with the data wherever possible. Yeah, great guidance there. Kind of going back to our core theme of how important it is to have the right data for AI systems to work effectively. You've said the transformation happens when data and AI move in lock step in an organization. Operationaly, what does that mean for people
planning their AI strategy? So we talk about the authentic enterprise, which is that every company, if not now, soon will be a company that has the humans that work there and have the agents that work. So what we're finding and we now have agents all through the business. We have agents helping our HR team. We have agents helping our sales teams, our marketing teams, our product teams, you know, our security teams, everybody. So as you go through this transformation, you have to say, well, what are these agents going to need? What data do I have today? And sometimes the data we had for the humans is not good enough for the agents. And so that's why they have to be in sync. And so we spent a lot of time and I'm part of the broader technology organization. We're saying, what are we transforming? What use cases are we trying to light up? What do we want these agents to do? And then we look and say, what data do we need for that to be effective? And I think
I was saying earlier in the talk, a lot of where I'm hiring teams is to now build the data for these agents, right? For these autonomous workflows. And that's new data. And you know, if you'd asked me three to four years ago about unstructured data, think all the documents on the notes on the Slack, I got kind of interesting, but I was a lot more focused on your metrics, your data science, your prediction, your forecasting, all the kind of classic data science. Suddenly, I'm a lot more worried about documents. I'm a lot like, well, how do we get the agent to understand a document? There's lots of ways we have the 360. How you read it, how you summarize it, how you give the right tokens to the agents, how it works for the models. This has a whole new class of work, but it kind of matters, right? If you say, well, let's say you have a sales play, the agent has to know what is a sales play. Has to read the sales play. Well, maybe the sales play today are in a Google slide. Well, if you give it a Google slide, can it understand in the same way that our actual sales team does? Well, probably not without a little bit of magic. So that's a whole new class
of work that is now changing where we're prioritizing. And frankly, we're having to scale up to meet the demand we're seeing in our own company. And I expect every company is going to see this as they also transform. It's interesting going back to the beginning of the interview. You talked about how you could never hire enough data people, data engineers, data analysts, data scientists. And the agents kind of served as a potential solution to that problem where you can know, okay, now we can multiply people's impact. But it seems like also at the same time with the line that you're kind of just telling there, it seems like all this agentic stuff is also creating more opportunity and more need than ever before. More need than ever before. And frankly, they need to create a lot of new data. Because if you think about, and by the way, one of the bigger kind of things I've started to realize lately is that actually what we're producing is not data, it's intelligence, right? That the raw data itself is not what's useful. It's after you've
processed it and you've structured it in a way. We might talk about trusted contacts, but I think of it is like, what is the intelligence of your business, right? The way you understand a customer, the way you go to market, all of that is actually intelligence that the humans in your business have built over time. But the moment it kind of lives in their brains, at some point, they have to become data. And these agents, as they do things, they're generating a lot of data. So our customers that are adopting agents pretty early, but we're already seeing six times that you said to Salesforce and they did with just the humans. So now they're writing more transactions, they're doing more things. They're actually making the data in Salesforce better and they're creating new data. Well, what did the agents say? What did it know? And one of the things that I think is for people out there, you can now actually measure the value of the quality of your data. You can see if I give a bad piece of information to the agent, how well does it answer?
If I give my wonderful, clean, happy process data or intelligence to the agent, oh my gosh, it answers better. It does it more efficiently. It costs less. So we're realizing you can now quantify the ROI of the quality of the data in a way you never could before. Yeah, that's got to be the theme of this episode for sure. The data are invaluable for having your AI systems work for my technical listeners, data scientists, AI engineers, data analysts, what have you? The consumers of their work are increasingly agents rather than people. What should they be building or learning right now in order to be better prepared for that world? I think one of the most important things to do is just use AI a lot, use different models, learn how the models work. I start to test, start to test, sending different contexts with different data to the model for the same prompt, learn about eVALs, evaluations. How do you write a good eVAL? Because agents are never going to answer exactly the same way twice, but you can build
ways that you look and measure it for sure. I found, personally, the more I've worked with them, and then I take the same piece of data, I try it against multiple models, I try different variations of that. It starts to build your intuition. Then of course, at scale, you're building testing, you're automating it, just like you would run an eVTest or multi-variate test, and you had to learn things of like, well, what's my sampling strategy? What's my control group? Now you have to do the same thing, but you're able to test these different agents and literally see how they think differently with the same piece of information and how do you need to change that information to get them to work more effectively. Great advice. Thank you for providing that. That is the end of our technical questions for the episode, but I always end my episodes with the same two questions. So the penultimate one, the second last one is, do you have a book recommendation for us? And this doesn't need to be something technical. It can be anything that you like. Okay.
So I think we're in a unique moment in history, where we're building what I would call my machine. That for the first time ever, we can see the way a mind works. One of the personal things my life is I've been a meditator for a long time. Twenty-two years of a deep contemplated practice. You do seem super zen. There is a real zenness about you that carries through. Probably even in your voice for people who aren't watching the video version, but I feel super calm sitting with you because of the calmness that you've exhibited. Well, thank you, but yeah. And so meditation is a way for you to witness your own mind. So in a world in which we all have to learn to control these artificial minds, it all starts with how do you learn how to balance your own mind. So I would recommend a book called The Book of Secrets. The Bicostials. It was an Indian Mystic. And he goes over the 112 ways to meditate. And it's a commentary on an ancient book,
this thousands of years old. And the thing he insight, one of those 112 ways will work for you. So if you try to meditate, it didn't work. There's 112 other ways to meditate. So that is like your guidebook or your Bible to all the techniques of meditation ever discovered. And so keep trying until you find one that unlocks the key. And it will help you have more calms than mind, which will help you manage these artificial minds. Great recommendation. Thank you. And then for listeners who want calming and insightful information from you after this episode, how should they follow you? I'm probably sure the most on LinkedIn. I think things like this are going to be a little more on YouTube. I have invested a nice microphone at home, but I spent a lot of my time working at Salesforce. So LinkedIn's probably the best place to follow me. I'm beginning to share kind of more. And it's the broadcast vehicle. So find me on LinkedIn and hopefully you'll find it
helpful for sure. Nice. Michael, thank you so much for taking the time. You know, you're the Chief Data Officer at this gigantic enterprise. We're at your flagship conference for you to take time out of your schedule to speak to me and my listeners. Greatly appreciate it. Thank you so much. Really great to be with you today. Excellent episode today in it. Salesforce's Chief Data Officer, Michael Andrew graced us with his presence to let us know how Salesforce's data team now serves two kinds of customers, humans and agents. Why agents need roughly 10 times more trusted data than humans? How Salesforce acting as customer zero used data 360 to connected snowflake Amazon and multiple CRM instances without moving the data, collapsing 266 million fragmented profiles into 141 million unique individuals. You talked about how agents let you quantify the ROI of data quality for the first time since you can measure how much better, faster and cheaper an agent performs when fed clean data versus bad. And he left us with advice for practitioners to use many models to send
the same prompt different context learned to write e-vals and build testing intuition the way you once learned a b testing. All right, I hope you enjoyed the conversation to be sure not to miss any of our exciting upcoming episodes. Subscribe to this podcast if you haven't already, but most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science podcast with you very soon.
More episodes
More from Super Data Science: ML & AI Podcast with Jon Krohn

1031: Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It), wit...
Super Data Science: ML & AI Podcast with Jon Krohn

1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gau...
Super Data Science: ML & AI Podcast with Jon Krohn

1029: How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Kati...
Super Data Science: ML & AI Podcast with Jon Krohn

1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell
Super Data Science: ML & AI Podcast with Jon Krohn