Skip to content
TrackPodcasts
businessSep 15, 202654:55

Inside Google's Billion Dollar Bet To Win The AI Race | Logan Kilpatrick

About this episode


Logan Kilpatrick is a member of the technical staff at Google DeepMind. In this conversation, we break down whether Google is actually behind in the AI race, the strategy behind Gemini 4's massive pre-training run, and how DeepMind decides between chasing general intelligence versus building specialized products. We also discuss China's open-source AI labs and why measuring real progress toward AGI might be harder than building the models themselves.

=====================

Arch Public is an agentic trading platform that automates investment strategies across Stocks, Commodities, ETFs and Crypto. Whether you’re rotating into AI & Gold, allocating to the S&P 500, or accumulating Bitcoin, Arch Public executes your plan 24/7 without ever taking custody of your assets or funds. Sign up today at https://www.archpublic.com, and start your FREE automated trading strategy!

=====================

TOKEN2049 returns to Singapore on October 7–8 at Marina Bay Sands. The world's largest crypto event. 25,000 attendees, 300 speakers, 1,000 side events and the whole industry in one place for two days, into the F1 weekend. Get 10% off your ticket with code POMP10 at https://token2049.com/singapore

=====================

Simple Mining makes Bitcoin mining simple and accessible for everyone. We offer a premium white glove hosting service, helping you maximize the profitability of Bitcoin mining. For more information on Simple Mining or to get started mining Bitcoin, visit https://www.simplemining.io/pomp

=====================

  • 0:00 - Intro
  • 0:51 - Is Google behind in the AI race?
  • 2:47 - Frontier commitment & the Gemini 4 pre-training run
  • 8:31 - Build vs. buy: Google's AI acquisition strategy
  • 13:13 - Chinese AI labs, competitors & filtering the hype
  • 19:03 - The ambition problem & where Google chooses to compete
  • 21:39 - General intelligence vs. specialized AI products
  • 35:01 - Inside DeepMind: Genome research & the innovation flywheel
  • 41:52 - Kaggle & the race to actually measure AI progress
  • 48:25 - The AI data economy: why data is the new bottleneck

Get every episode summarized

Each time The Pomp Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

495 searchable segments. Every word is indexed and playable.

Inside Google's Billion Dollar Bet To Win The AI Race | Logan Kilpatrick

The Pomp Podcast

0:00
54:55

Full transcript

The Pomp PodcastInside Google's Billion Dollar Bet To Win The AI Race | Logan Kilpatrick. Machine-transcribed; use the interactive transcript above to jump the player to any line.

College football is back. So, Hilton called to me the superstition concierge to make your fan rituals a reality. Need a room to match your lucky number? We got you. Want to make sure our team doesn't wash your lucky jersey? Oho, that smells lucky. Hilton's unmatched hospitality can keep up with any superstition. Even a marching bandwink up call it 555 and 55 seconds. Hit it! When you need a team that will do whatever it takes on game day, it matters where you stay. Hilton, for this day. Good things come in force. Like ukulele strings, boy band members. Oh yeah! And PayPal payin' for. No fees, no interest, no impact on your credit score. Just the flexibility to pay the way that works for you. Strum it, hum it, pay it in for. Download the PayPal app to get started. Subject to approval. Learn more at paypal.com slash payin' for. Paypalling. NMLS 910457. I think we are like laser focused right now at the frontier. We are seeing all these early signs of recursive self-improvement.

I think the other labs are seeing this as well. And so I think it's like underscoring the value of like... Today's conversation is with Logan Killpatrick. He's a member of the technical staff at Google DeepMind. In this conversation, we talk about the AI industry, the model lab wars. What's going on at Google? Are they actually committed to the frontier? How are they investing capital internally? What are the specific products and the decision making that's going on inside of Google DeepMind? How are the AI efforts? Other labs actually affecting their decision making. And then how should you as an individual think about benchmarking, coding agents, different types of bots, and many other aspects of the AI industry that everyone's talking about? Logan is somebody who has worked at a number of different companies in the industry. He's very well versed in what Google is doing, but how the industry is developing. And I think you'll find this conversation very valuable. Here's my conversation with Logan Killpatrick. All right, Logan, everyone thinks that Google is behind an AI race. What do you think? What is your response to the critics who believe that a Google maybe is not... The Google maybe is not near where they should be? Yeah, it's a good question. And I think it is... I think my reflection of the last two and a half years is like...

I think it's a fair criticism because people expect a lot of Google. This is what I try to remind myself. It's like, you know, it's not people trying to be rude saying that Google is behind. It's like, Google's an incredible company. Have such a story. But I guess the people expect high things of us. I think the tension point for us is you look at this portfolio of stuff that we're doing. We actually talk on camera about this, everything from like, you know, genomics work to weather and science to, you know, new frontier models, which Gemini 4 or 2, we just released a bunch of new audio models, etc. Like, I have this firm conviction that like Google and DeepMind, we have the world's best portfolio of stuff. The tension is like the portfolio spikes in different ways. And sort of this is a very natural thing. I don't think that's like an excuse to not be at the frontier. And so I think we are like laser focused right now at the frontier. We're seeing all these early signs of recursive self improvement. I think the other labs are seeing this as well. And so I think it's like, underscoring the value of like, being at the frontier.

And so hopefully we'll see that with with Gemini 4. But in immense amount of progress. And I think you've seen this with the Gemini 3.5, 3.6, 3.7, 3.8, lined up of like, in literally like, 3 to 4 week increments. Sometimes less, sometimes a little bit more. We're seeing like very reasonable progress. And again, this is like the early signs of this recursive self improvement loop. So hopefully we'll see that sort of like translate over in the same way to Gemini 4. And it'll be our largest most ambitious pre-training run so far. So I think you'll sort of get us back in contention with some of the frontier labs. Talk a little bit more about this like commitment to the frontier. I think that is one of the things that people have always wondered is, you can go after general intelligence, you can go after specialized workflows, you guys have a business to run, you get a lot of cash, but also if you're making a lot of bets on, you know, this kind of being the future of the company. But I do think that when I speak to people at Google, maybe there's more of a quibbid to the frontier than people realize kind of outside of the company.

Yeah, it's part of why I had this conversation with Korai, and he was sort of like making this verbal commitment to the frontier. Korai's SVP of DeepMind leads the organization now and is incredible. I love working with him. And he was sort of making this verbal commitment. It's obviously been everyone's perspective internally for a long time. So I think he was sort of doing it to answer a lot of these questions of like, do we even care about it? And I think this stems from like folks asking the question about, you know, why haven't we landed the pro model, why are we pushing on flash? Do we only care about flash models now because we're shipping flash? And the reality is like we were just seeing a lot of progress. There was a lot of juice to squeeze. The recipe that we had was like working extremely well with the flash models. And so there was this like really tight iteration loop. People were getting excited about the progress. And they're like, hey, let's focus and make that happen. But I do think like one of the things that makes me most excited about Google and the reason that I'm here and the reason that I'm doing this work is because

the front having a frontier model and the commitment to the frontier so deeply permeates across Google's business. You look at like every single thing that we're doing as Google, the way that you know are 13 plus billion user products are sort of like touching the world in different ways from workspace to everything else we're doing across search. Having a frontier model is core and fundamental and a direct accelerant at every single component of that business. Not to mention the things that are like more exploratory. Like what is having, I mean, overseeing science of this from OpenAI and others. Like what is having a frontier model mean for frontier drug discovery and being able to cure cancer? And like obviously there's a correlation between those things. It's like a not immediate correlation as having a bunch of products with billions of users. But like there is going to be some correlation. And I think that correlation is going to increase over time. So it is like the business is set up to like fundamentally depend on having a frontier model. And so you go ask everyone, I think this is to your comment.

Like it is the most important thing. There's nothing else. We're not thinking about like, hey, let's go make small cheap models that are great for everyone. Also like we do that because there's use cases in Google that support it. But the aim is to be at the frontier and being at the frontier will enable us to do all the other things. As you go through that decision making, it is kind of an interesting thing. Most businesses, they would look at like what's the problem we face. Okay, let's go build technology to solve that problem. When you're committed to the frontier, though, there's some version of just like let's go build the smartest model possible. And then we'll figure out how to apply it. And then once you get into the application of that, you know, kind of general intelligence, it becomes okay. Do we do it in a general way where people all throughout the company can just, you know, ping it or customers can do that. Or do we go build like specialized workflows and kind of go through that path. Maybe just walk me through like the decision tree, like bring us in the room. As you guys are thinking through some of this stuff, like how do you guys decide where to put resources? And like maybe what the sequence of events is, at least, you know, aspiration for you. Yeah, I'll make a general comment, which is, you know, Sam Altman has that famous interview where he's like being interviewed by people a long time ago.

And they're like, so what do you, what's the plan to make money off this thing? And he's like, I don't know, we're going to make the smartest thing possible. And they were going to ask it how to make money or something like that. And you know, a little bit of a tongue in cheek answer, but I think that's actually like less of what we've been trying to do at Google from the sense of like we actually know where the model is. We'll create tons of value for our customers in the world. We have all these products. We have all these services. We have the fastest growing cloud business in the world, et cetera, et cetera. So it's very obvious what the commercial application is. You don't have to think too deeply about that. I think to answer the question specifically about like what are the set of tradeoffs and how is the resource allocation being made? I think that's where this like really sign of sort of like recursive self improvement is coming from. And I think there's a lot of resources being focused on like how do we actually make the models better at coding and research and science and sort of the core work needed to make like further breakthroughs and accelerations of model progress. And I think folks are very interested in that because there's just like such a large economic opportunity and then having a great model will then enable us to do all the other things that we want to do.

So there's a huge I think everybody is very code-pilled, very science-pilled, very focused on that right now. And I think we were like a little something we relate to the game because folks I think knew it was important. But I think you, it's like yeah, in hindsight everything is much more clear. Like I think in hindsight now is obvious like we should have you know from an order of magnitude of resource allocation probably put more into coding sooner. And like that makes sense. And we had a bunch of other stuff that we were doing which those things actually turned out quite well. Like you know good example of this is you know nano banana like a great incredible image model that sort of took the world by storm had this massive impact for our consumer products and a bunch of other parts of the business. And like you know to make that model it took research and compute and time and sort of and had this huge impact. And like was it in hindsight right to do that versus doing something on coding like I don't know that there's actually a clear right answer. But like those are the types of trade-offs that are actually a lot easier to analyze in hindsight.

And it's harder to know in the moment whether you're making that right decision. But we do that reflection and sort of introspection which is important. If you look at some of the companies I think open AI has been very inquisitive in trying to go after some of these verticals obviously a SpaceX AI recently went back cursor. And then they launched rock bot and they've got a lot of coding stuff. It does feel like there's different strategies. There's a little bit of chest getting played. Some people say hey look we're going to focus on certain things we're going to build it from scratch. And that's kind of our ethos and DNA Google and other areas outside of AI has done both they built some things but they've also acquired things. How are you guys thinking about you know if you feel like you're behind maybe encoding or other areas like will you guys go buy stuff or is the focus internally on like let's go build. We know how to do it. We've got the resources. It's just a focus and that kind of strategy thing. Yeah, we've definitely done both in this context. And so we had I think the Gemini CLI launched early maybe like almost like a year and a half or two years ago. And sort of had some early traction. I think a few million users using that product.

And then we also went and did the sort of acue higher of the windsurf team which is now part of cognition who we were talking about off camera. Yeah, and so we have a bunch of those folks they've sort of been building this anti gravity product internally both for our internal engineers but also for external users and have seen actually I think like one of the biggest impacts has been the internal acceleration. And so they're really really deeply focused on like how do we build a great product for for engineers and actually even non engineers now inside of Google make it work really well. And then that sort of will translate to a great product externally eventually. And so it's been cool to see us do both things. I think it's one of the interesting strategic advantages of Google is we get to take many shots on goal. And so yeah, I think having a bunch of different coding products. And it's also like I think what's most interesting on this threat is the ecosystem has evolved. I think about this all the time like the coding product that you would go to market with today or the sort of like developer or whatever like knowledge even more generally knowledge work products that you're going to market with today look so different than what it was even 12 months ago and actually like rockbox a perfect example of this like it's not obvious to me that like the rockbox style product would have worked 12 months ago.

I think it works now because the models are good enough but like 12 months ago you actually like needed all of this like developer UI scaffolding of like you know give me all these extra features and buttons and things because like the model is not really that smart and I need to wheeled it and turn it and sort of critique it in these very, very specific ways. And so I think one of the my observation of this and my point of this is like I think there's going to be many of these opportunities to take shots on goal as the model progress continues it like unlocks a new paradigm and it feels like you know rockbot and an instinct and news and a bunch of these new products that are coming to market right now are like the evidence of like the models have crossed another chasm where the product experience you can now build is fundamentally different than the one you were building 12 months ago in this this is true for developers. So I think it's true that a bunch of other verticals as well. With Gemini 4 you guys have talked about this large pre training run that you're doing what should we take away from that like are there specific things that you can share in terms of what that's going to look like.

Yeah I think the thing to take away from this is like the commitment to the frontier doing large pre training runs is extremely expensive it is like a large order of magnitude of investment it's like a I don't the numbers are quite large. You want to tell us I don't I don't actually don't even know what the numbers are at the top of my head just but I can I can do some of the math in my head and like it's a lot. And I think it's important I think the other point of this actually is like pre training has been a significant strength for deep mind in the past. And so I think this is like one of the areas where I think we have like some of the best talent in the world. We have like actually this is like the infrastructure scale of Google is an advantage in this context like the TPU fleet is an advantage in this context. A bunch of the data infrastructure stuff we have is an advantage. And so I think the points of telling people about this pre training running is to tell people that like it's not like we're rolling over and playing dead like we are pushing the frontier everybody's working as hard as humanly possible will hopefully see a bunch of incredible results from this new pre training run.

And actually interestingly like you do the model comparison today and like all of our you know I think this pre training run will very specifically like get us to the category that we need to be at to be competitive with where competitors are at. And so I think the proof will be in the pudding when we hopefully watch this model and customers get their hands on it. But I think that's the expectation and that's the hope right now. When you say competitors I think most people will think about open AI and thrott big you know GROC or SpaceX etc. Do you guys worry at all about like the Chinese open source open weight type model players do you worry about maybe other competitors that aren't one of those three companies. Yeah I think what's so interesting right now is like this space feels incredibly dynamic and so I think I mean I think I personally have never understood all these memes of like I don't think about the competitor I'm like I think about our competitors because they're all incredible companies and like I think we would be wrong to not be thinking about what they're doing and you know examining are they are they making the right decisions are the things that we could be doing differently still sort of like knowing what our core focuses.

And so I spend a lot of time looking at like what are the things that folks are doing and obviously the Chinese model labs have done incredible job so far it's like there's there's clearly a bunch of like question marks as far as IP stuff model trading stuff but like with the set of constraints they have all things considered they've seemingly done a pretty solid job and like there's clearly research innovation they're doing as well. They're just copy whenever one else is doing there's like actual frontier research happening. So don't want to discount them as a competitor. It's also clear that like startups and companies want to use models that they can host themselves like I think there's like a there's like a philosophical question of like oh how do you you know what's the what's the feeling about these labs in China open source in these models and doing the thing they're doing and there's like a practical business question of like customers want these type of models they want to like I talk to start a business model. I want to like I talk to startups all the time and startups want to be able to take the way to the models and customize them for the use cases that they care about and so there's a huge market there there's a huge opportunity and so I think the question is like will we see like US open source labs and like in videos you know spinning up these types of efforts we have some of this on on the smaller on device model side with Gemma.

I think we'll see like reflection AI a bunch of other folks like take shots at like can you actually produce frontier open weight models. But I think the cool actually the cool thing for all of us is that how just how competitive it is like the fact that those labs are able to like stand in a similar regard at all to these large companies in the US is actually I think a good thing for all of us right now and so there's a huge amount of competition that's pushing everyone to be better. Today's episode is brought to you by token 2049 the largest conference in crypto is back token 2049 will host 25,000 people 300 speakers and a thousand plus side events in Singapore on October 7th and 8th and Marina Bay. The speaker list is absolutely stacked Shane Copeland from polymarket Jeff Yann from hyper liquid a dean of freedom from Nasdaq Arthur Hayes Bologi and Eric Trump crypto and traditional finance in the same building which tells you a lot about where this is going. And the conference runs right into f1 weekend so the whole thing turns into one giant week. If your head is Singapore use code pop 10 for 10% off your ticket token 2049 October 7th and 8th in Singapore go check them out in the link in the description.

When you think about playing chess you definitely got to understand what your opponent is doing I agree with you that you know kind of only focusing on your pieces does not help you win the game. But with that said though it does feel like there is a lot of question marks about how some of this stuff is getting done. You know if you think of some of the math problems that have recently been solved it was hey did the models train on you know other people's questions was there some peaking at you know data that people thought was private. There's some questions now about was you know Kimi actually passing some of their queries just a claw to answer versus Kimi doing it themselves and it's very difficult you know at least for me but I think many people to understand like what is real and what is just like Twitter fodder or X fodder and where people just say you know they like they like the drama is like the TMZ of the AI industry right like what's the new thing that we can all you grab hold of for the day. How do you personally think through you know where to spend your time in terms of like your attention because it's happening so fast there's so many different things you know even if you we just think over the last week or so you've got everyone from Paul to their Jones putting out off edge you've got you know Jensen talking about a G.I.

You've got Astro you've got like all these components. Unless you figured out how to get one 24 hours in a day you know you don't have as much time so what is your process to do that. Yeah I think this is actually an interesting point that you're making which is and I think this has I think the the trend has changed over time in like the level of signal to noise. I think there's actually just a lot more noise these days and so I do think it is like a it is a muscle that you have to build to sort of filter as much of this stuff as possible and like for me it means like I am definitely passively consuming a bunch of this stuff and trying to engage in things but like I'm trying to stay focused like we need to be at the frontier the most important thing that I can do is like help us go build better models. There's a bunch of stuff to stay on top of and make sure that like we're reacting to the right things that are happening and being proactive where it's needed but like I think is really easy to get caught up in all the crap that's happening in the world right now and like I think my advice to people is like filter out as much as possible be more intentional about how you spend your time because like there's a lot of noise and it's not always clear to me that like the noise is actually translating to any amount of signal and so I think I think I'm going to be able to do that.

And so yeah it's like there's yeah there's very specific cases where this is true like you know the hugging face open AI situation is like good example of like lots of noise there's definitely signal there there's something to be learned there's something to understand there's a lot of cases though where like this is not the case. And so yeah trying to train me intentional about the places where there's actual signal. If you almost take that same issue or challenge and flip it the other side of that is like there's a lot of opportunity cost given that the cost or barrier to build things has come down so much. You now have access to superhuman intelligence you can vibe good things you know you can have the bots go and build companies or you know kind of run parts of your business like that it does feel like not only are there more distractions but if you get distracted the opportunity cost is higher than ever. And so how do you think about you know your role internally you've worked inside of open AI you've worked at Google like maybe like what are some of the things you've picked up and how you're navigating the productivity side of this as well.

Yeah it's so true and I think actually the thing that I struggle most with now is like it's this like level of ambition problem which is like I used to be able to be like oh I'll just go like do this thing it's going to be small and concise and well scoped and now it's like shit I actually it's like I'm going to be able to do this. I actually like if I go do this like this could be a billion dollar opportunity for us and like so I have to take it like quite seriously like that like weighs on me and I'm like you know having to spend more time to be thoughtful about like is this the opportunity that like we really want to go after as a team because there's so much opportunity everywhere. I think for me this goes back to like I try to be extremely principled about like what are the things that Google is well positioned to compete in and there's a lot of things that were not well positioned to compete in there's definitely some that we are well positioned to compete in and we're like we have structural advantages with you know Google workspace and Google Cloud and distribution and things like that. And so that's sort of my filtering mechanism on the on the product side when we think about like what are the opportunities to go after like I don't want to go after everything I want to go after things in which there's like a natural lift because we have other assets inside of Google that will actually contribute to the success of these things and so this is what we've done in a studio we have all these deep integrations with Google Cloud and all this stuff that like no other product team in the world can actually do because they're not inside of Google building this product.

And so it means that we're it means that we're competing in some of these categories but we're running a playbook that only we can actually run and so we'll see in the fullness of time was that the right playbook does it actually make sense maybe we should have just been doing the things everybody else we're doing. But I think it's I'm trying to keep that filtering mechanism very tough of mine as we're making the decisions and there's just so many cool things that Google has that make this like fun and so I feel like I'm not limited by like by this at the moment. One of the aspects of the AI industry that is just intellectually you know stimulating I think for you me many other people is there's a level of strategy that is being played out so it's not just like can you get the hardware and the software to do certain things you know kind of create magic or turn sand and intelligence like that that is obviously very difficult and plenty of challenges there. But the strategy side you know we were talking previously that most of the frontier models are pursuing general intelligence some form or fashion right but then there's a bunch of startups that are saying well what if I take a specialized workflow approach and if you think of like what we've been building with Soviet you know this idea of well if we go and we build a bunch of proprietary technology that from model routers to harnesses to you know data pipelines and our models et cetera.

It does feel like there's almost a point on each application of AI where you kind of have to decide do we go after general intelligence or do we go after the specialized workflows and Harvey Sylvia there's many players I think that are seeing a lot of traction specialized workflows. But what if I'm fastening about Google is you guys have multiple applications where you have to make that decision over and over again like do we go and build a specialized workflows or can we just use the general you know purpose model. How are you navigating that like for each one of these use cases is almost like you you guys maybe making the decision more than anyone else in the world. Yeah actually I've got two points for you on this like one of the thought exercises that I am continually proposing to our team internally is in five years do we expect to maybe five years is the wrong time horizon but like in five to 10 years do we expect Google to have 10,000 products or two or three years. And I think this gets to this like vertical workflows versus sort of like general intelligence and so I think there's like clear signal in the market that customers want vertical applications.

They don't like you know you think about like why apps are so successful and like that sort of there's this this user behavior pattern which is like hey I think as a user in terms of like using a particular application or a tool in real life. I want to go swat a fly I go get a fly swatter I want to go drink water I get a cup I don't like go to this like Oracle all encompassing tool that can morph to do is like it's not something that like we intrinsically have grown up and sort of evolved this humans to understand. And so I do think there's this like really deep rooted muscle memory and I think the tension point will be given that extremely deep rooted muscle memory does it. Is that enough of a a sticking point that will like keep these vertical applications alive in a world where alive and thriving in a world where like the general purpose thing can actually do the same stuff. And so I think that will be the most interesting and this is where I think this like you know what's the intersection of like AI and new hardware consumer hardware devices I think is going to be really interesting like as people change the way that they work with software and with technology like you imagine you will want to do that.

Imagine you will want like new form factors because the form factors we have right now are sort of a little bit more of these like verticalized experiences. I do think there's a separate edge of this which is and this is true in all these vertical domains like the vertical domains are successful also because somebody is focused. And I like deeply believe this like you know startups are always worried about like oh is the big company going to come after me and it's like you can always do a better job in the big company with like very few exceptions because you're focused and more deep on some vertical that your problem that your customers have that like no one else is going after. And so I do think it's like an edge to have that vertical mess which is really interesting. And so I think the other point that I wanted to make and I'm curious actually what you think about this I think there's all this conversation of like general intelligence and something that's been like that's been very top of mind is I think the labs have historically described general intelligence as if like they would build the general intelligence themselves.

And that like seemingly the general intelligence would then be powered by like end to end almost the models that one of these labs creates I think there's something really interesting about this future where like if we really had general intelligence like you wouldn't expect that the general intelligence that Google creates is only using Google product services and models you'd expect like hey if that's generally intelligent enough to know that like there's other thing that some other company created can do something that our thing can't do or can do it better. Humans are generally intelligent not to figure that out and so I think it actually it's going to add a lot of I think on this path to general intelligence is like the model labs go and continue down this direction I think it's going to add a lot of like confusion to even like understand what that really ends up becoming because I think it's going to look a lot. It's going to look a lot more chaotic I think and how this like these general intelligence systems play out and I think this like beautiful vision of something that can just like do anything and everything for you it's interesting talk about this so before to have the general tone let's talk about like a microcosm of this and you know the problem that I've been thinking the most about for last year and half is Sylvia but for those that don't know the product you basically come in you touch your financial accounts you put up your private investments you start talking to Sylvia but we have chosen to go specialized workflows and we have done a whole bunch of

you know very innovative things that think in terms of the harness the memory and file system the model routers that you know fine tuning etc but if you take like the model router one of the perceived advantages of not being a frontier lab is that you should be able to route queries to any of the model lab you know models right and so what are the odds that Google is going to route to open air and traffic it's not zero but it's not a you know 90% either right and so same thing I think with each one of the labs is like what is the incentive for them to keep the queries within their family of models versus the ability to act more as like a third party and actually route across. I don't think we've really seen how everyone's going to play that and so as a third party you're like well I don't really care right I just want the best level of intelligence at the lowest cost that answers the question for you know our our user and so I do think there's some of those things also where people are trying to figure out not just like you know if you then extrapolate this out to like general intelligence.

The user doesn't care if it's a Google product or not maybe there's some like ethical or moral things that maybe they align more with but for the most part they just use a product because it's the best one. And is that going to actually be built by one company or is it going to be you know kind of a bundling. What's the saying is like the world is just bundling and unbundling over and over and over again and so you know it is a very difficult thing to predict because I think we're so early in this journey that like every week so it's got a new model that seems to leapfrog everybody else and then you're trying to predict how consumers are going to interface with this stuff. And maybe like the last example give is in my own life I have been using for the last couple of weeks Grock bot professionally and instinct personally. They actually do a lot of the same stuff what's your point about like you get the cup for water and you know you get the fly swatter to just what the fly like I just kind of have in my head you know okay instinct when I got a personal question and Grock when I'm doing something professionally.

It's probably pretty dumb you know like if they do the same thing like why don't you just use the same product but it then goes to like okay well now you're starting to see four or five others come to market. And as somebody who likes to be an early adopter like why should try those but then what about the context what about the memory how do I port that over and there's like this you know kind of like user journey we're all learning of is it worth the time to go try the new thing. If I don't have some kind of shared memory or you know especially get network effects were like your wife and you have shared memory somewhere then how do you interface with something and so like. It's almost like more questions than answers right now and I think that's probably why you know you and I and so many other people are so excited about this right. Yeah yeah I think you're right and I love this like the bundling and unbundling analogy because I think there's another version of this is like what's old is new again. Or what's new is old and whatever the expression is and like actually you see this with what's so interesting about crockpot and instinct is like this like form factors like what's old from a form factor is not like message like chat was one of the original ones you had all these like meh check thoughts like in the 80s and 90s or whatever it was and then like that came back and then boom all of a sudden what's old is new again.

And then the same thing is now true for messaging it's like we all use all these messaging apps and then it's like on now all of a sudden the hottest form factor for AI is like the messaging and so I think it's an interesting exercise of like it actually in all of these domains. I think the reason people feel this way is because like there are custom to this experience getting some customer to like adapt to some futuristic new thing is like actually extremely difficult to do and takes a really long time you want people to go to some form factor they're familiar with. And so I think about this all the time for like how do you actually get consumers or users to go and adopt new technology is like you want to make it feel familiar and I have this hypothesis that like messaging like in actual like chat apps like has it seems so unlikely that that's not going to be the dominant form factor. And a few years from now like it's surprising to me it hasn't been more dominant and I think it's because actually the operating system like messaging app owners are like just at the cost of this but like you'd expect you know like Apple and Android and WhatsApp etc to like really lean into that form factor and they already have where all the communication is happening and you throws

the questions in there and like you know it it feels like it's a natural place. It does feel like there could be some platformers now I think that the people developing these products are obviously partnering with and trying to prevent that but you know if you wake up and you've got a chat box that's very popular and all of a sudden you're blocked on Apple's system that would be a big problem right and so you know I don't know if that's really an Apple's best interest to do that stuff but I do think that there's some folks who are kind of thinking through that. Another aspect though around you know the kind of the chat bots and the messaging I do agree that it's an interface that we all are very comfortable with you know I use it on a daily basis with these bots but I have not yet become a very big voice user and I have a lot of friends that are like voice build you know they're walking around the like whisper in their microphones or whatever sitting at their desk do you use voice or like what is that maybe the adoption if you had to predict it inside of like the AIT at Google in terms of people who are you know that fingers on a keyboard versus using voice.

That's a good question actually and I I've been flowed between this there's like what actually what I found is like the best use case for me for voice is when I'm like doing some sort of demo in front of other people and that way I don't want to like fumble typing things and I can't spell and all that stuff and so just voice like straight is like way faster it makes the point it's like much much more succinct but there is something about I think it's a and I'm sure there's like good you know neuroscience research out there that explains this like I think it is like a people manifest thoughts in different ways and like the physical manifestation of thoughts either coming audibly or like through tactile typing or even writing like to me I feel like I have like different thoughts depending on the sort of expression form and so like the way that I spend a lot of time talking and doing stuff at work and I spend less time writing sometimes and so it's like actually quite helpful for me to like pulls me into a different mode of thinking when I start to write just because of the form factor and so I think we'll see actually more of that as well and that's I think the you know obviously people like audibly speaking and it's a helpful way to think through things as well

but I think we'll see the sort of buckets of these different types of thinking manifest from like how you interact with AI as well. It is interesting I think the science shows the single best way to remember something is to physically write it down like with your hand next would be typing right and third is just kind of hear it and don't do anything but maybe there is something about not just the memory but also the ideation right you know we definitely know from science that walking outside kind of the act of moving you know forward does a lot of ideation showering right as many kind of examples throughout history people who just want to shower and thought of things challenge twice a day for this reason it's not for hygiene it's just for I mean not tongue in cheek but sometimes honestly because like you do just have head space it's great. Yeah and look part of it is like are we just so all terminally online that just like the showers in a place that the phone doesn't go or is it like there is something about you know even 50 years ago before people had you know super computers in their pocket the shower did lead to new ideas right I think the shower is just cold 50 years ago and so people are just being shocked and

and you know having new ideas probably I love it let's talk about deep mind more specifically you know the work there obviously has has been very broad for a very long time we mentioned a little bit about the genome kind of project and the work that's being done there I'm pretty surprised at just how large it is but it still doesn't get maybe the respect that it deserves can you talk a little bit about some of what's going on there yeah no 100% I think I'm not an expert on all the science stuff that we're doing but it's incredible to see the progress I think across the way that I would frame this is deep mind is split up in a couple of different ways they're sort of like foundational Gemini and there's like we want to make the best frontier model and a bunch of different sizes of that model and sort of all of the different modalities that work in mainline Gemini and then there's a there's a whole science unit and inside the science unit there's everything from the genome project a bunch of the alpha-fold stuff there's a bunch of like science things related to like biology there's

a bunch of other science stuff related to like weather and mathematics and that whole portfolio is like also at the frontier of doing all these really interesting problems that no one else is doing and then has all these very unique collaborations actually with folks like isomorphic labs which is the are sort of like drug discovery company inside of google and deep mind that demises the is the CEO of and you know they do all these deep collaborations and so it's a really interesting way for them to like not only solve the problem and make progress in the problem from a foundational research perspective but then actually have like the applied side of it as well and so i think it's this like unique flywheel that exists inside of deep mind itself where like we're creating a bunch of the frontier innovation we're doing all this interesting science work and then it actually has an application it's not like we're just doing it for the sake of doing it and i think this was actually the lesson from from alpha-fold which was like hey we did all this really really interesting work it was super interesting

but we're actually doing it like to solve a scientific grand challenge less because like we had somewhere where it was an immediate commercial application of but it's like it became very clear like hey there's all these commercial applications we're opening this up the scientists are all using it and so i think the general philosophy is like do this frontier science work across genome across weather etc and then actually have a place to apply it to inside of google and then actually most interestingly take a bunch of the lessons in learning and data and other things and upstream those back into the mainline gemini model because ultimately like the mainline gemini model is going to become better at a bunch of those things than those individual domain specific models and so you need to make sure the flywheel also goes back to there and so that's why having it under like a single roof actually makes sense and we see the cross-pollination between these things we've seen historically like all of these interesting like alpha-proof with mathematics trickle back to directly increasing the reasoning capabilities of the model for mathematics in the mainline

gemini model we've seen this for cyber now it's it's not an alpha project but it's a similar domain where like cyber capabilities as we push the frontier on cyber directly correlate to like models having better coding capabilities and so there's all these other examples where like this flywheel spins and i'll make one comment which is i think people talk about the flywheels stuff like this as if it's like a magical thing that just like works there is an immense amount of effort and energy that is required to actually this the flywheel does not like you don't spend it and then it's like a hamster wheel it's like you are manually pulling it and like forcing it to work because like you know that the outcome is going to be great but like i have to remind myself this and our teams this internally because you think of this this magical thing that's always spinning and you just throw things into it that's not how it works it's a lot of effort and energy to make the thing actually today's episode is brought to you by arch public arch public has just expanded its agentic trading platform beyond crypto so pay attention this is a big one now they are

automating strategies across stocks commodities and ETFs and i think that this is going to be huge you can now automatically take profits when one market hits new all time highs and rotate that capital into other markets showing more opportunity whether you're rotating capital into AI stocks gold if you're investing in the smp 500 or you're accumulating bitcoin arch public brings real discipline and automation to your investment strategy additionally they've launched a powerful new tax loss harvesting tool with crypto being so volatile and its exemption from the wash sale rule arch public can offset gains with losses without compromising your long term positions it's exactly what every serious investor does institutional grade automation that works across every major asset class there's no more emotional trading no more missing tax opportunities just smarter hands-free execution of your preferred strategies go to arch public dot com right now connect with their team set up a time bring your account if you'd like then you can learn what automated trading can do for you arch public dot com today's episode is brought to you by simple mining bitcoin mining has a reputation for being complicated risky and hard to evaluate as a real

investment if you're considering mining in 2026 what actually matters isn't headline profitability it's uptime repairs and whether the operation is run like a real business that's why I've been using simple mining they're based in sear falls iowa and they run a white glove hosting operation where you own your miners you choose your own pool and you have bitcoin sent directly to your wallet they were featured on the ink 5000 lists as the fastest growing company in iowa with over 40 thousand machines under management what stands out to me is execution they have the number one rated a sick repair center and for the first 12 months repairs are included if mining margins get tight you can pause with no penalties and if you want to resize or upgrade your fleet there's a marketplace to resell equipment instead of being stuck to help people think it through with their mining actually makes sense right now they put together a short resource called the 2026 bitcoin mining blueprint it walks through the five mistakes investors make when allocating the mining and they also explain how to avoid them before deploying capital if it sounds interesting to you you can get it for free at simple mining dot i.o slash pop that simple mining dot i.o slash

pop go check it out today and see if you should get into the mining game it's funny like uh i thought AI was just going to solve all our problems we would have no jobs um you know always be like hanging out at the beach but uh i have said it over and over again every single person I know is working harder today than they've ever worked in their career and a lot of that I think is just they feel like it's a big moment you got to kind of accelerate to be able to to capture uh kind of your piece of it but but at the same time I think that people are inspired right that there is this element of imagine if you can be part of a team that accomplishes you know xyz thing and not see deep mind is a big part of that before we before we let you go uh let's talk about is it cacole or cacle how do you actually pronounce this cacle cacle all right well explain a little bit as to what it is you're now running cacle um and it may be kind of like what your vision for for the product yeah i think what the the sort of um the perspire actually historical context cacle is sort of a startup google acquired and i think 2016 or 2017 it's done a bunch of interesting stuff inside of google

we sort of brought the team over um into uh to be part of my team earlier this year and the sort of the basic hypothesis for this is like model progress itself and like this is like such a important point to underscore model progress is gated by our ability to measure progress like you cannot make progress on something that you aren't able to measure and so actually as you see one of the most interesting things in the last like three or four weeks is like you look at fable 5.1 you look at astra you look at hopefully jimin i4 as it lands in the market like these models are saturating all of the available benchmarks and so now you sort of sit there and you're like okay well where where do we go like we don't there isn't a bunch of problems that are difficult that sort of that we can actually measure and like scientifically continue to hill climb and so the mission for the for the cacle team and for building this platform is like we want to build the most open benchmark and evaluation platform in the world so that people can come together and collaborate on

on like all of these extremely difficult frontier benchmarks and challenges and competitions so that we can actually see difficult problems that models can't yet have not yet saturated and we know what we can actually measure and i think there's like a bunch of new on spits of this like the the everyday and i say the everyday person in quotes because like i've heard the everyday person is not going to but like people who care about this technology being able to show up on a platform and have a voice and like and a say in you know how are we measuring progress towards a g i what are the things that we should care about what are the types of tasks and benchmarks that like actually prove these things what are the for my comfort company what are the things that you as a company building Sylvia actually care about what are the capabilities you wish you had in a model that would unlock entirely new sectors entirely new geographies entirely new use cases for your customers and having a place where you can articulate that in a way that this is my i had this epiphany a year

and a half ago where i sat in years of like customer conversations where sort of you take a customer and they'd say hey i wish the models could do this and here's an example of that and then you'd see a researcher sort of you know with a blank stare because like it's the the work to translate sort of one anecdotal example into something that like an a i researcher can actually take action on is like it's on two ends of the sparrows impossible it's not it's not capable and so you have all this great feedback coming in from customers and you can't take action on it from a model perspective and so the exercise is like how do you get people to speak the same language the language that the researchers speak the language of model improvement is in the form of benchmarks you need to be able to measure something to make like scientifically rigorous progress on it um and i think the world is slowly starting to wake up to this fact and i want to help accelerate this because progress is bounded by our ability to measure progress and so yeah excited we're like definitely in the early

stages of this we'll have lots more stuff to share soon but trying to get the world building more benchmarks so that we can make progress for the stuff that like real people actually care about not like a bunch of academic stuff that people don't care about but like use cases that like you personally your company have and every other startup and company has um is really important i think this is like one of the problems of the of the decade we're we're talking earlier about you know one of the things that we've been talking internally quite a bit about that uh i still don't have an answer for i don't know if anyone does but when you look at these benchmarks you know uh you may see on a scale one to one hundred uh somebody comes in an eighty two and somebody comes in at seventy eight and you're like all right well i know eighty two is a higher number than seventy eight so like their quote-unquote better what's at the naked eye does that actually a difference that the human user can even tell you know what does that mean the four percentage points is kind of like a benchmark we invented right it and like is it real is it not is it noticeable it does it improve accuracy or like whatever the thing is and to me you know the work you guys are doing

there but but just more broadly as an industry ten years from now we'll probably have an excellent answer we'll be like i think the new ones to this and this is this is why the building the platform and the transparency matter so much because all of the detail is in exactly what are the four tasks that are different between those things and so here's a great example of this like if you haven't spent any time looking at benchmarks before like you the more time you spend the more you realize the things that we're measuring quality on and the things that people are talking about are fucking crazy like none of it makes sense like for example and here's like one specific example there's some of these new coding benchmarks i won't name names because i said that's crazy not a lot of disparage these folks because i think they're doing a reasonable job but like some of the new coding benchmarks have like six percent of tasks on like a programming language called zig zig nobody's ever heard of zig before this is not a programming language that any engineer at any

company is actually using i'm sure some people are using it but like the six percent difference could be like the quality on zig and like maybe some model happen to get access to some data and whatever this language is but like that doesn't matter for 99.9 percent of startups nobody cares about this thing and i'm sure they have a good reason for including that data but like being able to do this introspection of like not just there's five percentage points difference between these two models but like why is there a five percentage point does that actually matter for me as a business for me as a developer for me as a user of this model i think it's the whole game and like the products and surfaces and like even the benchmarks themselves don't do this right now they sort of show as in this this empirical thing that you know your 75 percent should be the same way that i perceive 75 percent which is completely not true and so i think it's like a fundamental problem with the way the things are set up right now it does feel like on one hand personalized benchmarks are going to become a thing yeah i don't know how somebody smarter than me will figure that out but like that

obviously is going to be you know important for businesses that are kind of like hey here's my specific ramifications the second thing though is we came out at sovia and we showed that a lot of the harness work and things that we had done made the soviet product more accurate than the frontier labs at answering tax-related questions and to me i'm like you know more business minded not as technical as as the engineer and team like great you know we we are higher we are better we are more accurate you know etc and immediately the engineer and team was like we better publish the evals the you know the rubric we better publish the user parameters like there was this entire effort as to like how much can we publish an open source without actually giving away things that would be considered you know kind of very important IP related you know type things and we went through a strategic you know kind of debate internally as to like there was definitely some stuff that we published that we could have not published and it would have maybe given us a little bit more of an

advantage but it was like if you're not a frontier model and you cannot you say that you're you know more accurate on something you almost have to like open source more to let people validate it themselves and so i do think that if you're a frontier model like you kind of don't care if people believe you are not because you're just like you know here's the evals whatever but the users care right like the kind of what you're talking about i think is a is a different situation and so it's less the academic application for you know who's got the best model and it's more about like i'm a business or i'm a user and i'm trying to evaluate which one of these things i should use for my specific use case i mean the benchmarking industry is going to be you know significantly bigger than it is today not you know obviously you guys kind of have a lead there in not in what you're doing yeah same thing with the data industry that's what's most interesting is all this this data moment is and i don't we don't need to talk deeply about it but like it's having this like crazy i'm sure you're seeing this on the start-up side it's just like absolutely ridiculous and i think the framing of this is like 2023 the question was like does the recipe work

do we have the recipe to get to general intelligence to get to this sort of like a g i thing in the future um i think we de-risk the recipe and we've made a few tweaks over the last few years but like generally de-risk the recipe then it was obvious like oh shit there's not enough compute in the world what's blast hundreds of billions of dollars into getting compute online um that is the going to be the blocker and sort of like to keep scaling up we need more compute etc etc all the all the labs have now done that we now have enough compute it's coming online there'll be further investment but like generally people know that that's something that needs to be solved it's now all data bound like the data to make progress on the model does not exist in the world like it is data that has to actually be created net new that doesn't exist or is coming from like even startups i think there's like a huge like wave of like startups that are going into these like exclusive data licensing agreements with like data providers or model labs and like the pew you know if you want to make progress it's all data bound um and so it's like this like data business is very tied to this benchmark

ecosystem is like very tied to ultimately model progress at the end of the day um and so it's very interesting to see like how quickly these things are like spiking uh up into the right i um in very bias but uh i'm an investor in micro one i think they've done a fantastic job on the data side um but i'm also an investor in a company called sunset um in sunset uh they started out as a company to help other companies shut down so if you have a startup it doesn't work it's a pain in the ass right like you got to get the lawyers involved you got to figure out how do i say as much money as possible to get back to investors but i also have like you know kind of a responsible way to wind down and that's where they started they i don't even know if they could spell a i at the time i love them but you know that was not their focus well all of a sudden they realized like there is a unmonetized asset that these companies have which is like all the slack messages in google drive you know just like the corporate data and could they basically at the point of shutdown buy that data from the company which creates a new asset that then can help them get more

money back for their investors they have to clean it and structure it and you know kind of do all these things to make it a usable form but then they can turn around and they can then sell it to the model apps and so you almost have this like beautiful thing where like a normal company would never want to sell that data because they were worried about all the competitive components and all the stuff but if you're shutting your business down you're like looking under the couch cushions for a couple you know pennies right you're like hey wherever we can find oh you want to buy our slack messages we were just going to delete them so like here knock yourself out right and so to your point like i do think that this has happened i've seen a ton of startups in all color work you know medical etc they're just trying to figure out how do we go and find data sets and no one else has has yet and then turn around and let's go and use it for robotics you know model training whatever i don't know how big that thing can be but it feels like we haven't even scratched the surface of what that whole industry is going to look like it's going to be massive and i think actually the the most difficult part of this and this is the part that's still like a dark art is like having data is not necessarily the problem at the moment the problem is like getting the data into

a format that the model labs can actually use or that the data vendors can actually some of the data vendors are not doing a bunches of this stuff but like it's the most difficult part because like the raw slack messages like there's like an immense amount of work and labor that's involved in like taking that and finding some way to take that data and like make it actually usable from a model improvement perspective and like rigorously it can go and like increase quality in some dimension so there's like i think there's even just like that business of like helping companies understand how they can actually make what's the value of their data and all that stuff i think is a is a really really difficult problem that feels like we're still early in trying to solve so and is like fundamentally correlated with like we can do that we'll see more model progress what 100 percent i think my takeaway from this conversation google the frontier commitment is real you guys got a lot of stuff going on i think you guys are doing a great job if people want to connect with you or or find you online we're sure you send them uh x i'll see you on x ping me i'm also on linked

in if you want to ask more boring questions so hi my friend thank you very much for doing this we'll do it together future i love it thank you for having me as a fun conversation

More episodes

More from The Pomp Podcast

View all episodes →