
Austin Lyons on NVDA, ARM, Google TurboQuant & New AI Innovations
About this episode
Austin Lyons says Nvidia (NVDA) and other GPU makers are "changing the dynamics" when it comes to how their chips improve LLM inference. Arm Holdings (ARM) is capitalizing on CPUs, something Austin considers a bottleneck for GPUs. He explains how Arm aims to fill in the efficiency gap. Turning to the AI memory space, Austin says Alphabet's (GOOGL) TurboQuant isn't a replacement for Micron (MU) and SanDisk (SNDK) but rather a tool to utilize their hardware more.
======== Schwab Network ========
Empowering every investor and trader, every market day.
Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribe
Download the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185
Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7
Watch on Sling - https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watch
Watch on Vizio - https://www.vizio.com/en/watchfreeplus-explore
Watch on DistroTV - https://www.distro.tv/live/schwab-network/
Follow us on X – / schwabnetwork
Follow us on Facebook – / schwabnetwork
Follow us on LinkedIn - / schwab-network
About Schwab Network - https://schwabnetwork.com/about
Get every episode summarized
Each time Schwab Network publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Transcript ready
140 searchable segments. Every word is indexed and playable.
Full transcript
Schwab Network — Austin Lyons on NVDA, ARM, Google TurboQuant & New AI Innovations. Machine-transcribed; use the interactive transcript above to jump the player to any line.
So joining us now is Austin Lyons, Senior Analysts at Creative Strategies, with a closer look at some of the latest tech developments. And so Austin, thank you so much for joining us on this Friday. Always appreciate your time and all of your thoughts. And you have said that the GPUs for everything era is maybe over. And so what does that mean for say, like, in video's mode, as more and more of these AI-ASIC players like Groc and Serupras try to slot into inference? Sure. Yes. Historically, NVIDIA was the first, and they were the best, and they were the only. We saw AMD come and try to play in the GPU space. But now, NVIDIA's GTC, and the narrative was always just, like, by GPUs, they're the best for everything. They're the most flexible. Last week at GTC, NVIDIA and Jensen Wong showed that if you take LLM inference, which is the defining workload of our era, you can actually break it down into sub-tasks if you want to think about it. For example, pre-fill and decode is what they call them. And you can run some of these workloads on GPUs, but actually to get really fast, like
what they call ultra low latency and high interactivity for the user, so really fast, clawed coding, for example. You can actually slot in Groc chips and unlock this new performance level. And so, you know, that sort of changes the dynamics, because if you can slot in Groc, or cerebral, it starts to say, like, oh, who else could slot in here and why should it be NVIDIA? Now, NVIDIA, for their mode, they're trying to take it up a level higher and say, hey, it's not about the individual components, it's about the whole system, all of the racks put together, who can give you the best optimized inference across the whole system. So NVIDIA will continue to try to say, once you put it all together, it's our data center, something branded NVIDIA that matters, even if you slot in other people, we believe that we can optimize every little piece and still unlock the best inference. All right, awesome. One of the companies that's trying to slot themselves in there's ARM, they're shifting from, like, licensing IP purely to actually selling these full racks for this AI.
Why is that important? How is it meaningful to earnings power? Yeah, totally. So in the data center that Jensen showed, Jensen and NVIDIA actually had racks of CPUs that were ARM licensed CPUs that NVIDIA makes, and the Jensen's rationale was, hey, actually CPUs are becoming the bottleneck for GPUs, because there's still a lot of work that needs to be done on the CPU to keep the GPUs fed. ARM saw this as an opportunity for them to also not just license their intellectual property to people like NVIDIA for those CPUs, but to actually make silicon themselves. And this is a big opportunity, huge opportunity, historically ARM for the last 35 years has only licensed their intellectual property. And then within the last handful of years, they took a step further and they, instead of just licensing the CPU design, they said, hey, we've got this thing called the compute subsystem. So we can give you more, so the CPU and some of the system IP around it.
And those have, they make 5% royalties or 10% royalties if they do more for you. But when they move into silicon, now they can sell the chip, they can make the money. Of course, the gross margins on a chip might only be 50% whereas on that IP that they're licensing, the gross margins were like 97%. So their blended margins will come down, but it'll be a lot of incremental margin dollars that are at play. And that's what investors are excited about for ARM is just, it's obviously a huge market to be able to sell into the data center to sell silicon, gives them an opportunity to drive a lot of top line and bottom line revenue. Yeah, and I mean, the market definitely agrees with you because we saw a huge move in ARM following that news, of course, but we have to discuss now to shift a bit the turbo quant announcement, I mean, which hit Micron pretty hard on fears of reduced high bandwidth memory demand. And so we've seen this not just to impact Micron, also Sandisk, a lot of these memory exposed names that have been a huge, I mean, boon for the overall AI trade.
So what did you make of some of the developments that we've seen just in the last few several sessions with the memory names? Yeah, Jenny, the, so, you know, memory's been super hot memory and storage over the past six months, these companies, they're stock of just skyrocketed. And obviously, there's always questions around memory. Is it cyclical? Is it, is this the top? And so I think when the market saw this turbo quant algorithm, the blog from Google that basically said, hey, all of this, like KV cash is what it's called, but it's just like all this working memory that's stored in this expensive high bandwidth memory, that can actually be compressed and stored more efficiently. And so the market saw that is like, oh, we must not need as much HBM anymore. Maybe I should get out of this trade. Maybe this is the top and I should sell and just take my profits. Now of course, I actually think my interpretation, and I think which would be the right interpretation in the long run, we can debate it, is that by being able to store more memory efficiently,
that just means we can do more. So you know, as over the past couple of years, when we first started with chat, GPT, like it was kind of helpful and kind of useful, but once it could start to reason, once we could, it had longer context, you could give it like all these PDFs to read, it got more and more valuable. And ultimately, what this algorithm from Google does is just unlock, hey, you can now attach even more PDFs, you can give it even more context. And so you can do more, and I think that's the right framing here, is that we'll be able to do more with the hardware that's already deployed. And so I don't think it is, it means, you know, Nvidia or AMD or anyone's going to start buying less memory. I think it just means they're going to continue to design their chips with as much memory as possible. And AI model labs will be able to do more with that hardware. Yeah, and that's one of those paradoxes, right? What is it? Javans or something like that? Yes. Exactly. But I do want to push back on so we can debate it. So here's my final question for you. What do you say to those who say, hey, all of this is great academically.
It reads well on paper. But the NASDAQ hasn't made a new all-time high since October of last year. It's now over 11% off its highs, and things seem to be accelerating in the other direction. Yes. So, great question, Alex. We're in this, like, churn moment right now where academically, there's a lot of really interesting things taking off, agent AI, cloud code. Now the barrier for writing software, anyone can write software now. All you have to do is speak in natural language, and legitimately you can make amazing apps. My kids make video games just on demand. They just have an idea. I want to make plant tycoon, and they make it, right? And this will really unlock productivity and the amount of what employees can do. If you hate writing expense reports and wasting your time doing that, awesome. Use cloud code and write software to do that for you now. Anyone can do that. And yet, when we zoom out like that productivity, it still hasn't quite manifested in companies
unlocking new lines of revenue or reducing their costs, right? And so I think the market is just waiting for that to get sorted out where it's like, hey, we've got this next level of AI. It's super useful. It's pretty expensive. CapEx is through the roof. But when will we actually see that generate new revenue for more than just the hyperscalers, but like the John Deers or caterpillars of the word, or insurance companies of the world? When will this actually benefit them? So that's my take on where we are, is we're kind of waiting for that adoption to happen. And for investors to see like, oh, yes, this is benefiting everyone. And I totally agree. And I think it's the fine line between, you know, like, investing in these companies and understanding what these companies do. And it does feel like there's a bit of a disconnect right now because I agree, like, you see the turbo-quant news. And I was a bit confused by the market's reaction, I mean, huge reaction. And I think there's still a lot of like proof of concept that we need to see play out in real time to obviously see maybe more favorability come into this space.
But we so appreciate your time today, Austin Lines, always great conversation. It's senior analyst at Creative Strategies.
More episodes
More from Schwab Network

JNJ Balances Growth, Patent Transitions, and Legal Overhangs
Schwab Network

nCino (NCNO) CEO on Earnings, AI in Banking & Overcoming "SaaS-pocalypse"
Schwab Network

The Big 3: NET, TWLO, ASTS
Schwab Network

Tie Lasater on Inflation Floors, Energy Risk, and Private Credit Stress
Schwab Network