
About this episode
Model sizes are crazy these days with billions and billions of parameters. As Mark Kurtz explains in this episode, this makes inference slow and expensive despite the fact that up to 90%+ of the parameters don't influence the outputs at all.
Mark helps us understand all of the practicalities and progress that is being made in model optimization and CPU inference, including the increasing opportunities to run LLMs and other Generative AI models on commodity hardware.
Get every episode summarized
Each time Changelog Master Feed publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Changelog Master Feed

Forking Cal.com to closed source (Changelog Interviews #685)
Changelog Master Feed
Sep 3, 20261:54:32pending

Postgres at PlanetScale (Changelog Interviews #684)
Changelog Master Feed
Aug 25, 20261:42:17pending

Canary tokens and digital tripwires (Changelog Interviews #683)
Changelog Master Feed
Jul 21, 20262:06:48pending

From open source hits to OpenAI (Changelog Interviews #682)
Changelog Master Feed
Jun 5, 20261:46:28pending