Skip to content
TrackPodcasts
technologyOct 14, 202557:37pending

Dataflow Computing for AI Inference with Kunle Olukotun - #751

About this episode

In this episode, we're joined by Kunle Olukotun, professor of electrical engineering and computer science at Stanford University and co-founder and chief technologist at Sambanova Systems, to discuss reconfigurable dataflow architectures for AI inference. Kunle explains the core idea of building computers that are dynamically configured to match the dataflow graph of an AI model, moving beyond the traditional instruction-fetch paradigm of CPUs and GPUs. We explore how this architecture is well-suited for LLM inference, reducing memory bandwidth bottlenecks and improving performance. Kunle reviews how this system also enables efficient multi-model serving and agentic workflows through its large, tiered memory and fast model-switching capabilities. Finally, we discuss his research into future dynamic reconfigurable architectures, and the use of AI agents to build compilers for new hardware. The complete show notes for this episode can be found at https://twimlai.com/go/751.

Get every episode summarized

Each time The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Dataflow Computing for AI Inference with Kunle Olukotun - #751

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

0:00
57:37

More episodes

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

View all episodes →