
About this episode
Clockwork began with a narrow goal—keeping clocks synchronized across servers—but soon realized that its precise latency measurements could reveal deeper data center networking issues. This insight led the company to build a hardware-agnostic monitoring and remediation platform capable of automatically routing around faults. Today, Clockwork’s technology is especially valuable for large GPU clusters used in training LLMs, where communication efficiency and reliability are critical. CEO Suresh Vasudevan explains that AI workloads are among the most demanding distributed applications ever, and Clockwork provides building blocks that improve visibility, performance and fault tolerance. Its flagship feature, FleetIQ, can reroute traffic around failing switches, preventing costly interruptions that might otherwise force teams to restart training from hours-old checkpoints. Although the company originated from Stanford research focused on clock synchronization for financial institutions, the team eventually recognized that packet-timing data could underpin powerful network telemetry and dynamic traffic control. By integrating with NVIDIA NCCL, TCP and RDMA libraries, Clockwork can not only measure congestion but also actively manage GPU communication to enhance both uptime and training efficiency.
Learn more from The New Stack about the latest in Clockwork:
Clockwork’s FleetIQ Aims To Fix AI’s Costly Network Bottleneck
What Happens When 116 Makers Reimagine the Clock?
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Get every episode summarized
Each time The New Stack Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The New Stack Podcast

How Microsoft is governing thousands of Kubernetes clusters without manual inter...
The New Stack Podcast

Why long-running AI agents break on HTTP and how Ably is fixing it
The New Stack Podcast

Why the Linux Foundation adopted MCP, with Jim Zemlin and Mazin Gilbert
The New Stack Podcast

Fresh data has us asking, does AI demand Kubernetes?
The New Stack Podcast