
Hyperscaling natural language processing
About this episode
In this episode of the Data Exchange I speak with Edmon Begoli, Chief Data Architect at Oak Ridge National Laboratory (ORNL). Edmon has developed and implemented large-scale data applications on systems like Open MPI, Hadoop/MapReduce, Apache Calcite, Apache Spark, and Akka. Most recently he has been building large-scale machine learning and natural language applications with Ray, a distributed execution framework that makes it easy to scale machine learning and Python applications.
Our conversation included a range of topics, including:
- Edmon’s role at the ORNL and his experience building applications with Hadoop and Spark.
- What is distributed online learning?
- Why they started using Ray to build distributed online learning applications.
- Two important use cases: suicide prevention among US veterans and infectious disease surveillance.
Detailed show notes can be found on The Data Exchange web site.
Join Michael Jordan, Manuela Veloso, Azalia Mirhoseini, Zoubin Ghahramani, Wes McKinney, Ion Stoica, Gaël Varoquaux, and many other speakers at the first Ray Summit In San Francisco, May 27-28. Tickets start at $200.
Get every episode summarized
Each time The Data Exchange with Ben Lorica publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The Data Exchange with Ben Lorica

Your AI Agent Is Costing You More Than You Think
The Data Exchange with Ben Lorica

Reasoning Doesn't Start With Language
The Data Exchange with Ben Lorica

An Agent Is Just an LLM in a For-Loop
The Data Exchange with Ben Lorica

The Bloomberg Terminal for AI Compute
The Data Exchange with Ben Lorica