
technologyAug 29, 202541:28pending
Episode 57: AI Agents and LLM Judges at Scale: Processing Millions of Documents (Without Breaking the Bank)
About this episode
While many people talk about “agents,” Shreya Shankar (UC Berkeley) has been building the systems that make them reliable. In this episode, she shares how AI agents and LLM judges can be used to process millions of documents accurately and cheaply.
Drawing from work on projects ranging from databases of police misconduct reports to large-scale customer transcripts, Shreya explains the frameworks, error analysis, and guardrails needed to turn flaky LLM outputs into trustworthy pipelines.
We talk through:
- Treating LLM workflows as ETL pipelines for unstructured text
- Error analysis: why you need humans reviewing the first 50–100 traces
- Guardrails like retries, validators, and “gleaning”
- How LLM judges work — rubrics, pairwise comparisons, and cost trade-offs
- Cheap vs. expensive models: when to swap for savings
- Where agents fit in (and where they don’t)
If you’ve ever wondered how to move beyond unreliable demos, this episode shows how to scale LLMs to millions of documents — without breaking the bank.
LINKS
Shreya's website (https://www.sh-reya.com/)
DocETL, A system for LLM-powered data processing (https://www.docetl.org/)
Upcoming Events on Luma (https://lu.ma/calendar/cal-8ImWFDQ3IEIxNWk)
Watch the podcast video on YouTube (https://youtu.be/3r_Hsjy85nk)
Shreya's AI evals course, which she teaches with Hamel "Evals" Husain (https://maven.com/parlance-labs/evals?promoCode=GOHUGORGOHOME)
🎓 Learn more:
Hugo's course: Building LLM Applications for Data Scientists and Software Engineers (https://maven.com/s/course/d56067f338) — https://maven.com/s/course/d56067f338
Get full access to Vanishing Gradients at hugobowne.substack.com/subscribe
Drawing from work on projects ranging from databases of police misconduct reports to large-scale customer transcripts, Shreya explains the frameworks, error analysis, and guardrails needed to turn flaky LLM outputs into trustworthy pipelines.
We talk through:
- Treating LLM workflows as ETL pipelines for unstructured text
- Error analysis: why you need humans reviewing the first 50–100 traces
- Guardrails like retries, validators, and “gleaning”
- How LLM judges work — rubrics, pairwise comparisons, and cost trade-offs
- Cheap vs. expensive models: when to swap for savings
- Where agents fit in (and where they don’t)
If you’ve ever wondered how to move beyond unreliable demos, this episode shows how to scale LLMs to millions of documents — without breaking the bank.
LINKS
Shreya's website (https://www.sh-reya.com/)
DocETL, A system for LLM-powered data processing (https://www.docetl.org/)
Upcoming Events on Luma (https://lu.ma/calendar/cal-8ImWFDQ3IEIxNWk)
Watch the podcast video on YouTube (https://youtu.be/3r_Hsjy85nk)
Shreya's AI evals course, which she teaches with Hamel "Evals" Husain (https://maven.com/parlance-labs/evals?promoCode=GOHUGORGOHOME)
🎓 Learn more:
Hugo's course: Building LLM Applications for Data Scientists and Software Engineers (https://maven.com/s/course/d56067f338) — https://maven.com/s/course/d56067f338
Get full access to Vanishing Gradients at hugobowne.substack.com/subscribe
Get every episode summarized
Each time Vanishing Gradients publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Vanishing Gradients

The Rise of the AI Scientist
Vanishing Gradients
Sep 4, 20261:10:55failed

If Developers Build on Chinese Open-Weight Models, Who Leads AI?
Vanishing Gradients
Aug 3, 20261:18:16pending

Four Months Inside a Production AI Agent: What Real Users Changed
Vanishing Gradients
Jul 25, 20261:04:39pending

Building an Enterprise AI Agent for Healthcare
Vanishing Gradients
Jul 17, 20261:08:54pending