
technologyJan 23, 202446:16pending
Collaboration & evaluation for LLM apps (Practical AI #253)
About this episode
Small changes in prompts can create large changes in the output behavior of generative AI models. Add to that the confusion around proper evaluation of LLM applications, and you have a recipe for confusion and frustration. Raza and the Humanloop team have been diving into these problems, and, in this episode, Raza helps us understand how non-technical prompt engineers can productively collaborate with technical software engineers while building AI-driven apps.
Get every episode summarized
Each time Changelog Master Feed publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Changelog Master Feed

Forking Cal.com to closed source (Changelog Interviews #685)
Changelog Master Feed
Sep 3, 20261:54:32pending

Postgres at PlanetScale (Changelog Interviews #684)
Changelog Master Feed
Aug 25, 20261:42:17pending

Canary tokens and digital tripwires (Changelog Interviews #683)
Changelog Master Feed
Jul 21, 20262:06:48pending

From open source hits to OpenAI (Changelog Interviews #682)
Changelog Master Feed
Jun 5, 20261:46:28pending