Skip to content
TrackPodcasts
technologyJan 23, 202446:16pending

Collaboration & evaluation for LLM apps (Practical AI #253)

About this episode

Small changes in prompts can create large changes in the output behavior of generative AI models. Add to that the confusion around proper evaluation of LLM applications, and you have a recipe for confusion and frustration. Raza and the Humanloop team have been diving into these problems, and, in this episode, Raza helps us understand how non-technical prompt engineers can productively collaborate with technical software engineers while building AI-driven apps.

Get every episode summarized

Each time Changelog Master Feed publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Collaboration & evaluation for LLM apps (Practical AI #253)

Changelog Master Feed

0:00
46:16

More episodes

More from Changelog Master Feed

View all episodes →