
Decoding LLM Quality: From Unit Testing to User Feedback
About this episode
Dive deep into the nuances of measuring the quality of Large Language Model (LLM) prompts, as we explore past methodologies, and evaluate both qualitative assessments and large-scale testing techniques. Join us as we discuss the challenges of traditional metrics, the role of user feedback, and brainstorm new ways to gauge generative model performance.
—
Continue listening to The Prompt Desk Podcast for everything LLM & GPT, Prompt Engineering, Generative AI, and LLM Security.
Check out PromptDesk.ai for an open-source prompt management tool.
Check out Brads AI Consultancy at bradleyarsenault.me.
Add Justin Macorin and Bradley Arsenault on LinkedIn.
Please fill out our listener survey here to help us create a better podcast: https://docs.google.com/forms/d/e/1FAIpQLSfNjWlWyg8zROYmGX745a56AtagX_7cS16jyhjV2u_ebgc-tw/viewform?usp=sf_link
Hosted on Ausha. See ausha.co/privacy-policy for more information.
Get every episode summarized
Each time The Prompt Desk publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The Prompt Desk

What we learned about LLM’s in a year
The Prompt Desk

Validating Inputs with LLMs
The Prompt Desk

Why you can't automate everything with LLMs
The Prompt Desk

Data Preparation Best Practices for Fine Tuning
The Prompt Desk