
When AI Benchmarks Lie: A Better Way to Evaluate Ft. Chris Hay
About this episode
This episode explores the world of AI evaluation, with insights from Chris Hay on why benchmarks are "stupid" and how to effectively evaluate AI models. Get the tools pip install tool-use-ai Check out Chris' Channel https://www.youtube.com/@chrishayuk Links https://github.com/EleutherAI/lm-eval... Lessons from the Trenches on Reproducible Evaluation of Language Models - https://arxiv.org/pdf/2405.14782
https://github.com/confident-ai/deepeval Connect with us https://x.com/ToolUseAI
https://x.com/chrishayuk *The opinions of Chris are purely Chris's opinions and don't represent the opinions of his employer
Get every episode summarized
Each time Tool Use - AI Conversations publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Tool Use - AI Conversations

Hermes Agent has won. Here's why
Tool Use - AI Conversations

How To Make Your Websites Fully Autonomous (ft rtrvr)
Tool Use - AI Conversations

How To Make Your A.I. Product Go Viral (ft Mano Tsiris)
Tool Use - AI Conversations

How To Build a Hybrid AI System with Any-LLM (ft Nathan Brake)
Tool Use - AI Conversations