Skip to content
TrackPodcasts
technologyNov 26, 202457:46pending

When AI Benchmarks Lie: A Better Way to Evaluate Ft. Chris Hay

About this episode

This episode explores the world of AI evaluation, with insights from Chris Hay on why benchmarks are "stupid" and how to effectively evaluate AI models. Get the tools pip install tool-use-ai Check out Chris' Channel https://www.youtube.com/@chrishayuk Links https://github.com/EleutherAI/lm-eval... Lessons from the Trenches on Reproducible Evaluation of Language Models - https://arxiv.org/pdf/2405.14782

https://github.com/confident-ai/deepeval Connect with us https://x.com/ToolUseAI

https://x.com/MikeBirdTech

https://x.com/FieroTy

https://x.com/chrishayuk *The opinions of Chris are purely Chris's opinions and don't represent the opinions of his employer

Get every episode summarized

Each time Tool Use - AI Conversations publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

When AI Benchmarks Lie: A Better Way to Evaluate Ft. Chris Hay

Tool Use - AI Conversations

0:00
57:46

More episodes

More from Tool Use - AI Conversations

View all episodes →