Skip to content
TrackPodcasts
scienceApr 4, 202526:11pending

AI Benchmark Deep Dive: Gemini 2.5 and Humanity's Last Exam

Deep Papers

About this episode

This week we talk about modern AI benchmarks, taking a close look at Google's recent Gemini 2.5 release and its performance on key evaluations, notably  Humanity's Last Exam (HLE). In the session we covered Gemini 2.5's architecture, its advancements in reasoning and multimodality, and its impressive context window. We also talked about how benchmarks like HLE and ARC AGI 2 help us understand the current state and future direction of AI.


Join us for the next live recording, or check out the latest AI research

Learn more about AI observability and evaluation, join the Arize AI Slack community or get the latest on LinkedIn and X.

Get every episode summarized

Each time Deep Papers publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

AI Benchmark Deep Dive: Gemini 2.5 and Humanity's Last Exam

Deep Papers

0:00
26:11

More episodes

More from Deep Papers

View all episodes →