
Measuring Brilliance in Generative AI: Perplexity, Precision, and Faithfulness
About this episode
We unpack how to evaluate AI that writes and creates, not just predicts. Why perplexity captures surprise, why a low perplexity score isn’t a guarantee of correctness, and how precision, recall, and the harmonic F1 balance model performance. We compare BLEU and ROUGE, explore Retrieval-Augmented Generation to stay faithful to private data, and discuss out-of-domain challenges, agentic AI, and the guardrails shaping the future.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
Sponsored by Embersilk LLC
Get every episode summarized
Each time Intellectually Curious publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Intellectually Curious

Claude Commerce: The One-Brain AI Reimagining Digital Shopping
Intellectually Curious

GPT-6 Astra: The Autonomous AI Operator Redefining Science and Workflows
Intellectually Curious

Zero-Friction Innovation: AI, Activation Energy, and the Long-Tail Frontier
Intellectually Curious

Momentum Exchange Tethers and Orbital Skyhooks
Intellectually Curious