Skip to content
TrackPodcasts
scienceJul 13, 20265:21pending

Measuring Brilliance in Generative AI: Perplexity, Precision, and Faithfulness

About this episode

We unpack how to evaluate AI that writes and creates, not just predicts. Why perplexity captures surprise, why a low perplexity score isn’t a guarantee of correctness, and how precision, recall, and the harmonic F1 balance model performance. We compare BLEU and ROUGE, explore Retrieval-Augmented Generation to stay faithful to private data, and discuss out-of-domain challenges, agentic AI, and the guardrails shaping the future.


Note:  This podcast was AI-generated, and sometimes AI can make mistakes.  Please double-check any critical information.

Sponsored by Embersilk LLC

Get every episode summarized

Each time Intellectually Curious publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Measuring Brilliance in Generative AI: Perplexity, Precision, and Faithfulness

Intellectually Curious

0:00
5:21

More episodes

More from Intellectually Curious

View all episodes →