Skip to content
TrackPodcasts
newsJan 22, 20249:15pending

Token Extravaganza: Unveiling the World's Largest Open-Source LLM Dataset - 3T Tokens

Strict Scrutiny

About this episode

In this episode, we explore an extravaganza of linguistic data as the world's largest open-source LLM dataset, featuring an unprecedented 3 trillion tokens, is unveiled, opening new frontiers in language model research.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Get every episode summarized

Each time Strict Scrutiny publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Token Extravaganza: Unveiling the World's Largest Open-Source LLM Dataset - 3T Tokens

Strict Scrutiny

0:00
9:15

More episodes

More from Strict Scrutiny

View all episodes →