
Exposing the Rotten Reality of AI Training Data
About this episode
In a report released December 20, 2023, the Stanford Internet Observatory said it had detected more than 1,000 instances of verified child sexual abuse imagery in a significant dataset utilized for training generative AI systems such as Stable Diffusion 1.5.
This troubling discovery builds on prior research into the “dubious curation” of large-scale datasets used to train AI systems, and raises concerns that such content may contributed to the capability of AI image generators in producing realistic counterfeit images of child sexual exploitation, in addition to other harmful and biased material.
Justin Hendrix spoke the report’s author, Stanford Internet Observatory Chief Technologist David Thiel.
Get every episode summarized
Each time The Tech Policy Press Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The Tech Policy Press Podcast

AI Governance Reaches Crisis Point Ahead of Trump-Xi Summit
The Tech Policy Press Podcast

Across the US, School Districts Wrestle With Classroom AI Policy
The Tech Policy Press Podcast

Rep. Zoe Lofgren on FISA, Surveillance and the Fourth Amendment
The Tech Policy Press Podcast

Virginia Eyes AI Chatbot Rules as Observable Risks Warrant Action
The Tech Policy Press Podcast