Skip to content
TrackPodcasts
technologyDec 31, 202340:16pending

Exposing the Rotten Reality of AI Training Data

About this episode

In a report released December 20, 2023, the Stanford Internet Observatory said it had detected more than 1,000 instances of verified child sexual abuse imagery in a significant dataset utilized for training generative AI systems such as Stable Diffusion 1.5.

This troubling discovery builds on prior research into the “dubious curation” of large-scale datasets used to train AI systems, and raises concerns that such content may contributed to the capability of AI image generators in producing realistic counterfeit images of child sexual exploitation, in addition to other harmful and biased material. 

Justin Hendrix spoke the report’s author, Stanford Internet Observatory Chief Technologist David Thiel.

Get every episode summarized

Each time The Tech Policy Press Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Exposing the Rotten Reality of AI Training Data

The Tech Policy Press Podcast

0:00
40:16

More episodes

More from The Tech Policy Press Podcast

View all episodes →