
Data Augmentation in Natural Language Processing
About this episode
This week’s guests are Steven Feng, Graduate Student and Ed Hovy, Research Professor, both from the Language Technologies Institute of Carnegie Mellon University. We discussed their recent survey paper on Data Augmentation Approaches in NLP (GitHub), an active field of research on techniques for increasing the diversity of training examples without explicitly collecting new data. One key reason why such strategies are important is that augmented data can act as a regularizer to reduce overfitting when training models.
Subscribe: Apple • Android • Spotify • Stitcher • Google • RSS.
Detailed show notes can be found on The Data Exchange web site.
Subscribe to The Gradient Flow Newsletter.
Get every episode summarized
Each time The Data Exchange with Ben Lorica publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The Data Exchange with Ben Lorica

Your AI Agent Is Costing You More Than You Think
The Data Exchange with Ben Lorica

Reasoning Doesn't Start With Language
The Data Exchange with Ben Lorica

An Agent Is Just an LLM in a For-Loop
The Data Exchange with Ben Lorica

The Bloomberg Terminal for AI Compute
The Data Exchange with Ben Lorica