
Aug 5, 2026
21 Million Songs Were Stolen to Train AI
Somewhere in your playlist right now is a song sitting inside an AI training dataset and the artist never signed off on it. The Atlantic's
Alex Reisner built a free tool that lets anyone search 21 million songs pulled into datasets shared across the AI-development world. But the real story isn't "AI companies are stealing your stuff." It's a chain-of-custody problem and if you're a founder building on someone else's model, you may have already inherited legal exposure you can't even see.
In this episode we break down the AI Watchdog investigation, the $1.5 BILLION Anthropic settlement (the largest copyright recovery in history), the labels coming for 61,000+ recordings, and the 3-step "dataset receipts" system that protects you before you build.
⏱️ TIMESTAMPS
0:00 – The song hiding in a spreadsheet
0:45 – The question no one's asking: where did AI's homework come from?
1:05 – Alex Reisner, Books3, and 183,000 pirated books
1:40 – 7.5M books, Hollywood scripts, then music
2:15 – 4 datasets, 21 million songs, and the names inside
2:50 – The truth: being in the pile ≠ being used
3:30 – The $1.5 billion Anthropic settlement
4:00 – Universal & Sony vs. the AI music companies
4:25 – Why this should terrify Gen Z founders
5:00 – Clean data as a competitive advantage
5:15 – STEAL THIS MOVE: the 3-step dataset receipts system
6:00 – Who Alex Reisner really is + your move
👉 Subscribe for more Gen ZEO Playbook breakdowns: https://www.youtube.com/@GenZEOPlaybook
Comment: which AI tool in your stack have you NEVER checked the data source on?
No comments yet. Be the first to say something!