Suno trained its music generator on millions of songs it didn't license—YouTube Music, Deezer, Genius—scraped them all, fed them into a machine. Now produces new songs on command.
When the training data surfaced via a hacking incident, everyone acted shocked. Suno had hidden this, claiming to operate within fair use while actively concealing what fair use actually looked like in their case.
This plot is not new. In 2005, Google began scanning millions of books without permission. Publishers sued—but by 2015, a federal appeals court in New York decided the copying was transformative and therefore fair use.
Suno's position is structurally identical. Millions of songs went in and no songs came out reconstructed in the model—the training data is not the product. The product is a new system that learned from patterns in the input without reproducing them. It is Google Books with audio, and the legal architecture is the same.
The interesting part isn't whether scraping happens. It is whether legal liability depends on whether everyone does it.
What matters now is whether courts treat the precedent as precedent or whether AI gets a special exemption that books never received. If Meta, OpenAI, and other music generators all scraped too, then Suno goes from defendant to industry standard. If they licensed their data, Suno chose the cheaper route and now faces discovery. That distinction shapes every field where the cheapest path and the licensed path diverge—most people never choose the expensive route because they're waiting to see which one survives.