The Danger of “Just Scrape It” in AI Strategy | HackerNoon
By ai_poster · 7/28/2026, 2:13:47 AM
In AI, the phrase "just scrape it" creates products whose foundations are difficult to defend, as collecting data first and asking about provenance, rights, and consent later has become a dangerous habit. In 2025, Anthropic agreed to pay $1.5 billion to settle a class action brought by authors — the largest publicly reported copyright settlement in US history, covering roughly half a million books at an implied rate of around $3,000 per work. The court granted final consideration of the deal in 2026, and the terms include destruction of the pirated dataset itself. The judge ruled that training on copyrighted books can be transformative fair use, but what was not fair use was the acquisition: downloading over seven million books from pirate libraries. Music publishers filed a $3.1 billion claim against the same company in January 2026 built on the same acquisition theory. As of mid-2026, more than seventy AI copyright cases are active or recently resolved across US and international courts, with cumulative claimed damages estimated beyond $50 billion. In *Thomson Reuters v. Ross Intelligence*, a US court rejected fair use for training on proprietary legal content outright. Data provenance—the record of where data came from, what rights attached to it, how it was transformed, and whether it can be reused responsibly—is not administrative overhead.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.