US Publishers Demand Common Crawl Stop Scraping Their Content via @sejournal, @MattGSouthern

General News

Summary

Digital Content Next has sent Common Crawl a cease-and-desist letter demanding that it stop scraping publisher content and remove protected material from its datasets. The dispute centers on whether Common Crawl can keep previously collected pages in its public archive and share them with AI companies. Publishers argue that copyright should require permission first, while Common Crawl says it removes affected URLs from future crawls and cannot edit published archive files without breaking integrity. The issue affects publishers that want to block training bots and could reshape how AI data collection is handled across the web.

Classifications

industries
No industries detected
applications
No applications detected

AskAI Classifications

Labels
No AI classifications detected

Linked Companies