US Publishers Demand Common Crawl Stop Scraping Their Content via @sejournal, @MattGSouthern
Summary
Digital Content Next has sent Common Crawl a cease-and-desist letter demanding that it stop scraping publisher content and remove protected material from its datasets. The dispute centers on whether Common Crawl can keep previously collected pages in its public archive and share them with AI companies. Publishers argue that copyright should require permission first, while Common Crawl says it removes affected URLs from future crawls and cannot edit published archive files without breaking integrity. The issue affects publishers that want to block training bots and could reshape how AI data collection is handled across the web.
Classifications
industries
No industries detected
applications
No applications detected
AskAI Classifications
Labels
No AI classifications detected