Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets. DCN CEO Jason Kint announced the legal notice in a blog post , and Press Gazette reported additional details from the letter this week. Common Crawl has crawled several billion new pages each month since 2007 to build a free public archive. That archive has been used to train many of the AI models in use today. OpenAI’s GPT-3 paper listed filtered Common Crawl as 60% of the model’s training mix. The dispute matters for any site that blocks AI crawlers. Blocking Common Crawl’s crawler, CCBot, stops future collection but doesn’t touch content already in the archive, which anyone can still download.…