— Alex ReisnerI also discovered more than 1,000 other domains that produce this incorrect “no captures” result for at least several of the crawls, and most of these domains belong to publishers, including the BBC, Reuters, The New Yorker, Wired, the Financial Times, The Washington Post, and, yes, The Atlantic. According to my research and Common Crawl’s own disclosures, the companies behind each of these publications have sent legal requests to the nonprofit. At least one publisher I spoke with told me that it had used this search tool and concluded that its content had been removed from Common Crawl’s archives.
Replicated under Fair Use from The Nonprofit Feeding the Entire Internet to AI Companies by Alex Reisner.