Resource

The Nonprofit Feeding the Entire Internet to AI Companies Highlight

posted on in: Quote, media and ai.

In his 2024 report, Baack, the ex-Mozilla researcher, pointed out that Common Crawl could require attribution whenever its scraped content is used. This would help publishers track the use of their work, including when it might appear in the training data of AI models that aren’t supposed to have access. This is a common requirement for open data sets and would cost Common Crawl nothing. I asked Skrenta if he had considered this. He told me he had read Baack’s report but didn’t plan on taking the suggestion, because it wasn’t Common Crawl’s responsibility. “We can’t police that whole thing,” he told me. “It’s not our job. We’re just a bunch of dusty bookshelves.”

— Alex Reisner

Replicated under Fair Use from The Nonprofit Feeding the Entire Internet to AI Companies by Alex Reisner.

Copy this link to share with your friends.

https://aramzs.xyz/resources/quotes/the-nonprofit-feeding-the-entire-internet-to-ai-companies/in-2024-reporthttpswwwmozillafoundationorgenresearchlibrarygenerative-ai-training-data-baack-ex-mozilla-45b74/