AI news story
Cloudflare’s new policy pushes AI companies to pay for publishers’ content
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI traini…
Editor's take
Cloudflare is now requiring AI companies to differentiate their web crawling activities, specifically distinguishing between data gathered for search engine indexing and that used for training large language models, with a September 15 deadline. Failure to comply will result in AI crawlers being blocked from accessing publisher content by default.
This policy shift directly impacts AI developers like OpenAI and Google, forcing them to reconsider their data acquisition strategies for model training. It addresses a long-standing grievance among publishers who have seen their content scraped without compensation or attribution to fuel the development of AI models that now compete with them. This is a crucial moment in the ongoing debate over fair use and compensation in the age of generative AI, potentially reshaping how AI companies access the open web.
The immediate next step to watch is how AI companies respond by September 15. Will they implement the technical separations Cloudflare demands, or will they face widespread content denial? Furthermore, this could catalyze a broader industry push for standardized APIs or licensing frameworks for AI training data, moving beyond the current ad-hoc scraping model and potentially leading to new revenue streams for content creators.