AI news story
News Organizations Push Back Against Web Archive Used For AI
Major news organizations, including CNN, NBC and USA Today, have joined an effort to curb the storage of their content in a web archive used by artificial intelligence companies for training chatbots.
Editor's take
News publishers, including CNN, NBC, and USA Today, are taking action to restrict the use of their content within the Internet Archive's Wayback Machine by AI developers for model training. This move stems from concerns over unauthorized data scraping and the potential dilution of their intellectual property value as AI models increasingly incorporate copyrighted material without compensation.
The significance lies in the ongoing tension between content creators and AI developers over data rights. This collective pushback by prominent media outlets highlights a growing industry-wide demand for clearer frameworks around fair use and licensing for AI training data, potentially impacting the accessibility and diversity of information available for future model development.
Future developments to monitor include the legal precedents set by any potential lawsuits, the specific technical mechanisms news organizations might employ to block archive access, and whether AI companies will proactively develop more ethically sourced or licensed datasets in response. The outcome could reshape how AI models are trained and the economic models for news content in the digital age.
Signal score: 5
This event was corroborated by 2 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by Bloomberg. Read the original article at Bloomberg.