AI news story

Porn, dog poo and social media snaps: the ‘taskers’ scraping the internet for Meta-owned AI firm

Scale AI gig workers describe desperation of using people’s personal profiles and copyrighted work to train AI Tens of tho…

  • AI
  • Source: The Guardian AI
  • Published: 2026-04-07

Editor's take

Scale AI, a firm partially owned by Meta, has reportedly employed thousands of gig workers to scrape publicly available data, including personal social media content and copyrighted material, for AI training purposes. This practice highlights the ethically precarious foundation upon which many large language models are built, raising questions about consent and intellectual property rights for the vast datasets used to develop systems like Meta's Llama models.

The reliance on such "taskers" underscores the immense, often invisible, labor required to fuel AI development, while simultaneously exposing the potential for misuse of personal information. This situation is particularly relevant as AI companies face increasing scrutiny over data privacy and the origins of their training data, following similar criticisms leveled against OpenAI's ChatGPT.

Future developments to monitor include how regulatory bodies respond to these data scraping allegations and whether Scale AI or Meta implement more robust ethical oversight for data acquisition. The long-term impact on public trust in AI, and the willingness of individuals to share content online, will also be crucial indicators.