AI news story

LLM text data is drying up, but Meta points to unlabeled video as the next massive training frontier

A research team from Meta FAIR and New York University trained a multimodal AI model from scratch and found that several com…

  • LLMs
  • Source: The Decoder
  • Published: 2026-03-08

Editor's take

Meta researchers have demonstrated a multimodal AI model trained on unlabeled video data, challenging conventional wisdom regarding LLM development. This work suggests that vast quantities of readily available, unstructured video content could serve as a substantial new resource for training future AI systems, potentially alleviating concerns about the scarcity of high-quality text datasets.

The implications are significant for AI development, particularly as models like Meta's Llama 2 and OpenAI's GPT-4 continue to push the boundaries of capability. If unlabeled video proves to be a viable and scalable training source, it could democratize access to powerful AI models by reducing reliance on expensive and curated text corpora. This shift could also accelerate progress in areas requiring a deeper understanding of the physical world, moving beyond purely textual reasoning.

Future research should focus on the efficacy of different video encoding and processing techniques for AI training, and whether this approach can yield models with comparable or superior reasoning abilities to those trained on text. It will be crucial to observe if other major AI labs, such as Google DeepMind or Anthropic, adopt similar strategies and what new architectural innovations emerge to leverage this multimodal frontier.