AI news story

Scaling Human Judgment: How Dropbox Uses LLMs to Improve Labeling for RAG Systems

To improve the relevance of responses produced by Dropbox Dash, Dropbox engineers began using LLMs to

  • AI
  • Source: InfoQ
  • Published: 2026-03-07

Editor's take

Dropbox engineers are now leveraging large language models (LLMs) to refine and augment human judgment in the labeling process for their Retrieval Augmented Generation (RAG) system, Dropbox Dash. This initiative aims to enhance the accuracy and relevance of the information retrieved and subsequently used to generate responses.

This development is significant as it addresses a critical bottleneck in RAG systems: the quality of training data. By using LLMs to assist human labelers, Dropbox is not only attempting to scale its data annotation efforts but also to improve the nuanced understanding required for high-fidelity RAG performance, impacting how effectively users can query and interact with their documents.

Future observations should focus on the measurable impact of LLM-assisted labeling on Dash's response quality metrics, such as retrieval precision and answer relevance. It will be crucial to see if this approach can consistently outperform purely human-driven labeling at scale and whether the LLMs themselves require ongoing fine-tuning to adapt to evolving data patterns and user needs within Dropbox's ecosystem.