AI news story
Why Aren’t We Measuring How AI Affects Humans?
As AI systems become more capable, a lot of resources and effort are being put toward measuring their abilities. Researchers look at technical evaluation metrics, subject AIs to reasoning tests, track their throughput, and much more. But there’s one
Editor's take
The AI community is heavily invested in benchmarking model capabilities through technical evaluations and reasoning tests, yet neglects to systematically measure the real-world human impact of these systems.
This oversight is critical as AI, from Meta's Llama 3 to OpenAI's GPT-4, increasingly permeates daily life, influencing decisions and shaping interactions. Without dedicated metrics for human well-being, fairness, and societal effects, we risk deploying powerful technologies with unforeseen negative consequences, potentially exacerbating existing inequalities or creating new ones.
Future research must prioritize developing and standardizing methodologies to quantify AI's influence on human users and society. Key questions include how to reliably attribute observed societal shifts to specific AI deployments and whether organizations like the Partnership on AI will champion such human-centric evaluation frameworks.
Signal score: 5
This event was corroborated by 6 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by IEEE Spectrum. Read the original article at IEEE Spectrum.