AI news story

AI models fail at robot control without human-designed building blocks but agentic scaffolding closes the gap

A new framework from Nvidia, UC Berkeley, and Stanford systematically tests how well AI models can control robots throug…

  • Robotics
  • Source: The Decoder
  • Published: 2026-04-02

Editor's take

Nvidia, UC Berkeley, and Stanford researchers demonstrated that advanced AI models, including large language models, struggle with direct robot control without pre-defined functional components. Their systematic evaluation revealed that even sophisticated models like Google's PaLM 2 or OpenAI's GPT-4, when tasked with generating robotic manipulation code, produced insufficient or erroneous outputs without explicit, human-engineered "building blocks" representing object properties or task primitives.

This finding underscores a critical bottleneck in achieving truly generalizable robotic intelligence. Current AI excels at pattern recognition and natural language understanding, but translating that into precise, real-world physical actions remains a significant hurdle. The reliance on human-designed abstractions highlights the need for either more intuitive ways for AI to acquire common-sense physical reasoning or improved methods for integrating learned behaviors with foundational robotic understanding, impacting industries from manufacturing to logistics.

Future developments should focus on how these "agentic scaffolding" techniques, which involve adaptive test-time computation, can be integrated into more autonomous learning pipelines. It will be crucial to observe whether this approach can scale to more complex tasks and a wider variety of robotic hardware without requiring extensive human annotation for each new scenario, moving towards more robust and adaptable robotic agents.