AI news story

Alibaba's Qwen team makes AI models think deeper with new algorithm

Reinforcement learning hits a wall with reasoning models because every token gets the same reward. A new algorithm from Alibab…

  • AI
  • Source: The Decoder
  • Published: 2026-04-05

Editor's take

Alibaba's Qwen team has developed a novel algorithm that significantly enhances the reasoning capabilities of large language models by introducing a token-level reward weighting mechanism. This innovation addresses a fundamental limitation in applying reinforcement learning to complex AI reasoning tasks, where uniform rewards have previously hindered deeper, multi-step thought processes.

The breakthrough is significant because it promises to unlock more sophisticated problem-solving and longer, more coherent generative outputs from models like Qwen-VL. This could have broad implications across industries requiring advanced AI analysis, from scientific discovery to complex code generation, by enabling models to tackle more intricate chains of logic than previously feasible.

Future developments to monitor include independent verification of these performance gains across diverse reasoning benchmarks and the integration of this technique into other leading LLM architectures, such as those from Google or OpenAI. The extent to which this weighted token reward system can generalize beyond Qwen's current model families will be a key indicator of its lasting impact.