AI news story

Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena leaderboard. The model supports 16 languages and lets…

  • AI
  • Source: The Decoder
  • Published: 2026-07-21

Editor's take

Alibaba's Qwen Audio 3.0 TTS Plus has achieved the top position on Artificial Analysis' Speech Arena leaderboard for text-to-speech models. This advancement is significant as it showcases a leap in natural language generation for audio, offering enhanced controllability over vocal expression across 16 languages. The development signals a growing maturity in multimodal AI, moving beyond mere transcription to nuanced content creation, impacting voiceover industries and accessibility tools.

The key differentiator for Qwen Audio 3.0 TTS Plus appears to be its sophisticated style control, allowing users to dictate tone and emotion through natural language prompts or specific tags, a feature that addresses a perceived limitation in prior TTS systems. However, its processing speed of 16 characters per second, while functional, lags considerably behind real-time human speech and potentially other emerging models. Future developments will likely focus on bridging this speed gap without sacrificing the nuanced control Alibaba has demonstrated.