AI news story
China’s LongCat-2.0 is a 1.6T Model Trained Without NVIDIA
Meituan trained a 1.6T-parameter sparse MoE on 50,000 domestic AI chips with zero Nvidia GPUs, the architecture trick tha…
Editor's take
Meituan has successfully trained a 1.6 trillion parameter sparse Mixture-of-Experts (MoE) model, dubbed LongCat-2.0, utilizing 50,000 domestic AI chips and entirely bypassing NVIDIA hardware. This achievement underscores China's growing capability to develop and deploy large-scale AI models independent of Western chip manufacturers, a critical development given ongoing geopolitical tensions and export controls on advanced semiconductors.
The significance lies in demonstrating that sophisticated, massive AI models can be built using alternative hardware, potentially accelerating China's independent AI ecosystem development and challenging NVIDIA's dominance in the high-performance computing space for AI training. This could have material implications for global AI hardware supply chains and research collaborations.
Future attention should focus on the performance benchmarks of LongCat-2.0 compared to similarly sized models trained on NVIDIA infrastructure, like Meta's Llama 3 or Mistral AI's Mixtral 8x22B. Additionally, the long-term scalability and cost-effectiveness of this domestic chip ecosystem for ongoing AI research and deployment will be crucial indicators of its sustained impact.