AI news story

Grok 4.20 trails Gemini and GPT-5.4 by a wide margin but sets a new record for not hallucinating

xAI's Grok 4.20 is cheap, fast, and hallucinates less than any other tested model, but it can't keep up with the top tier in…

  • LLMs
  • Source: The Decoder
  • Published: 2026-03-12

Editor's take

xAI's Grok 4.20 demonstrates a significant reduction in factual inaccuracies, outperforming established models like Google's Gemini and OpenAI's GPT-5.4 in zero-hallucination metrics. This development is crucial as it addresses a core limitation of current large language models, potentially making them more reliable for sensitive applications where truthfulness is paramount, even if raw performance on traditional benchmarks lags.

The focus on hallucination reduction, rather than solely benchmark scores, signals a potential shift in LLM development priorities. Future advancements will likely hinge on whether this trade-off between accuracy and raw generative capability can be narrowed, and if companies like xAI can translate this reliability into commercially viable products that compete with the broader feature sets of their rivals.