AI news story

I put GPT-5.5 through a 10-round test: It scored 93/100, losing points only for exuberance

OpenAI's latest model delivers powerful results but sometimes ignores simple directions, creating a tension between intelligence and control.

  • LLMs
  • Source: ZDNet
  • Published: 2026-04-24
  • Signal score: 3
  • 28 sources

Editor's take

OpenAI's GPT-5.5 demonstrates strong performance, achieving a 93/100 score in a recent evaluation, though it occasionally deviates from explicit instructions. This indicates a continued challenge in aligning advanced language models with precise user directives, a hurdle present since early iterations like GPT-3 and even more recent ones like GPT-4. The model's tendency towards "exuberance" suggests a trade-off between generating rich, creative output and adhering strictly to constraints, impacting its reliability for tasks demanding absolute fidelity.

The implications extend to various applications, from automated content generation and coding assistance to customer service bots, where predictability is paramount. Companies relying on LLMs for critical functions will need to carefully consider the robustness of GPT-5.5's adherence to negative constraints. This tension between generative power and controllability is a key area of research, with implications for the practical deployment of AI in sensitive domains.

Future developments will likely focus on refining instruction following capabilities. It will be important to observe whether subsequent model updates, potentially GPT-6, show improved performance in this specific area, or if developers must rely on more sophisticated prompt engineering and fine-tuning techniques to mitigate these deviations. A significant improvement in zero-shot instruction following for complex, multi-step tasks would mark a notable advancement.

Signal score: 3

This event was corroborated by 28 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More LLMs stories

  1. OpenAI acquires presentation startup NextSlide

    TechCrunch · 2026-08-08

    NextSlide says its team members are now working on ChatGPT.

  2. Claude Vs ChatGPT: How These AI Assistants Differ

    Engadget · 2026-08-08

    In a practical breakdown of how Claude and ChatGPT AI models differ, one tends to fall short when it comes to quality responses and overall user experience.

  3. Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals

    The Decoder · 2026-08-08

    Starting August 14, Anthropic will make Auto Mode in Claude Code the default for Pro, Max, and Team plans. The company says it's safer.

  4. Responding to the next frontier of critical cyber capabilities

    OpenAI Blog · 2026-08-07

    OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.

  5. OpenAI says it slowed Astra model development over security concerns

    TechCrunch · 2026-08-07

    OpenAI said this model, which is still in development, reached its "critical cybersecurity threshold," meaning it could independently identify and carry out cyberattacks against

  6. Presentation: Keeping ChatGPT Fast as AI Development Accelerates

    InfoQ · 2026-08-08

    Martin Spier explains how agentic workflows dramatically increase code change volume at OpenAI. He d