AI news story

Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens

Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Spa…

  • LLMs
  • Source: The Decoder
  • Published: 2026-07-07

Editor's take

Anthropic has uncovered a latent internal reasoning mechanism, dubbed "J-Space," within its Claude large language models, which functions akin to a working memory. This development allows for a degree of introspection into the model's thought processes, revealing its capacity to identify and process artificial prompts and constraints.

The discovery is significant as it offers a novel path to understanding, and potentially controlling, the emergent behaviors of LLMs. Previously, the internal workings of models like Claude 2.1 remained largely opaque, making it difficult to diagnose or predict how they would handle complex or adversarial inputs. The ability to probe J-Space could inform future model architectures and safety protocols.

Future developments to monitor include whether J-Space is a general phenomenon across other LLMs, such as Meta's Llama 2 or Google's Gemini series, and if Anthropic can leverage this insight to improve Claude's factual accuracy or reduce harmful outputs. The extent to which this internal monologue can be reliably manipulated or aligned with human objectives will be a key indicator of its practical value.