AI news story

NVIDIA Introduces a 4-Bit Pretraining Methodology Using NVFP4, Validated on a 12B Hybrid Mamba-Transformer at 10T Token Horizon

NVIDIA introduces a 4-bit pretraining methodology built around the NVFP4 microscaling format — combining selective BF16 layers, 16×16 Random Hadamard Transforms on Wgrad inputs, 2D weight scaling, and stochastic rounding on gradients — validated on a

  • Hardware
  • Source: MarkTechPost
  • Published: 2026-05-18
  • Signal score: 4
  • 34 sources

Editor's take

NVIDIA has unveiled a novel 4-bit pretraining technique, NVFP4, which leverages a combination of mixed-precision training, Hadamard transforms, weight scaling, and stochastic rounding. This innovation is particularly significant as it demonstrates the feasibility of training large models, like a 12 billion parameter hybrid Mamba-Transformer, at a 10 trillion token scale using substantially reduced precision.

This development addresses the escalating computational and memory demands of modern AI training. By enabling 4-bit pretraining, NVIDIA offers a path to more efficient and accessible large-scale model development, potentially lowering the barrier to entry for organizations and researchers struggling with the immense resource requirements of current state-of-the-art models. This could accelerate progress in areas where massive datasets are crucial.

Future observations should focus on the actual inference performance of models pretrained with NVFP4. Specifically, assessing any degradation in accuracy compared to their full-precision counterparts, especially on downstream tasks, will be critical. Additionally, understanding the hardware support and widespread adoption of NVFP4 by other model developers will indicate its long-term impact on the AI ecosystem.

Signal score: 4

This event was corroborated by 34 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More Hardware stories

  1. Firebird Takes Its AI Factory Platform Global With a 2-Gigawatt Pipeline

    Unite.AI · 2026-08-08

    Firebird opened its first AI factory in Hrazdan, Armenia, on August 8, 2026, and used the ceremony to lay out the rest of the map: a second market in Kazakhstan with 125 megawatts

  2. NVIDIA AI Releases NOOA: An Object-Oriented Python Framework That Turns an AI Agent Into a Single Python Class

    MarkTechPost · 2026-08-07

    NVIDIA Labs has open-sourced NOOA (NVIDIA Object-Oriented Agents), a model-agnostic Python framework for building AI agents.

  3. Firebird Launches CIS Region’s Largest AI Factory in Armenia

    NVIDIA AI Blog · 2026-08-08

    The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia

  4. d-Matrix Buys Wallaroo to Orchestrate Inference Across Chips

    Unite.AI · 2026-08-03

    d-Matrix has acquired Wallaroo.ai, a maker of software for deploying and orchestrating AI inference, in a deal the Santa Clara chip company announced on August 3, 2026.

  5. ASML Supplier Zeiss Says It Can Handle Demand for Key AI Parts

    Bloomberg · 2026-08-03

    One of the critical suppliers in the semiconductor industry, Germany’s Zeiss Group, pushed back on investor concerns about bottlenecks in the AI supply chain and said it’s

  6. Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model

    MarkTechPost · 2026-08-02

    Inkling-Small matches Inkling at a quarter the size, and its NVFP4 checkpoint runs on one NVIDIA B300 GPU