AI news story

Qwen Team Releases FlashQLA: a High-Performance Linear Attention Kernel Library That Achieves Up to 3× Speedup on NVIDIA Hopper GPUs

The QwenLM team has released FlashQLA, a new kernel library that dramatically accelerates the forward and backward passes of Gated Delta Network (GDN) Chunked Prefill, targeting both large-scale pretraining and edge-side agentic inference scenarios.

  • Hardware
  • Source: MarkTechPost
  • Published: 2026-04-29
  • Signal score: 4
  • 43 sources

Editor's take

The Qwen team has introduced FlashQLA, a kernel library designed to significantly boost the inference speed of certain large language model architectures, particularly Gated Delta Networks, by up to threefold on NVIDIA's Hopper (H100) GPUs.

This development is crucial as it addresses a key bottleneck in deploying large models, especially for real-time applications like agentic AI. By optimizing the computationally intensive linear attention mechanism, FlashQLA could enable more complex and responsive AI agents to operate at the edge, bringing advanced AI capabilities to devices with more constrained resources, a strategic goal for companies like NVIDIA and its partners.

Future developments will hinge on FlashQLA's broader compatibility beyond GDN architectures and its actual integration into popular LLM frameworks. Demonstrating similar speedups for more widely adopted models like Llama 3 or Mistral's latest offerings, and showing scalability across different hardware generations, will be key indicators of its long-term impact.

Signal score: 4

This event was corroborated by 43 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More Hardware stories

  1. Firebird Takes Its AI Factory Platform Global With a 2-Gigawatt Pipeline

    Unite.AI · 2026-08-08

    Firebird opened its first AI factory in Hrazdan, Armenia, on August 8, 2026, and used the ceremony to lay out the rest of the map: a second market in Kazakhstan with 125 megawatts

  2. NVIDIA AI Releases NOOA: An Object-Oriented Python Framework That Turns an AI Agent Into a Single Python Class

    MarkTechPost · 2026-08-07

    NVIDIA Labs has open-sourced NOOA (NVIDIA Object-Oriented Agents), a model-agnostic Python framework for building AI agents.

  3. Firebird Launches CIS Region’s Largest AI Factory in Armenia

    NVIDIA AI Blog · 2026-08-08

    The global buildout of AI infrastructure reached a new milestone today — Firebird, an emerging AI cloud, launched the CIS region’s largest AI factory in Armenia

  4. d-Matrix Buys Wallaroo to Orchestrate Inference Across Chips

    Unite.AI · 2026-08-03

    d-Matrix has acquired Wallaroo.ai, a maker of software for deploying and orchestrating AI inference, in a deal the Santa Clara chip company announced on August 3, 2026.

  5. ASML Supplier Zeiss Says It Can Handle Demand for Key AI Parts

    Bloomberg · 2026-08-03

    One of the critical suppliers in the semiconductor industry, Germany’s Zeiss Group, pushed back on investor concerns about bottlenecks in the AI supply chain and said it’s

  6. Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model

    MarkTechPost · 2026-08-02

    Inkling-Small matches Inkling at a quarter the size, and its NVFP4 checkpoint runs on one NVIDIA B300 GPU