AI news story

Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

AMD has released Lemonade, an open-source server designed for running large language models locally on their hardware, lever…

  • LLMs
  • Source: Hacker News
  • Published: 2026-04-02

Editor's take

AMD has released Lemonade, an open-source server designed for running large language models locally on their hardware, leveraging both GPUs and NPUs. This development is significant as it offers a performant, accessible alternative for on-device AI inference, potentially lowering barriers for developers and enterprises seeking private, efficient LLM deployment without relying solely on cloud infrastructure. The integration of NPUs hints at AMD's strategy to differentiate its silicon for AI workloads beyond traditional GPU dominance.

Future developments to monitor include Lemonade's adoption by key hardware partners and its ability to support increasingly complex models like Llama 3 70B or Mistral Large with competitive latency and throughput. Performance benchmarks against established solutions like NVIDIA's Triton Inference Server and the evolving landscape of hardware-accelerated inference will be crucial indicators of its long-term impact on the local LLM ecosystem.