AI news story
Perplexity Splits AI Work Between PCs and Servers to Ease Strain
Perplexity AI Inc. is trying to manage the huge demand for artificial intelligence computing power by building a platform that diverts AI work between personal computers and cloud-based servers.
Editor's take
Perplexity is developing a hybrid computing approach, offloading some AI processing to user devices while retaining complex tasks on cloud infrastructure. This strategy aims to alleviate the significant strain on specialized AI hardware, like NVIDIA's H100 GPUs, which are currently a bottleneck for many AI companies. By distributing workloads, Perplexity seeks to improve responsiveness and potentially reduce its reliance on costly, scarce server resources, a challenge faced by nearly every player in the generative AI space, from OpenAI to Google.
The success of this distributed model hinges on efficient task segmentation and seamless data transfer. If Perplexity can effectively balance computational demands between client devices and their servers, it could offer a more scalable and cost-effective path to delivering AI services, potentially influencing how other AI providers manage their infrastructure. Future developments will reveal how this architecture handles varying client hardware capabilities and network conditions, and whether it truly mitigates the GPU supply chain constraints impacting the broader industry.
Signal score: 5
This event was corroborated by 5 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by Bloomberg. Read the original article at Bloomberg.