Alibaba Cloud Cuts Nvidia GPU Usage 82% | Pooling System

by priyanka.patel tech editor

Alibaba Cloud slashes GPU Demand by 82% with New AI Computing Solution

A new computing pooling system developed by Alibaba Cloud dramatically reduces the need for expensive Nvidia GPUs to power artificial intelligence models, cutting demand by 82 percent. The breakthrough, dubbed Aegaeon, addresses a critical inefficiency in the rapidly expanding AI landscape and promises meaningful cost savings for cloud providers.

Alibaba Cloud, the AI and cloud services division of e-commerce giant Alibaba Group Holding, unveiled Aegaeon after extensive beta testing in its model marketplace. The system’s core innovation lies in efficiently allocating computing resources, a challenge that has become increasingly pressing as demand for AI services surges.

Did you know? – GPU demand has skyrocketed due to the computational intensity of AI, particularly large language models. This has led to shortages and high costs for cloud providers and developers.

Addressing the AI Resource Bottleneck

The problem Aegaeon tackles is the uneven distribution of demand across the vast library of available AI models. While a select few, such as Alibaba’s own Qwen and DeepSeek, receive the bulk of user requests, the majority remain largely unused. This leads to significant resource inefficiency, with valuable GPUs sitting idle or underutilized.

“Aegaeon is the first work to reveal the excessive costs associated with serving concurrent LLM workloads on the market,” researchers from Peking University and Alibaba Cloud wrote in a research paper presented this week at the 31st Symposium on operating Systems Principles (SOSP) in Seoul, South Korea.

Pro tip: – Efficient resource allocation is key to making AI more accessible. Technologies like aegaeon can lower costs,enabling wider adoption and innovation.

How Aegaeon Works: A Dramatic Reduction in GPU Usage

During the three-month beta period, Aegaeon demonstrably improved resource allocation. The system reduced the number of Nvidia H20 GPUs required to support dozens of models – some with up to 72 billion parameters – from 1,192 to just 213. This represents an 82 percent reduction in GPU demand, a figure that underscores the potential for substantial cost savings.

The researchers found that 17.7 percent of GPUs in Alibaba Cloud’s marketplace were allocated to serve only 1.35 percent of requests. GPU pooling, a technique where a single GPU serves multiple models, is not new, but Aegaeon represents a significant advancement in its implementation and effectiveness.

Implications for Cloud Providers and the future of AI

The progress of Aegaeon has broad implications for cloud services providers like Alibaba Cloud and ByteDance’s Volcano Engine, which concurrently serve thousands of AI models to users. By optimizing resource allocation, Aegaeon allows these providers to deliver AI servi

Why was aegaeon developed? Aegaeon was created to address the significant resource inefficiency in AI computing, specifically the uneven demand across a large library of AI models. Valuable GPUs were frequently enough sitting idle or underutilized, leading to excessive costs.

Who developed Aegaeon? The system was developed by researchers from Peking University and Alibaba Cloud, the AI and cloud services division of Alibaba Group Holding.

What is aegaeon? Aegaeon is a new computing pooling system that dramatically reduces the need for expensive Nvidia GPUs to power AI models. It achieves this by efficiently allocating computing resources, optimizing GPU usage.

How did it end? Aegaeon completed a three-month beta testing period within Alibaba Cloud’s model marketplace, demonstrating an 82% reduction in GPU demand. The research detailing its capabilities was presented at the 31st Symposium on Operating Systems principles (SOSP) in Seoul, South Korea, signaling a commitment to sustainable and scalable AI infrastructure. The research paper detailing Aegaeon’s capabilities. The company’s commitment to innovation in this area signals a growing focus on sustainable and scalable AI infrastructure.

Leave a Comment