Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANs

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficient large-model inference in resource-constrained distributed edge environments by proposing a multi-cluster collaborative inference framework. In this framework, devices within each cluster employ lightweight local models to extract features, which are then uploaded to multi-GPU edge servers for fusion. The approach jointly optimizes model pruning ratios, task scheduling, bandwidth allocation, and transmit power to minimize inference distortion under constraints on latency, energy consumption, and server capacity. By innovatively integrating rate-distortion theory with partial information decomposition, the study reveals the fundamental trade-off between model pruning and collaborative inference, and establishes a unified optimization mechanism tailored for multi-cluster edge intelligence networks. Experimental results demonstrate that the proposed framework significantly outperforms existing methods, achieving simultaneous improvements in inference accuracy and resource efficiency.
📝 Abstract
The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments. In this paper, we propose a multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively. Within each cluster, devices capture data from diverse perspectives and employ lightweight on-device LAIMs to extract local features. These features are then transmitted to the edge server, where they are aggregated and fused to generate a more accurate inference outcome. To reveal the fundamental trade-off between model pruning and collaborative inference performance, we develop a theoretical framework that characterizes the impact of pruning ratios and device contributions using rate-distortion theory and partial information decomposition. Based on this analysis, we formulate a joint optimization problem that determines the model pruning ratio, the task scheduling strategy, the bandwidth allocation, and the transmission power, with the goal of minimizing the inference distortion while satisfying the constraints of latency, energy consumption, and server capacity. Extensive simulation results demonstrate that the proposed framework significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks.
Problem

Research questions and friction points this paper is trying to address.

large AI models
model pruning
collaborative inference
edge intelligence
resource-constrained environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

pruning-aware co-inference
multi-cluster edge intelligence
rate-distortion theory
partial information decomposition
joint resource optimization
🔎 Similar Papers
No similar papers found.