🤖 AI Summary
This work addresses the fundamental trade-off between energy efficiency and latency in AI-driven distributed baseband processing within O-RAN architectures. For the first time, AI inference overhead is explicitly incorporated into an end-to-end modeling framework that jointly accounts for computation, transmission, and inference energy consumption and delays across access, metro, and backbone network segments. A throughput-aware energy model, fine-grained latency analysis, and mixed-integer programming formulation are combined with real hardware emulation to derive a joint optimization model and propose a quality-of-service- and load-aware task placement strategy. Results demonstrate that optimal deployment configurations are jointly determined by user QoS requirements and network load, offering both theoretical grounding and practical guidance for realizing low-latency, energy-efficient, AI-native O-RAN systems.
📝 Abstract
The Open Radio Access Network (O-RAN) architecture introduces flexible functional splits and open interfaces that enable distributed and centralized deployment of baseband processing. While this flexibility offers opportunities for improved resource utilization, it also introduces fundamental trade-offs between energy efficiency and latency. In this paper, we develop a throughput-based end-to-end energy consumption model for O-RAN and extend it by incorporating detailed latency modeling and application-specific Artificial Intelligence/Machine Learning inference costs. The proposed end-to-end modeling framework provides a general representation of processing, transport, and inference-related energy and delay across the access, metro, and long-haul network segments. Building on this general model, we formulate an optimization problem that selects the placement of baseband processing and AI inference tasks across candidate O-RAN configurations to analyze energy-latency tradeoffs under network load, server frequency, and energy-budget constraints. Using representative hardware platforms and realistic traffic assumptions, we evaluate multiple baseband processing placements corresponding to different O-RAN functional configurations. Our results reveal how user quality of service requirements and network load conditions jointly determine the optimal placement of baseband processing and AI inference tasks, highlighting the inherent trade-off between energy efficiency and latency. The analysis provides practical insights for latency-aware and energy-efficient O-RAN deployments supporting emerging AI-driven services.