🤖 AI Summary
This work proposes a hierarchical wireless foundation model to overcome the limited generalization of existing task-specific AI methods in wireless communication. The architecture jointly optimizes multiple tasks through an upstream geometry-aware channel encoder and a downstream differentiable optimization decoder. It introduces a novel geometry-aware cross-attention mechanism and self-supervised masked reconstruction to learn high-fidelity, task-agnostic channel representations. A hybrid supervised–unsupervised training strategy and modular design further enhance adaptability. The model achieves performance comparable to task-specific baselines while significantly reducing inference latency, demonstrates strong generalization to unseen channel conditions and system configurations, and incurs minimal parameter overhead.
📝 Abstract
The increasing complexity of next-generation wireless networks has driven the integration of artificial intelligence (AI) into wireless communications. However, most existing studies focus on developing task-specific deep learning techniques for single scenarios, which limits their ability to generalize across diverse tasks, channel conditions, and system configurations. To address this generalization bottleneck, we propose a hierarchical wireless foundation model (WFM) for multi-task optimization. The proposed WFM couples an upstream foundation channel encoder (FCE) with a downstream foundation optimization decoder (FOD) via geometry-aware cross-attention. Specifically, the FCE extracts task-agnostic channel representations via self-supervised masked reconstruction while the FOD generates multi-task optimization decisions through differentiable output heads. Moreover, a hybrid supervised-to-unsupervised training strategy is employed to overcome the performance ceiling of purely supervised learning, and the modular architecture of the WFM enables efficient adaptation to unseen communication tasks with minimal parameter overhead. Simulation results show that the proposed WFM learns high-fidelity channel representations and achieves competitive multi-task optimization performance while substantially reducing optimization inference latency relative to numerical baselines. Furthermore, it exhibits robust generalization to unseen propagation environments, varying constraint parameters, and heterogeneous system configurations.