🤖 AI Summary
This study addresses the challenges of repetitive development and limited cross-scenario transferability in wireless resource management (RRM) models by proposing the RRM-GPT framework. This work pioneers a foundation model paradigm for RRM that maps network observations into tokens and leverages a Transformer architecture to generate decision fields via an autoregressive mechanism. By integrating pretraining with reinforcement learning-based fine-tuning, the framework effectively captures inter-functional dependencies. Evaluated in 5G New Radio (NR) scenarios, a single unified model is capable of generating complete scheduling grants while achieving zero-shot transferability across diverse deployments. Consequently, this approach effectively overcomes the generalization bottlenecks inherent in conventional RRM models, offering a scalable and adaptable solution for next-generation wireless networks.
📝 Abstract
Learning-based models for radio resource management (RRM) are typically built for a single function and deployment, so each new setting repeats the development pipeline. RRM decisions, however, share a common structure: each is assembled from interdependent fields, defined by the standard, whose values are selected in view of the network state. We propose RRM-GPT, an autoregressive framework for RRM foundation models that generate these decisions as a language model generates text. An encoder maps heterogeneous network observations into a common token representation, and a decoder emits the decision one field at a time, each conditioned on the network state and the fields already committed. Pretraining on unannotated network logs teaches the model what makes a decision valid and how controllers choose among valid decisions; post-training then adapts it to deployment-specific operator objectives through imitation or reinforcement learning. The framework targets two forms of reuse: a function-specific model reused across deployments, and a model shared across RRM functions that generates their interdependent decisions as one sequence. In a 5G New Radio (NR) case study, we demonstrate that a single model generates complete scheduling grants spanning user selection, timing, link adaptation, resource allocation, and control signaling. The model captures dependencies among grant fields and transfers learned behavior to an unseen scenario without adaptation.