🤖 AI Summary
Existing research on optimizing data freshness in wireless networks faces two key limitations: (i) Age-of-Information (AoI) theory primarily relies on classical models, lacking adaptability to B5G/6G’s heterogeneous, dynamic environments; and (ii) reinforcement learning (RL) surveys lack a systematic, freshness-centric taxonomy. Method: This paper introduces the first freshness-aware RL survey framework tailored for B5G/6G, proposing a unified four-category classification of freshness-oriented policies—update control, channel access, risk-sensitive decision-making, and multi-agent coordination—and jointly modeling AoI alongside its functional and application-specific variants. It integrates AoI theory, cross-layer optimization, and risk-sensitive Markov decision processes to formalize critical challenges: delayed decision-making, stochastic environment modeling, and decentralized multi-agent collaboration. Contribution/Results: The work establishes the first systematic learning-theoretic foundation for intelligent freshness management in 6G networks, enabling principled design of adaptive, robust, and scalable freshness-aware protocols.
📝 Abstract
The age of information (AoI) has become a central measure of data freshness in modern wireless systems, yet existing surveys either focus on classical AoI formulations or provide broad discussions of reinforcement learning (RL) in wireless networks without addressing freshness as a unified learning problem. Motivated by this gap, this survey examines RL specifically through the lens of AoI and generalized freshness optimization. We organize AoI and its variants into native, function-based, and application-oriented families, providing a clearer view of how freshness should be modeled in B5G and 6G systems. Building on this foundation, we introduce a policy-centric taxonomy that reflects the decisions most relevant to freshness, consisting of update-control RL, medium-access RL, risk-sensitive RL, and multi-agent RL. This structure provides a coherent framework for understanding how learning can support sampling, scheduling, trajectory planning, medium access, and distributed coordination. We further synthesize recent progress in RL-driven freshness control and highlight open challenges related to delayed decision processes, stochastic variability, and cross-layer design. The goal is to establish a unified foundation for learning-based freshness optimization in next-generation wireless networks.