🤖 AI Summary
This work addresses the challenge of modeling path-dependent cycle aging in heterogeneous battery fleets participating in grid frequency regulation, where frequent charge–discharge cycles render traditional modeling approaches inadequate. To tackle this, the authors propose a reinforcement learning–based real-time scheduling method that formulates the problem as a constrained Markov decision process with a bounded action space and introduces a dense surrogate reward to align short-term control actions with the long-term objective of minimizing battery aging. Efficient function approximation in high-dimensional state–action spaces is achieved by combining random nonlinear feature mapping via Extreme Learning Machines (ELM) with linear temporal difference learning. Experiments on both synthetic and real-world grid regulation signals demonstrate that the proposed approach significantly reduces cycle depth frequency and aging metrics, outperforming existing baseline strategies.
📝 Abstract
Battery energy storage systems are increasingly deployed as fast-responding resources for grid balancing services such as frequency regulation and for mitigating renewable generation uncertainty. However, repeated charging and discharging induces cycling degradation and reduces battery lifetime. This paper studies the real-time scheduling of a heterogeneous battery fleet that collectively tracks a stochastic balancing signal subject to per-battery ramp-rate and capacity constraints, while minimizing long-term cycling degradation. Cycling degradation is fundamentally path-dependent: it is determined by charge-discharge cycles formed by the state-of-charge (SoC) trajectory and is commonly quantified via rainflow cycle counting. This non-Markovian structure makes it difficult to express degradation as an additive per-time-step cost, complicating classical dynamic programming approaches. We address this challenge by formulating the fleet scheduling problem as a Markov decision process (MDP) with constrained action space and designing a dense proxy reward that provides informative feedback at each time step while remaining aligned with long-term cycle-depth reduction. To scale learning to large state-action spaces induced by fine-grained SoC discretization and asymmetric per-battery constraints, we develop a function-approximation reinforcement learning method using an Extreme Learning Machine (ELM) as a random nonlinear feature map combined with linear temporal-difference learning. We evaluate the proposed approach on a toy Markovian signal model and on a Markovian model trained from real-world regulation signal traces obtained from the University of Delaware, and demonstrate consistent reductions in cycle-depth occurrence and degradation metrics compared to baseline scheduling policies.