🤖 AI Summary
This study addresses the fragmented search behavior and difficulty in navigating complex non-convex landscapes inherent to static-architecture hybrid algorithms for high-dimensional structural model updating. To this end, we propose DRL-DCO, which employs a Deep Deterministic Policy Gradient (DDPG) agent as a decision-maker. By introducing progress-aware states and diversity-based rewards, the method dynamically coordinates the global exploration of Differential Evolution (DE) with the local exploitation of CMA-ES. It unifies population evolution and restart mechanisms, enabling autonomous strategy switching and real-time resource reallocation, thereby integrating discrete algorithms into a cohesive framework supporting unsupervised reasoning. Evaluations on benchmark functions and IASC-ASCE structural health monitoring tasks demonstrate that DRL-DCO significantly outperforms mainstream adaptive and hybrid evolutionary algorithms in both convergence accuracy and robustness.
📝 Abstract
Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.