Score
Designs, builds, and trains control policies and systems that produce continuous-valued actions, including simulation environments and benchmark suites for continuous control and the evaluation metrics that measure control performance. Implements and analyzes architectures and tooling for decentralized, distributed, hierarchical, stochastic, and real-time controllers, and establishes training and experiment pipelines (including continuous integration) for benchmarking and deploying continuous-action reinforcement learning agents.
This paper addresses the persistent gap between theoretical advances in reinforcement learning (RL) and their practical deployment in robotics and control systems. To bridge this divide, we propose a structured taxonomy tailored to real-world robotic applications, grounded in the Markov decision process (MDP) framework and systematically incorporating mainstream deep RL algorithms—including DDPG, TD3, PPO, and SAC—across canonical domains such as motion control, dexterous manipulation, and multi-agent coordination. The taxonomy explicitly integrates training paradigms and deployment maturity metrics. Crucially, we identify recurring design patterns and evolutionary trends in high-dimensional continuous control tasks, thereby unifying theoretical insights with engineering constraints. Our framework advances reproducibility, transferability, and robustness in RL deployment on physical robots, offering both a methodological foundation and actionable guidelines for practitioners. (149 words)
In continuous-time reinforcement learning, the discrete-time execution and performance evaluation of stochastic policies have long lacked rigorous theoretical foundations. This work introduces a piecewise-constant control framework and establishes, for the first time, the weak convergence of discretely sampled policies to their continuous-time stochastic counterparts in the fine-mesh limit. We derive the optimal first-order convergence rate and provide both high-probability and almost-sure convergence guarantees. Leveraging tools from stochastic analysis and weak convergence theory, we quantify bias and variance bounds for policy evaluation and policy gradient estimation under discrete-time observations. These results furnish a rigorous theoretical basis for exploratory stochastic control. The study bridges a critical gap in the convergence analysis of policy discretization in continuous-time RL, thereby enhancing the interpretability and reliability of algorithm design.
This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.
In autonomous control of multi-stage industrial processes, balancing local specialization with global coordination remains challenging. Method: We propose the first multi-agent reinforcement learning (MARL) benchmark environment tailored to sequential industrial recycling tasks (sorting + baling), bridging the gap between academic benchmarks and real-world industrial requirements. Built upon SortingEnv and ContainerGym, our scalable simulation platform enables systematic comparison of modular multi-agent versus monolithic agent architectures, augmented with an action masking mechanism to explicitly constrain the feasible action space according to industrial constraints. Results: Action masking substantially improves training stability and final performance for both architectures, significantly narrowing their performance gap. Under realistic action constraints, the advantage of specialization diminishes markedly. This work establishes a new paradigm for evaluating transferability in industrial RL and highlights the critical role of action-space modeling in control policy design.
To address coordinated control of multiple subprocesses in cyber-physical systems, this paper proposes a dual-timescale hierarchical decentralized control architecture: a global controller—formulated as an infinite-horizon discounted Markov decision process (MDP)—optimizes overall performance under budget constraints; meanwhile, (N) local controllers, each modeled as an MDP, autonomously make decisions under either constrained optimization (COpt) or unconstrained optimization (FOpt) frameworks. We establish, for the first time, a rigorous theoretical connection between COpt and FOpt: proving existence of optimal policies, deriving bounds on the difference between their optimal value functions, and characterizing equivalence conditions. We further prove that static deterministic optimal policies exist and identify precise budget–cost matching conditions under which COpt and FOpt become equivalent. These results provide a sound theoretical foundation and principled design guidelines for decentralized, scalable, federated autonomous control.
This work addresses the challenges of credit assignment and high gradient variance in reinforcement learning with hybrid discrete-continuous action spaces, where conventional policy gradient methods suffer significant performance degradation, particularly in high-dimensional continuous action settings. To overcome these limitations, the paper proposes Hybrid Policy Optimization (HPO), which integrates pathwise derivatives and score function gradients to construct an unbiased hybrid gradient estimator within differentiable simulators. Theoretical analysis reveals that the cross-term in the hybrid gradient becomes negligible near discrete optimal responses, justifying an approximately decoupled update strategy that effectively reduces variance. Empirical results demonstrate that HPO substantially outperforms Proximal Policy Optimization (PPO) on inventory control and switched linear quadratic regulator tasks, with performance gains increasing as the dimensionality of the continuous action space grows.
This work addresses the frequent failures of toolchains in robotic policy training and the absence of reliable evaluation and recovery mechanisms. It proposes AgenticRobotics—a backend-agnostic agentic control plane powered by large language models that dynamically orchestrates ephemeral worker nodes to establish a recoverable and verifiable “train–evaluate–refine” transactional loop. Key innovations include an evidence-gated promotion mechanism, commit-key-based crash recovery, and a signed skill repository with a verified tool registry. The system enables zero-loss, zero-duplication execution and valid decision-making at arbitrary points under unattended operation, reducing erroneous promotions from 0.005–0.021 to 0.001 and successfully detecting all six classes of artifact tampering.
This work addresses the limited generalization of reinforcement learning (RL) policies in real-world deployment, where environmental dynamics shift, action and observation spaces vary, and control objectives change. To tackle this challenge, the authors introduce the first large-scale, physically realistic continuous control benchmark for HVAC control, built upon EnergyPlus. Leveraging a parameterized building generator, the benchmark systematically produces diverse building configurations and defines standardized tasks to evaluate key generalization capabilities—including objective adaptation, dynamic shifts, action space variations, and cross-domain transfer. The platform supports heterogeneous observation and action spaces and integrates with the Gymnasium interface alongside a unified evaluation protocol, thereby providing a robust foundation for advancing both building energy efficiency and the robustness of RL algorithms.
研究通过使用编码代理生成和优化策略代码,解决闭环保策设计成本高的问题,利用存档的实现经验提高新任务的成功率。
本文提出REFINEPPO,通过迭代行动改进方法解决连续控制问题,利用共享精炼网络逐步修正动作提案,与PPO结合,在多个基准任务上展示了更快的收敛速度和更好的性能。
This study addresses the challenges of multi-unit collaborative modeling, information heterogeneity, and strict action constraints in complex engineering systems by proposing a large language model-based hierarchical collaboration framework. The core innovation lies in the design of a Continuity-Aware GRPO algorithm, which evaluates the evolutionary impact of decisions on subsequent control intervals to effectively resolve heterogeneous context reasoning and constraint enforcement difficulties. By integrating large language models with reinforcement learning, this approach achieves precise cross-level decision-making. Empirical validation in traffic flow control and virtual power plant tasks demonstrates that the proposed method comprehensively outperforms existing baselines, significantly enhancing collaborative control efficacy in complex systems.