Score
Design, adapt, and deploy a single control or decision policy intended to run on multiple robots, including training or fine-tuning one shared model so it performs across agent variations. Package and validate the policy so each robot can run an independent copy at inference without relying on inter-robot communication or per-robot bespoke controllers.
This work addresses the generalizable policy learning problem in multi-agent settings where agents possess heterogeneous action spaces but share a common observation space. We propose the first diffusion-based planning framework trained on joint multi-agent trajectories, integrating inverse dynamics modeling with cross-agent trajectory alignment to enable forward transfer and zero-shot adaptation to unseen agents. Evaluated on the BabyAI benchmark, our method achieves up to a 42.20% improvement in task completion accuracy, significantly outperforming single-agent training and conventional imitation learning baselines. Key contributions include: (1) the first application of diffusion models to cross-agent policy generalization; (2) the construction of an action-agnostic implicit planning representation that decouples observation understanding from action execution; and (3) effective policy knowledge transfer across heterogeneous action spaces via joint trajectory training.
Existing multi-robot path planning approaches exhibit limited generalization when scaled to larger systems and incur prohibitive training costs. This work introduces diffusion models to multi-robot path planning for the first time, proposing a novel “train small, deploy large” paradigm that supports dynamic robot counts. By training a shared diffusion model on a small number of robots and incorporating a dedicated inter-agent attention mechanism with temporal convolutions, the method enables efficient deployment to significantly larger-scale systems. Experimental results demonstrate that the proposed approach substantially outperforms current reinforcement learning and heuristic baselines across diverse scenarios, achieving notable advances in both scalability and planning accuracy.
This work addresses the scalability challenge in robot policy evaluation and deployment, where fragmented models, datasets, and interfaces necessitate O(NM) independent integrations for N policies and M environments. To overcome this, the authors propose XPolicyLab—a unified, open framework that reduces integration complexity to O(N+M) by decoupling policy inference from environment execution via a client-server architecture. Central to this approach is a standardized schema for observations, actions, and trajectories, complemented by lightweight adapters that confine heterogeneity to the policy side. The framework enables consistent evaluation across both simulation and physical robots. In experiments, XPolicyLab integrated 42 diverse policies, reducing average per-policy integration time from five hours to thirty minutes, while maintaining compatibility with RoboTwin, RoboDojo, and real-world robotic platforms.
In multi-agent reinforcement learning (MARL), independent policy gradient methods suffer from suboptimal convergence due to joint sampling error—i.e., the deviation of the empirical joint action distribution from the true joint policy distribution caused by independent action sampling over finite trajectories. This error stems from uncoordinated action selection across agents. To address it, we propose MA-PROPS: a method that introduces an adaptive centralized behavioral policy to dynamically identify and compensate for under-sampled joint actions, coupled with joint action probability redistribution to enable robust on-policy sampling. Crucially, MA-PROPS does not rely on the centralized-training-with-decentralized-execution (CTDE) paradigm and supports fully decentralized deployment. Experiments on cooperative and non-conflicting multi-agent tasks demonstrate that MA-PROPS significantly reduces joint sampling error, improves convergence to optimal policies, and enhances training stability. The approach provides a theoretically grounded and empirically efficient pathway to robustify independent policy gradient methods.
Existing reinforcement learning (RL) approaches for drone racing exhibit poor generalization to unseen tracks, necessitating costly retraining. Method: We propose a novel bilevel RL framework featuring “environment-as-policy” adaptive environment shaping: an upper-level policy dynamically generates customized training environments in real time, while a lower-level policy learns agile track navigation; integrated with Proximal Policy Optimization (PPO), differentiable track parameterization, and online evaluation, the framework enables dynamic, curriculum-based difficulty adjustment. Contribution/Results: Our method achieves zero-shot generalization to diverse, previously unseen high-difficulty tracks—both in simulation and on physical hardware—using only a single learned policy. It improves generalization success rate by 42% over state-of-the-art environment shaping methods, effectively overcoming the generalization bottleneck imposed by static or hand-designed environments.
Existing adaptive agent-based regulatory simulations lack mechanisms to systematically integrate diagnostic feedback into policy controllers, resulting in delayed and opaque policy adjustments. This work proposes a lightweight machine-guided policy revision layer that represents policies as defeasible rules and combines symbolic control, defeasible logic, and policy prioritization to operationalize contestability at the controller level, thereby endowing policy decisions with explainability, contestability, and dynamic revisability. In an emissions regulation agent-based model, the approach significantly reduces problem recurrence under scenarios where the VPVA mechanism fails due to excessive conservatism, while effectively maintaining key performance indicators such as violation rates, overshoot, and volatility.
This work addresses the challenge of transferring motion policies across robots with diverse morphologies and dynamics. To enable cross-robot policy generalization, the authors propose representing actions as incremental changes in Cartesian-space states and introduce the State Prediction and Adaptive Command Execution (SPACE) framework. This framework leverages geometric end-effector displacement prediction combined with lightweight action adapters to bridge embodiment differences. Notably, it is the first approach to uniformly handle dynamic discrepancies across three levels: between distinct robot embodiments, among hardware variants of the same morphology, and within a single robot under runtime condition changes. Experiments demonstrate that behavior cloning trained on this universal action representation significantly outperforms direct control command prediction when deployed across platforms, maintaining robustness under variations in control frequency, payload, and controller gains.
Existing controller designs in multi-agent large language model systems typically support only one-shot routing and lack the capacity for critical evaluation and iterative refinement of intermediate outputs. This work reframes multi-agent collaboration as a sequential decision-making problem and introduces the first controller that integrates both critique and routing functionalities, iteratively assessing the current draft and dynamically selecting the optimal agent for refinement. The approach is grounded in a constrained finite-horizon Markov decision process, trained with a composite reward function and Lagrangian relaxation-based policy gradient methods. Evaluated across seven reasoning benchmarks, the proposed method significantly narrows the performance gap with the strongest single model while using fewer than 25% of its inference calls, achieving efficient and controllable collaborative reasoning.
This work addresses the challenge of enabling multi-robot systems to simultaneously adapt to unknown environments, unfamiliar teammates, and dynamically varying team sizes in open-world settings. To this end, it proposes the first hypergraph game–based open-adaptive teaming framework, which models team-level non-pairwise collaborative relationships through hypergraphs to support structural reasoning under dynamic team compositions. The framework incorporates a progressive diversity-aware training mechanism—termed HOLA—that enables zero-shot policy transfer across environments, partners, and team scales without fine-tuning. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines in cooperative pursuit tasks involving multiple drones and legged robots, with policies directly deployable on Crazyflie and Zsibot L1 hardware platforms to achieve robust collaboration in entirely unseen scenarios.