Score
Designs and implements differentiable planning and optimization methods that select whether to perform aerial suppression drops and tune continuous drop parameters (locations, orientations) so those decisions are represented as variables amenable to gradient computation. Builds mappings from continuous action representations to grid- or simulation-based environment models and uses gradient-based optimization to minimize expected-impact objectives (for example expected burned area) under the chosen drop strategy.
This study addresses the challenge of wildfire spread prediction and aerial firefighting deployment under environmental and operational uncertainties. The authors propose a novel integrated approach that combines a hybrid CNN–cellular automata fire propagation model with differentiable gradient-based intervention planning. By jointly leveraging terrain, fuel, and wind field data, the method generates binary aerial drop patterns specifying both location and direction, while explicitly distinguishing between the immediate and sustained suppression effects of water and fire retardants. This work presents the first end-to-end coupling of deep learning–driven fire prediction with differentiable intervention strategies, enabling joint quantification of aleatoric and epistemic uncertainties. Evaluated on the 2020 Bear Fire case, the proposed framework significantly reduces burned area and supports robust, uncertainty-aware firefighting decisions.
Trajectory planning for SWAP-constrained UAVs in 3D environments remains challenging: modular approaches suffer from local minima, while end-to-end methods exhibit strong data dependency, large Sim2Real gaps, and dynamic infeasibility. Method: We propose the first self-supervised framework jointly integrating deep perception and differentiable trajectory optimization. It leverages an unlabeled 3D cost map to guide planning and introduces a neural-network-driven time-allocation strategy, enabling fully end-to-end differentiable perception–planning training. Contribution/Results: The method balances generalizability with physical interpretability. Experiments demonstrate significant improvements over SOTA: 31.33% reduction in position tracking error and 49.37% decrease in control energy consumption. Robustness is validated across both simulation and real-world UAV platforms, confirming practical transferability and reliability.
This work addresses the fragmentation and lack of a unified differentiable framework in existing reinforcement learning approaches for quadrotors operating across multiple tasks. To overcome these limitations, the authors propose a unified differentiable simulation framework capable of supporting four distinct tasks: hovering, trajectory tracking, landing, and racing. They further introduce an amended backpropagation through time (ABPT) algorithm that incorporates differentiable unrolling optimization, value-augmented objectives, and visited-state initialization to mitigate gradient bias arising from insufficient state coverage and non-differentiable rewards. Experimental results demonstrate that ABPT significantly improves performance on tasks with partially non-differentiable rewards while maintaining competitive results in fully differentiable settings, and successfully enables preliminary policy transfer to the real world.
To address the challenge of simultaneously satisfying hard constraints and ensuring real-time performance in quadrotor trajectory optimization, this paper proposes a spatial-temporal decoupled iterative optimization framework. Methodologically, trajectories are represented using B-splines, and a novel control-point-level strict constraint enforcement mechanism is introduced; a guidance-gradient-driven alternating QP-LP solving strategy, combined with constraint linearization, ensures efficient convergence. The key contribution lies in breaking the inherent trade-off between safety and computational efficiency: the method achieves millisecond-scale generation of safe, high-speed trajectories in both simulation and real-world flight experiments—significantly outperforming state-of-the-art approaches (e.g., CHOMP, TrajOpt). The implementation is open-sourced to facilitate reproducibility.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This work addresses the challenge of ensuring safety in industrial cyber-physical systems when applying deep reinforcement learning, where black-box exploration may inadvertently violate hardware constraints and conventional reward shaping struggles to balance safety with task performance. To overcome this, the authors propose a physics-informed safety mechanism that embeds a differentiable dynamics model into the loss function of a Proximal Policy Optimization (PPO) policy network. By performing short-horizon forward simulations to predict trajectories, the method imposes soft penalties—decoupled from the task-specific reward—on potential safety violations. This approach regularizes the policy online without requiring intricate reward engineering. Evaluated on a one-degree-of-freedom helicopter simulation platform, the method significantly reduces pitch angle constraint violations while maintaining excellent trajectory tracking performance.
This work addresses the challenge of local suboptimality in differentiable planning for highly nonlinear and hybrid discrete-continuous systems, where pathological objective landscapes—characterized by flat regions and abrupt gradients—hinder effective optimization. To overcome this, the paper introduces the Model-Driven Policy Optimization (MDPO) framework, which injects noise into the action space of a differentiable simulator to enable stochastic exploration. Crucially, MDPO dynamically modulates the noise intensity at each time step based on gradient sensitivity, thereby achieving adaptive, time-varying allocation of exploration resources. By explicitly incorporating gradient information into the exploration mechanism, the method substantially outperforms deterministic differentiable planning and model-free baselines such as PPO, yielding significantly higher-quality solutions across multiple complex benchmark tasks.
This work addresses the low sample efficiency of traditional model-free reinforcement learning methods—such as Proximal Policy Optimization (PPO)—which rely on high-variance advantage estimates. The authors propose Analytic Policy Gradients (APG), a method that leverages differentiable environment dynamics to compute exact, end-to-end gradients of policy returns with respect to policy parameters. To mitigate gradient degradation in long-horizon tasks, APG incorporates a segment-wise backpropagation mechanism and combines Monte Carlo estimation with critic-guided bootstrapping for effective gradient guidance. Evaluated on four continuous control benchmarks under identical network architectures and training protocols, APG consistently outperforms PPO, demonstrating substantially higher sample efficiency and faster convergence.
This work addresses the challenge of simultaneously achieving high trajectory-tracking accuracy, stability, and tunable response speed in quadrotor reinforcement learning control. The authors propose a heuristic, adjustable control method based on the Proximal Policy Optimization (PPO) algorithm, incorporating a reward function with dual-bandwidth exponential terms and an episode truncation mechanism. Within six million environment steps, the method efficiently trains a policy exhibiting critically damped responses and approximately 2% steady-state error. By adjusting reward weights and exponential coefficients, the controller can flexibly switch between acrobatic-like and inspection-like operational modes while maintaining stable performance. Experimental results across 100 random initial conditions demonstrate that the proposed approach achieves precise, tunable, and sample-efficient control in both position and yaw tracking.