Score
Designs and implements automatic-differentiation components and integrations—e.g., reverse-mode/autograd implementations, custom gradient rules, and framework adapters—so numerical routines and training loops are end-to-end differentiable. Implements and tests gradient-rule registration, interface code, and numerical-stability fixes to ensure correct, stable backpropagation.
In neural ODE training, automatic differentiation (AD) with high-order ODE solvers—such as Leapfrog or linear multistep methods (LMMs)—often induces gradient distortion, leading to non-convergent oscillations that degrade training stability and generalization. This work identifies, for the first time, that standard AD breaks consistency between discrete gradients and the underlying continuous flow updates when applied to reversible or symmetric high-order discretizations—constituting the primary cause of oscillatory behavior. To address this, we propose a lightweight, solver-agnostic post-processing gradient correction method that enforces geometric consistency of discrete gradients with the flow map, thereby restoring gradient convergence without re-implementing the solver. Theoretical analysis establishes convergence guarantees under mild assumptions, and extensive numerical experiments demonstrate substantial suppression of training oscillations. Our approach consistently improves training stability and generalization accuracy across multiple benchmark tasks.
This paper addresses the lack of semantic foundations for automatic differentiation (AD) of algebraic data types (e.g., lists, trees) and arbitrary-order derivatives in higher-order functional languages. Methodologically, it introduces the first higher-order differentiable programming framework grounded in diffeological spaces, integrating categorical semantics and logical relations to formally define the semantics of forward-mode AD for higher-order functions and rigorously prove its structural preservation and Taylor-approximation completeness. Key contributions are: (1) the first complete correctness proof of AD semantics for languages with algebraic data types; (2) the first generalization of AD to arbitrary-order derivatives, unifying the behavior of derivative selection across primitive operations; and (3) the establishment of a unique macro-characterization of higher-order AD, providing a rigorous mathematical foundation for differentiable programming.
This work addresses the efficient and robust computation of gradients for numerical solutions of differential equations. We systematically survey four differentiable programming paradigms—adjoint methods, automatic differentiation (via source-to-source transformation and operator overloading), numerical perturbation, and symbolic-numeric hybrid approaches—and introduce, for the first time, a unified differentiability framework that bridges inverse problem solving and machine learning methodologies. We establish a cross-method comparative taxonomy and provide platform-specific best-practice guidelines for scientific computing libraries including SciPy, JAX, and TorchDiffeq. Our analysis rigorously characterizes trade-offs among accuracy, memory footprint, computational complexity, and applicability domains for each method. The results deliver both theoretical foundations and practical implementation pathways for differential-equation–data fusion modeling tasks, including parameter inversion, sensitivity analysis, and physics-informed neural networks (PINNs).
Neural ODE training suffers from high computational cost, excessive memory consumption, and numerical instability during backpropagation. To address these challenges, this paper introduces the Algebraically Invertible ODE Solver family, grounded in algebraically invertible numerical integration. The method integrates high-order implicit/explicit reversible schemes, adjoint-state techniques, and memory–computation co-optimization to achieve, for the first time, high-order accuracy, strict numerical stability, and exact gradient computation in backpropagation. Unlike recursive checkpointing, our approach achieves strictly superior time and memory complexity bounds. Extensive evaluation on multiple benchmark ODE tasks demonstrates a 2.1× reduction in training latency and a 68% decrease in GPU memory usage, while preserving gradient precision and numerical robustness.
This paper addresses the challenge of end-to-end differentiability in complex programs featuring nontrivial control flow and data structures. To this end, it introduces a probabilistic programming paradigm for differentiation, unifying optimization and probabilistic inference within a differentiable programming framework. Methodologically, it transcends conventional automatic differentiation (AD) by establishing, for the first time, a theoretical link between differentiability of control flow/data structures and uncertainty modeling—integrating AD, graphical models, convex optimization, and Bayesian inference into a cohesive differentiable program modeling framework. Key contributions include: (1) revealing that differentiable programming is fundamentally probabilistic programming—not merely gradient computation; (2) proposing the “program-as-model” design principle; and (3) establishing the first comprehensive knowledge system spanning theory, design, and applications, enabling the development of differentiable software infrastructure for large language models and foundation models.
This work addresses the challenge of implementing reverse-mode automatic differentiation for programs featuring algebraic effects such as finite discrete probabilistic choice. Building upon the Compositional Homomorphic Automatic Differentiation (CHAD) framework, it formulates automatic differentiation as a semantics-preserving program transformation and constructs a backward-pass mechanism tailored to the finite atomic distribution monad. The correctness of this construction is established using logical relations from category theory. This study presents the first systematic extension of reverse-mode automatic differentiation to effectful languages with discrete outputs, introducing a reusable differentiation scheme applicable to a broad class of algebraic effects—including nondeterminism, exceptions, and writer effects. The approach not only enables correct reverse differentiation of programs with finite discrete probabilistic structure but also lays a foundational theoretical groundwork for differentiating more general effectful languages.
This study addresses the challenges of extending differential programming, integration complexity, and insufficient evaluation within heterogeneous scientific software ecosystems by proposing an AI agent-driven framework for differentiable scientific software evolution. The framework introduces unified differentiation interfaces and shared resource mechanisms, leveraging AI coding agents to automate the implementation of automatic differentiation. Furthermore, it establishes a closed-loop quality assessment system integrating independent derivative verification, workflow testing, and performance benchmarking to drive recursive software improvement. Experimental evaluations across twenty scientific software packages demonstrate that the proposed approach significantly reduces gradient computation overhead while successfully enabling the efficient reuse of differentiable workflows in applications such as quantum control and thermal design.
This work addresses the challenge of selecting the regularization parameter (nugget) in ill-posed linear systems arising in machine learning, where existing adaptive methods lack compatibility with automatic differentiation and suffer from computational inefficiency. To overcome these limitations, we introduce autonugget, a lightweight Python package fully compatible with JAX’s automatic differentiation framework. Our approach uniquely integrates Richardson extrapolation with Tikhonov regularized solutions computed across multiple nugget values, thereby preserving end-to-end differentiability while avoiding the information loss inherent in single-solution strategies. Experimental results demonstrate that autonugget significantly enhances solution accuracy and training stability without compromising rapid prototyping capabilities.
This work addresses the high memory overhead and neglect of local structure in gradient computation for implicit nonlinear solvers within differentiable simulation. The authors propose a solver-level differentiation method that constructs an adjoint algorithm symmetric to the forward solve by reverse-scanning a block-structured implicit solver, entirely avoiding the assembly of a global Jacobian matrix. For the first time, adjoint computation is aligned with the block structure of the forward solver, combining vertex-block descent with reverse-colored Gauss–Seidel sweeps to enable efficient backpropagation using only local 3×3 adjoint solves. This approach leverages operator-view approximations of the inverse and its transpose. On a single GPU, it achieves a 33× speedup and 71× reduction in memory compared to unrolled automatic differentiation, enabling, for the first time, differentiable elastic dynamics simulation of million-contact coupled soft bodies with up to 8 million vertices.
This work addresses the high computational cost and memory consumption of automatic differentiation (AD) in physics-informed neural networks (PINNs), as well as its susceptibility to silent errors in architectures involving inter-sample dependencies such as BatchNorm or self-attention. The study presents the first systematic evaluation of finite differences (FD) as an alternative for derivative computation in PINNs, introducing a calibrated step-size strategy and a stochastic FD variant. It further proposes a sample-wise gradient approximation method that requires only forward passes. Experiments on three benchmark partial differential equations demonstrate that, in full-batch settings, FD achieves comparable accuracy to AD while being faster and using less memory. The proposed stochastic FD excels particularly in steady-state problems, and crucially, FD yields derivative errors an order of magnitude lower than AD in models with inter-sample dependencies.