Score
Implement RANSAC-compatible model solvers and the integration code used inside a RANSAC loop; this involves designing numerically stable, per-sample-efficient solver routines, detecting and handling degenerate sample configurations, and providing clean interfaces for hypothesis generation, verification, and model refinement so the solver can be called repeatedly by the RANSAC controller.
This study addresses the unclear practical efficacy of automatically generated polynomial symmetry-breaking constraints in integer linear programming across different solvers. The authors systematically evaluate the performance of mainstream mathematical programming and SMT solvers when handling such constraints, comparing three strategies: native quadratic handling, internal reformulation, and explicit linearization. Their experiments reveal that the effectiveness of symmetry breaking is highly solver-dependent, advocating for a solver-aware evaluation paradigm. The findings indicate that compact families of quadratic symmetry-breaking constraints generally enhance solver performance, whereas excessive linearization, overly large breaking sets, or inappropriate reformulations often lead to model bloating or search degradation, thereby diminishing or even reversing potential benefits.
To address the insufficient time-integration capabilities of the SUNDIALS numerical library in high-performance scientific computing, this project systematically extends its time-stepping solvers. Methodologically, it introduces three novel classes of single-step methods—low-storage Runge–Kutta (LSRK), symplectic structure-preserving block RK, and general-purpose operator-splitting schemes—alongside a new multi-rate adaptive step-size controller and explicit RK-based adjoint sensitivity analysis (filling a longstanding gap). It further enhances nonlinear solvers with Anderson acceleration and improves error handling and logging infrastructure. These contributions significantly improve efficiency, stability, and accuracy for large-scale transient simulations, enabling long-duration, high-fidelity, and multiphysics-coupled modeling. Validation across multiple HPC applications demonstrates speedups of 1.5–3× and markedly improved numerical robustness.
This work addresses a key limitation of conventional LLM-based PDE solvers, which implicitly embed numerical strategies within generated code, making pre-execution validation and post-failure correction challenging. To overcome this, the authors propose AutoPDE, the first framework to explicitly model solution strategies as revisable, decoupled objects separate from implementation code. AutoPDE employs a three-stage pipeline—PDE type identification, numerical method selection, and adaptive parameter tuning—augmented by low-overhead trial solves and a reusable skill library to construct and refine strategies prior to code generation. Evaluated on the PDE Agent Bench, AutoPDE achieves a 54.5% pass rate, outperforming the strongest baseline by 14.2 percentage points, thereby substantially improving both the reliability and interpretability of AI-driven PDE solving.
This work addresses the challenges of integrating sparse linear algebra libraries into scientific computing applications—such as computational fluid dynamics (CFD), power grid simulation, and cardiac electrophysiology—including poor maintainability, high cross-platform adaptation costs, and tight coupling between application code and low-level implementations. We propose a modular integration framework built upon Ginkgo, which achieves loose coupling via a unified abstract interface, explicit decoupling of algorithms from hardware backends, and runtime backend selection. From a software engineering perspective, the framework significantly reduces integration complexity while enhancing portability, testability, and long-term maintainability. Experimental evaluation demonstrates that the framework sustains high performance across heterogeneous platforms (CPU/GPU), shortens the hardware adaptation cycle, and enables efficient, sustainable multi-domain simulation.
This work addresses the frequent mismatch between user-specified physical intent and the actual behavior of multiphysics simulation code generated by large language models, often due to erroneous implementations of partial differential equations (PDEs). To bridge this gap, we propose a PDE-structure-based intent verification method that deterministically reconstructs the governing equations implicitly encoded in the generated code and compares them against the user’s intended PDEs, enabling semantic correctness validation and iterative refinement. We introduce, for the first time, a formal metric termed the Intent Fidelity Score (IFS) to quantify alignment with physical intent, establish a PDE-driven feedback loop, and demonstrate compatibility with major PDE frameworks including MOOSE, FEniCS, and FreeFEM. Evaluated on 220 cases in MooseBench, our approach substantially improves IFS—by 0.22–0.41 on challenging instances with initial IFS < 0.7—while audits reveal that execution-only repair strategies still yield physically incorrect results in 39–40% of cases.
This work addresses the limitations of traditional model-based reduced-order modeling in scenarios where high-fidelity model code is difficult to integrate, necessitating data-driven alternatives. It presents the first systematic integration of data-driven techniques—such as Dynamic Mode Decomposition—with model-driven approaches within the open-source library pyMOR. Built upon a unified interface of VectorArray, Operator, and Model abstractions, the proposed framework enables a flexible and efficient hierarchical reduction pipeline. The study demonstrates seamless interoperability between data-driven and model-driven methods, validates the efficacy of data-driven reduction through practical case studies, and highlights its performance advantages and complementary potential relative to conventional approaches.
This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.
This work addresses the challenge of selecting the regularization parameter (nugget) in ill-posed linear systems arising in machine learning, where existing adaptive methods lack compatibility with automatic differentiation and suffer from computational inefficiency. To overcome these limitations, we introduce autonugget, a lightweight Python package fully compatible with JAX’s automatic differentiation framework. Our approach uniquely integrates Richardson extrapolation with Tikhonov regularized solutions computed across multiple nugget values, thereby preserving end-to-end differentiability while avoiding the information loss inherent in single-solution strategies. Experimental results demonstrate that autonugget significantly enhances solution accuracy and training stability without compromising rapid prototyping capabilities.
This work addresses the absence of a standardized benchmark for evaluating code generation targeting partial differential equation (PDE) solvers, particularly with respect to numerical accuracy, computational efficiency, and compatibility with mainstream finite element libraries. To bridge this gap, the authors introduce the first multi-metric, multi-library benchmark for PDE solver generation, comprising 645 structured instances spanning six mathematical problem types and eleven PDE classes. The benchmark supports three major finite element frameworks—DOLFINx, Firedrake, and deal.II—and incorporates a staged evaluation framework that holistically assesses code executability, numerical correctness, and performance. Experimental results demonstrate that while current large language models can produce executable code, their success rate drops substantially when stringent accuracy and efficiency requirements are imposed, thereby underscoring the necessity and effectiveness of the proposed benchmark in advancing reliable and efficient automated PDE solver generation.
This work proposes the first native Rust implementation of a Modelica compiler that seamlessly bridges Modelica with modern scientific computing ecosystems such as CasADi, JAX, and Julia, eliminating the need for error-prone model rewrites and preserving full semantic fidelity. By structuring compilation into well-defined phases that translate Modelica into a generic algebraic intermediate representation, the system enables multi-backend code generation and real-time simulation. It uniquely supports zero-install execution in web browsers via WebAssembly and facilitates hardware- and software-in-the-loop control. The compiler covers core functionalities of the Modelica Standard Library and demonstrates superior compilation and simulation performance compared to existing open-source alternatives, with successful validation in quadrotor real-time control and cross-platform unified modeling scenarios.