Score
Designs and implements closed‑loop optimization systems that integrate computational optimizers (metaheuristics, parameter tuners) with physical hardware—commonly optical devices such as SLMs and SPIMs—to run measurement‑feedback experiments and compile and execute control on the device. Builds experimental pipelines for hardware‑in‑the‑loop tuning, uses hardware feedback to guide updates, and evaluates empirical performance under iteration, memory, power, and other hardware constraints.
This work addresses the challenge of deploying AI models on heterogeneous microcontrollers under stringent physical constraints—including memory, power consumption, and thermal limits—while preserving model accuracy, a task where manual tuning proves inefficient. The authors propose the first hardware-in-the-loop co-optimization framework driven by a large language model (LLM) agent, which automatically jointly optimizes both the neural model and firmware through a closed-loop cycle of compilation, flashing, and real-device measurement. Requiring only three iterations for deployment and surpassing human expert performance within seven, the method enables battery-free operation on solar-powered MCUs. Experiments demonstrate 250× compression of vision models (accuracy loss <3.3%) and 400× compression of audio models (feature error rate <6%), achieving efficient deployment in a moose-monitoring camera (96.7% accuracy) and a child speech wearable (8.44% FER).
Existing auto-tuning approaches suffer from poor optimizer generalizability and high manual design costs in high-dimensional, irregular parameter spaces. Method: This paper proposes a large language model (LLM)-driven paradigm for adaptive optimization algorithm generation. It jointly encodes problem descriptions and structured search-space representations, employs prompt engineering to guide the LLM in dynamically synthesizing customized search strategies, and refines these strategies via an iterative validation framework. Contribution/Results: Evaluated on four real-world high-performance computing scenarios, the generated optimizers outperform state-of-the-art methods by 72.4% on average. Ablation studies quantify the individual contributions of problem description modeling (30.7%) and search-space modeling (14.6%) to overall performance gain. To our knowledge, this is the first work to leverage LLMs for end-to-end optimization algorithm synthesis, overcoming the adaptability limitations inherent in fixed, hand-crafted optimizers.
Reconfigurable optical interferometers face significant challenges in implementing arbitrary unitary transformations when analytical phase decomposition methods—such as the Clements decomposition—are unavailable, particularly for nonstandard or novel circuit architectures. Method: This paper proposes a data-driven automated calibration and programming framework: first, a device-specific end-to-end response model is constructed via supervised learning, eliminating reliance on architecture-dependent analytical models; second, phase control parameters are jointly optimized to directly approximate the target unitary matrix. Contribution/Results: This work pioneers the tight integration of data-driven modeling with physical-layer control, circumventing conventional analytical decomposition algorithms. It substantially enhances programmability for unconventional interferometric architectures. Experimental results demonstrate high fidelity (>99.5%) and strong robustness even in absence of analytical solutions, establishing a general-purpose calibration paradigm for large-scale programmable photonic integrated circuits.
Joint optimization in compound-lens computational imaging systems suffers from heavy reliance on manually initialized optical designs, hindering simultaneous achievement of global optimality and physical realizability. Method: We propose Quasi-Global Synthetic Optimization (QGSO), a novel optical design paradigm comprising two stages: (i) OptiFusion automatically discovers diverse initial optical configurations; (ii) EPJO enables multi-initialization parallel, physics-constrained end-to-end joint optimization, integrating differentiable optical modeling, neural rendering, and gradient-driven co-updating of optics and computation. Contribution/Results: QGSO significantly outperforms conventional stepwise design and existing joint optimization approaches across multiple imaging tasks—achieving substantial PSNR and SSIM gains while eliminating manual initialization bottlenecks. The implementation is open-sourced, facilitating reproducible research in intelligent, physics-informed optical design.
This work addresses the low reusability and insufficient GPU acceleration support of existing stochastic optimal control algorithms. We propose the first modular, plug-and-play open-source CUDA library for real-time stochastic model predictive control. The library decouples controller cores—including MPPI, Tube-MPPI, and Robust MPPI—from user-defined dynamics models and cost functions, enabling seamless integration of custom components via a unified C++/CUDA API. Built upon a tight fusion of model predictive control and path integral optimal control theory, it achieves millisecond-level closed-loop execution on multiple generations of NVIDIA GPUs. Experimental evaluation demonstrates 10–50× speedup over state-of-the-art CPU- and GPU-based implementations, while significantly improving cross-task transferability and development efficiency. The library establishes a general-purpose, high-performance infrastructure for real-time stochastic optimization control of dynamic systems.
This study addresses industrial job shop scheduling by developing a customized model that incorporates hardware constraints and systematically evaluates the performance of quantum annealing (D-Wave), digital annealing (Fujitsu), quantum-inspired algorithms, and classical approaches—including mixed-integer linear programming (MILP) and exact solvers—on platforms such as IBM Quantum. Through a hardware-software co-design methodology, the work demonstrates the practical utility of quantum and quantum-inspired techniques in enhancing the quality of approximate solutions, guiding solver selection, and integrating into classical computational workflows. The research establishes a scalable hybrid solving paradigm for industrial scheduling and validates its feasibility and potential during early-stage proof-of-concept demonstrations.
To enhance the predictive capability of digital twins, efficient acquisition of high-quality experimental data is imperative. This work proposes a novel integration of eigenvalue- and condition number-based computations into the Pyomo.DoE framework via a callback mechanism, enabling rigorous support for eigenvalue-oriented optimal design criteria such as E-optimality and ME-optimality. By establishing a unified abstraction for experimental modeling, the approach selectively targets dimensions in the parameter space that exhibit insufficient information content or numerical instability, seamlessly combining first-principles models with intrusive uncertainty quantification. The proposed method substantially broadens the range of design criteria supported by Pyomo.DoE, reduces user modeling effort, and improves both the efficiency and accuracy of constructing high-fidelity digital twins.
Solving large-scale sparse linear systems in high-performance computing faces significant challenges due to high latency and energy consumption. This work proposes a novel optical analog computing paradigm by, for the first time, mapping general sparse linear systems onto the phase dynamics of coupled laser cavities within an optical Laser Processing Unit (LPU), where the steady-state optical field directly encodes the solution. By integrating laser cavity modeling with LPU simulation, the method demonstrates superior performance over GPU-based Krylov subspace solvers—such as Conjugate Gradient (CG) and GMRES—on standard SuiteSparse multi-band sparse matrices, achieving substantially lower solution latency, higher parallelism, and improved energy efficiency.
Traditional accelerator tuning relies on manually crafted procedures, which are difficult to reuse efficiently after lattice modifications, thereby hindering rapid early-stage design iteration. This work proposes the first end-to-end autonomous algorithm discovery framework driven by a language model agent, integrating large language models, a particle accelerator simulation platform, and an automated feedback loop to iteratively generate and refine tuning algorithms starting from minimal initial code. For the first time, an intelligent agent directly participates in exploring accelerator tuning strategies, successfully producing 16 non-dominated algorithms in the ALS-U accumulator ring model. These algorithms exhibit diverse physical trade-offs between beam capture speed and error correction performance and significantly outperform expert-designed solutions.
This study systematically evaluates the practical applicability of quantum computing in industrial optimization and machine learning. Building upon the QCHALLENGE initiative, we establish a unified benchmarking framework encompassing both superconducting and trapped-ion architectures. We introduce three standardized metric categories and a traffic-light–style decision mechanism to quantitatively compare the performance boundaries of quantum, hybrid, and classical approaches across dimensions including model formulation, scalability, solution quality, runtime, and portability. Our analysis identifies the most promising near-term quantum application scenarios, delineates domains where hybrid strategies offer the greatest feasibility, and clarifies areas where classical methods remain superior, thereby providing a clear roadmap for industrial deployment.