Institution profile

Institute for Advanced Algorithms Research

Academic institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

Oct 07, 2026

This study addresses the absence of evaluation benchmarks for assessing AI capabilities in solving fundamental open problems in science by constructing a benchmark comprising 82 unsolved challenges in mathematics and physics. Methodologically, it introduces a novel evaluation framework grounded in authentic scientific literature that operates without predefined ground-truth answers, alongside a multi-evaluator model ensemble and a problem-context modeling mechanism to objectively quantify solution progress. Experimental results demonstrate that GPT-6-Astra achieves the highest resolution rate of 14.0%, significantly outperforming existing open-source and lightweight models. By bridging the gap in evaluating AI-driven frontier scientific exploration, this work establishes a reliable paradigm for assessing the reasoning limits of large language models on open-ended problems.

0 citationsRead paper

Recursive Improvement of a Differentiable Scientific Software Ecosystem

Oct 03, 2026

This study addresses the challenges of extending differential programming, integration complexity, and insufficient evaluation within heterogeneous scientific software ecosystems by proposing an AI agent-driven framework for differentiable scientific software evolution. The framework introduces unified differentiation interfaces and shared resource mechanisms, leveraging AI coding agents to automate the implementation of automatic differentiation. Furthermore, it establishes a closed-loop quality assessment system integrating independent derivative verification, workflow testing, and performance benchmarking to drive recursive software improvement. Experimental evaluations across twenty scientific software packages demonstrate that the proposed approach significantly reduces gradient computation overhead while successfully enabling the efficient reuse of differentiable workflows in applications such as quantum control and thermal design.

0 citationsRead paper

Text2Mem: A Unified Memory Operation Language for Memory Operating System

Sep 14, 2025

Existing LLM agent memory frameworks support only basic operations (e.g., encode, retrieve, delete) and lack advanced capabilities such as merging, promoting, or demoting memories; moreover, memory commands lack formal specification, leading to unpredictable behavior and poor cross-system interoperability. Method: We propose Text2Mem—the first standardized memory operation language—featuring a JSON Schema–defined instruction set and semantic invariants that establish an end-to-end path from natural language directives to deterministic execution. Its three-layer architecture (parsing–validation–adaptation) decouples instruction generation from execution, supporting the full spectrum of operations (encode, retrieve, merge, promote, etc.). A unified execution contract integrates embedding and summarization models, backed by an extensible SQL prototype. Contribution/Results: Text2Mem ensures safety, determinism, and portability across heterogeneous backends. We further introduce Text2Mem Bench, a benchmark suite enabling systematic evaluation of memory operation frameworks.

0 citationsRead paper
Recent publications

Latest Papers

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

Oct 07, 2026

This study addresses the absence of evaluation benchmarks for assessing AI capabilities in solving fundamental open problems in science by constructing a benchmark comprising 82 unsolved challenges in mathematics and physics. Methodologically, it introduces a novel evaluation framework grounded in authentic scientific literature that operates without predefined ground-truth answers, alongside a multi-evaluator model ensemble and a problem-context modeling mechanism to objectively quantify solution progress. Experimental results demonstrate that GPT-6-Astra achieves the highest resolution rate of 14.0%, significantly outperforming existing open-source and lightweight models. By bridging the gap in evaluating AI-driven frontier scientific exploration, this work establishes a reliable paradigm for assessing the reasoning limits of large language models on open-ended problems.

0 citationsRead paper

Recursive Improvement of a Differentiable Scientific Software Ecosystem

Oct 03, 2026

This study addresses the challenges of extending differential programming, integration complexity, and insufficient evaluation within heterogeneous scientific software ecosystems by proposing an AI agent-driven framework for differentiable scientific software evolution. The framework introduces unified differentiation interfaces and shared resource mechanisms, leveraging AI coding agents to automate the implementation of automatic differentiation. Furthermore, it establishes a closed-loop quality assessment system integrating independent derivative verification, workflow testing, and performance benchmarking to drive recursive software improvement. Experimental evaluations across twenty scientific software packages demonstrate that the proposed approach significantly reduces gradient computation overhead while successfully enabling the efficient reuse of differentiable workflows in applications such as quantum control and thermal design.

0 citationsRead paper

Text2Mem: A Unified Memory Operation Language for Memory Operating System

Sep 14, 2025

Existing LLM agent memory frameworks support only basic operations (e.g., encode, retrieve, delete) and lack advanced capabilities such as merging, promoting, or demoting memories; moreover, memory commands lack formal specification, leading to unpredictable behavior and poor cross-system interoperability. Method: We propose Text2Mem—the first standardized memory operation language—featuring a JSON Schema–defined instruction set and semantic invariants that establish an end-to-end path from natural language directives to deterministic execution. Its three-layer architecture (parsing–validation–adaptation) decouples instruction generation from execution, supporting the full spectrum of operations (encode, retrieve, merge, promote, etc.). A unified execution contract integrates embedding and summarization models, backed by an extensible SQL prototype. Contribution/Results: Text2Mem ensures safety, determinism, and portability across heterogeneous backends. We further introduce Text2Mem Bench, a benchmark suite enabling systematic evaluation of memory operation frameworks.

0 citationsRead paper