SRHarness: A Harness for Agentic Symbolic Regression

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of long-horizon search in LLM-based agents for symbolic regression, which stems from the lack of structured runtime support. To overcome this limitation, we propose a domain-specific runtime infrastructure that operates independently of the underlying model architecture. By designing composable scientific action interfaces alongside persistent state and trajectory management mechanisms, our approach systematically optimizes the agent search process. Experimental evaluations demonstrate that the proposed framework achieves 93.69% accuracy on the LSR-Transform benchmark, significantly outperforming baseline methods such as SR-Scientist and Codex. These results confirm its effectiveness in enhancing complex symbolic recovery capabilities.
📝 Abstract
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.
Problem

Research questions and friction points this paper is trying to address.

Agentic Symbolic Regression
Runtime Infrastructure
Large Language Models
Search Trajectories
Scientific Discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Symbolic Regression
Runtime Infrastructure
Composable Scientific Actions
Persistent Scientific State
Trajectory Lifecycle Management
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.