Scaling Scientific Discovery Environments for Turn-Level Agentic RL

📅 2026-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of process supervision in scientific discovery agents, which stems from the absence of verifiable environments grounded in real-world data. To overcome this limitation, the authors propose SciDisco, a framework that integrates hypotheses, datasets, and implicit evidence graphs via SciThèque and embeds a verifier to construct a process-verifiable scientific discovery environment. Building upon this infrastructure, they introduce DAG-guided trajectory synthesis to generate multi-turn demonstrations and develop DiscoPO, a novel turn-level policy optimization algorithm, enabling reinforcement learning driven by fine-grained, verifiable evidence. This approach establishes the first turn-level process supervision mechanism for scientific discovery tasks, substantially enhancing the accuracy and verifiability of large language model agents in complex scientific reasoning. The resulting SciDisco-14B model achieves state-of-the-art performance on hypothesis-driven data analysis benchmarks.
📝 Abstract
Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciThèque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.
Problem

Research questions and friction points this paper is trying to address.

scientific discovery
process supervision
long-horizon reasoning
real-world scientific data
agent training environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

process-verifiable environments
turn-level credit assignment
DAG-grounded trajectory synthesis
scientific discovery agents
verifier-filtered demonstrations
🔎 Similar Papers
No similar papers found.
Y
Yucheng Xu
Shanghai AI Lab
Keyi Zhang
Keyi Zhang
Lead Compiler Engineer at Efficient Computer
Y
Yuyang Yu
Shanghai AI Lab
M
Min Zhang
Shanghai AI Lab
S
Shiyuan Meng
Shanghai AI Lab
P
Pei Chu
Shanghai AI Lab
Z
Zhongying Tu
Shanghai AI Lab