The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the insufficient robustness and latent error propagation of scientific agents in multi-turn interactions by proposing SciARP, a benchmark that reformulates 620 scientific problems into multi-turn tasks. Through 13 types of perturbation injection and trajectory-level comparative analysis, it systematically evaluates the reliability of eight large language models across reasoning, evidence processing, and conclusion formation. This work is the first to quantify the multi-turn robustness of scientific agents, revealing a decoupling between baseline performance and robustness. Furthermore, it demonstrates that distinct perturbations yield differentiated robustness profiles, while failures exhibit delayed triggering and cascading propagation characteristics that are difficult for models to recover from autonomously.
πŸ“ Abstract
Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations), a benchmark for evaluating scientific agents under scientifically plausible perturbations throughout multi-turn problem solving. \textsc{SciARP} transforms 620 scientific problems into interdependent tasks of 3--13 turns and defines 13 perturbation types spanning problem understanding, evidence processing, reasoning, and conclusion formation. Clean and perturbed versions of each task are independently executed under matched settings, producing paired live trajectories for evaluating both task success and process reliability. Experiments across eight LLMs from four model families reveal three key robustness characteristics. First, different classes of scientific perturbations exhibit distinct robustness profiles and can decouple task progression from scientific reliability: agents may continue advancing through the task even after their information or reasoning has become unreliable. Second, stronger clean-task performance does not necessarily translate into stronger robustness, as models with higher clean-task accuracy can exhibit larger degradation under perturbation. Third, perturbation effects exhibit strong temporal dynamics: they may remain latent for multiple turns before emerging and subsequently propagate through downstream dependencies. Together, these findings show that current scientific agents remain insufficiently robust to scientifically plausible perturbations, with failures often remaining undetected, propagating, and resisting recovery.
Problem

Research questions and friction points this paper is trying to address.

Scientific agents
Multi-turn interactions
Robustness
Perturbations
Large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Agents
Robustness Benchmark
Multi-Turn Interactions
Perturbation Analysis
Large Language Models
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Xiaoting Lyu
Xiaoting Lyu
Beijing Jiaotong University
X
Xinbo Ma
Xi’an Jiaotong University, China
Y
Yufei Han
Inria, France
Hangwei Qian
Hangwei Qian
CFAR, A*STAR, Singapore | Lund University | NTU | USTC
Artificial IntelligenceTrustworthy AITransfer LearningTime SeriesAI for Science
Z
Ziyang Lin
Xi’an Jiaotong University, China
Bin Wang
Bin Wang
Ocean University of China
Time Series ModelingTrajectory MiningHealthcare ManagementAir Traffic Control
B
Bin Wang
Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security, China
W
Wei Wang
Xi’an Jiaotong University, China