Self-Supervised Scaling of Terminal Environments for Scientific Domains

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high construction costs, poor reusability, and absence of semantic verification in training environments for scientific terminal agents by proposing a software-in-the-loop reconstruction framework. The method leverages existing workflows to automatically generate reference outputs and introduces, for the first time, the derivation of verification criteria from execution results to replace manually authored reference solutions. Furthermore, it integrates hierarchical semantic validation with anti-shortcut checking mechanisms to enable unsupervised scaling. Using this framework, 3,000 high-quality trajectories were generated to perform supervised fine-tuning on Qwen3-27B, improving its performance to 53.56% on Terminal-Bench 2.
📝 Abstract
Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs. For each workflow, we execute multiple input configurations and partition cases into public observations and hidden evaluations. Given the instruction, input schema, and public input--output observations, an agent constructs an editable program without access to the source workflow. The candidate is evaluated on hidden configurations against workflow outputs. A hierarchical verifier combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, while public feedback supports iterative revision. The construction admits additional workflows and configurations without authoring a reference solution for each task. We instantiate SWR with 500 workflows and 46 software families across six domains. Across three attempts per task, Qwen3.8-Max solves 838 tasks and produces 1,422 verified trajectories, which we oversample to 3,000 reconstruction-only training examples. Supervised fine-tuning of Qwen3.8-27B improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. These results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents.
Problem

Research questions and friction points this paper is trying to address.

terminal agents
training environments
self-supervised scaling
scientific domains
verifier
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised learning
Software-in-the-loop reconstruction
Terminal agents
Hierarchical verifier
Scientific software workflows
🔎 Similar Papers
No similar papers found.