🤖 AI Summary
This study addresses the unreliability of model-generated actions in terminal agents, which frequently disrupt execution environments and cause task failures. To mitigate this, we propose Mid-Harness, a framework that introduces an intermediate verification mechanism between the model and the execution layer. Specifically, it employs multi-action sampling with pairwise verifiers to filter candidate actions and pioneers the concept of "action scaling," leveraging test-time compute to dynamically optimize boundary allocation. This approach enhances decision reliability without modifying underlying models or frameworks, and its core capabilities can be transferred to lightweight models via knowledge distillation. Experiments demonstrate that Mid-Harness improves Pass@1 success rates from 50% to 68% on TerminalBench-Lite, achieving significant performance gains with reduced token overhead.
📝 Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.