CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a learnability criterion for assessing solver-task compatibility in existing terminal-based task training, where reliance solely on executability and verification fails to ensure appropriate task difficulty. The authors propose CalibForge, a novel system that introduces the concept of a “solver-relative learnable zone” and employs an adversarial task calibration mechanism. By leveraging behavioral divergence across multiple solvers and contrastive relationships between strong-pass and weak-fail outcomes, CalibForge automatically synthesizes high-quality tasks residing within this learnable zone. Integrating multi-solver ensembles, contrastive learning, and verifiable task generation, the method produces 5,431 calibrated tasks, achieving 47.57% accuracy on Terminal-Bench 2.0—surpassing baseline methods by up to 30.04 percentage points—and demonstrating significant improvements on SWE-bench Pro and Doc2Repo.
📝 Abstract
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.
Problem

Research questions and friction points this paper is trying to address.

terminal tasks
solver calibration
learnability
adversarial synthesis
agent training
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial solver calibration
terminal-task synthesis
solver-relative learnability
multi-solver disagreement
contrastive calibration
🔎 Similar Papers
2024-08-10AAAI Conference on Artificial IntelligenceCitations: 30