Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive evaluation costs of foundation models during agent post-training and the stringent tool-calling demands of end-to-end testing. To overcome these limitations, it proposes a three-dimensional, cold-start-free screening mechanism that identifies decisive steps from successful trajectories by leveraging decisive action probability, patch multiple-choice questions, and prefix-conditioned pass rates. By integrating trajectory replay with Bits Per Byte metrics, this approach transcends conventional single-step or end-to-end evaluation paradigms to efficiently predict model potential. Experiments across ten open-source models demonstrate that all proposed screening metrics exhibit strong rank correlation with post-training SWE-bench performance, effectively translating low-cost benchmarks into reliable proxies for high-cost downstream tasks.
📝 Abstract
How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@$K$ evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@$1$. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.
Problem

Research questions and friction points this paper is trying to address.

base model evaluation
agentic coding
post-training prediction
tool use
SWE-bench
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic coding
base model evaluation
decisive step
trajectory replay
post-training prediction
🔎 Similar Papers
No similar papers found.
Tan Yu
Tan Yu
NVIDIA
LLMRAGCross-modal searchadvertisingvision backbone
A
Alexander Bukharin
NVIDIA
K
Khushi Bhardwaj
NVIDIA
Jennifer Williams
Jennifer Williams
Assistant Professor at University of Southampton (UK)
Speech ProcessingMachine LearningSecurity and Privacy
Zirui Liu
Zirui Liu
Peking University
SystemsAlgorithmsData Structures
J
Jonathan Lingjie Li
NVIDIA
Soumye Singhal
Soumye Singhal
NVIDIA
Deep LearningNLPArtificial Intelligence
J
Joseph Jennings
NVIDIA
Sanjeev Satheesh
Sanjeev Satheesh
Stanford University
Deep learning
Yash Jain
Yash Jain
Essential
Foundation ModelsComputer VisionMulti-modal learning
Ashish Vaswani
Ashish Vaswani
Startup
Deep Learning
V
Venkat Krishna Srinivasan
NVIDIA
M
Matthew Papakipos
NVIDIA
Hyunwoo Kim
Hyunwoo Kim
NVIDIA
Language ModelReasoningCognitionSocial Intelligence
J
Jian Zhang
NVIDIA
Oleksii Kuchaiev
Oleksii Kuchaiev
NVIDIA
machine learningdeep learninggraph theorybioinformatics
Markus Kliegl
Markus Kliegl
NVIDIA
deep learningmachine learningartificial intelligencefluid mechanicsPDE
Mostofa Patwary
Mostofa Patwary
Director, Applied Deep Learning Research, NVIDIA
Natural Language ProcessingLarge Scale Deep LearningHigh Performance ComputingParallel
Mohammad Shoeybi
Mohammad Shoeybi
Senior Director of Applied Research at NVIDIA
Large Language ModelsNLPMulti-Modal ModelsGenerative AI
Bryan Catanzaro
Bryan Catanzaro
NVIDIA
Parallel ComputingMachine Learning
Jonathan Cohen
Jonathan Cohen
NVIDIA
Computer graphicsparallell programmingCUDAdeep learning
Jiantao Jiao
Jiantao Jiao
Director of Research & Distinguished Scientist, NVIDIA; Professor of EECS and Stat, UC Berkeley
Generative AILarge Language ModelsReinforcement LearningMachine LearningAI security