TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of test generation in SWE agent ensembles, specifically the absence of explicit selection losses, incomplete optimization directions, and feedback-free iteration. To overcome these challenges, this work proposes TestGRAD, a framework that formulates test generation as a spatial optimization problem. By introducing differential losses, full CRUD gradients, and failure-mode momentum mechanisms, TestGRAD evolves executable test suites to effectively discriminate among candidate patches. Furthermore, it integrates gradient descent-inspired automated test optimization, memory mining, and behavioral separation objectives to transcend conventional unidirectional generation paradigms. Evaluated on SWE-bench Verified, TestGRAD achieves an 84.2% Pass@1 score, surpassing baseline methods by 3.6 percentage points while compressing historical context by over two orders of magnitude.
📝 Abstract
SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over $100\times$.
Problem

Research questions and friction points this paper is trying to address.

SWE-agent ensemble
test-based selection
test-space optimization
patch selection
test generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Space Optimization
Differential Loss
CRUD Gradients
Failure Pattern Momentum
SWE-Agent Ensemble