LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive token consumption incurred during candidate trajectory verification in test-time scaling for software engineering agents. To this end, it proposes a zero-execution filtering mechanism grounded in policy hidden states. Specifically, the method leverages hidden-state representations obtained during the generation phase to construct an execution-free filter. By integrating positive-negative sample bank comparisons with a distance-linear scoring fusion scheme, this filter effectively replaces large language models for preliminary screening, enabling efficient candidate retention without additional inference overhead. Evaluations on the SWE-bench benchmark demonstrate that the proposed mechanism reduces verification tokens by 66.6%–81.0% and halves total token consumption, while achieving a Best@16 accuracy of 60.06%.
📝 Abstract
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.
Problem

Research questions and friction points this paper is trying to address.

Software Engineering Agents
Test-time Scaling
Token Efficiency
Trajectory Verification
Execution-free Filtering
Innovation

Methods, ideas, or system contributions that make the work stand out.

LatentSift
Hidden States
Token-Efficient Verification
Software Engineering Agents
Policy-State Filtering