Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "verification trap" in code generation, a phenomenon wherein generators and verifiers sharing erroneous premises cause test-time selection mechanisms to fail. This work is the first to formally define and quantify this issue, revealing the limitations of coupled evidence and proposing decoupled evidence as a core mitigation strategy. The approach is systematically validated through multi-benchmark evaluations, comparisons across five code models, and a lightweight predictor that operates without gold-standard labels. Results demonstrate that the verification trap significantly degrades generation accuracy, while the proposed predictor achieves an AUROC of 0.846. Furthermore, introducing premise-agnostic auditors effectively recovers oracle-level performance, establishing a new paradigm for enhancing the robustness of code generation systems.
📝 Abstract
Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.
Problem

Research questions and friction points this paper is trying to address.

code generation
test-time selection
verification trap
false premises
verifier-generator coupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verification Trap
Test-Time Selection
Code Generation
False Premises
Decoupled Evidence