Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of distribution-free selective guarantees and low certificate reliability in chain-of-thought verifiers under small calibration budgets. To this end, it investigates abstention-based verifiers and proposes fixed-sequence certificates to achieve distribution-free guarantees in small-sample regimes. Methodologically, this work elucidates the principles underlying effective abstention and introduces certified lower bounds alongside Benjamini-Hochberg lattice conditions. By integrating conformal prediction, residual stream probing, and cross-fitting techniques, a novel certificate algorithm is developed that eliminates the need for monotonicity assumptions. Experimental results demonstrate that the proposed certificates consistently outperform Bonferroni methods in coverage across diverse model signals, significantly improving coverage levels for non-trivial targets.
📝 Abstract
Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an $(α,δ)$-valid procedure that issues a certificate with probability $P_{\rm fire}$ bounds the failure probability of an issued certificate only by $δ/P_{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75π_0$ from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $δ$, because abstention absorbs the failures.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought verification
distribution-free guarantees
small calibration budgets
selective prediction
certified abstention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought verification
distribution-free guarantees
conformal selection
abstention
calibration budget
💼 Related Jobs
No related jobs found.
A
Arjun Balaji
Columbia University