Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic overconfidence exhibited by large reasoning models following alignment training, which undermines uncertainty quantification (UQ) in black-box settings. We provide the first theoretical proof that alignment suppresses output variability. To overcome this limitation, this work proposes a prompt-level relaxation operator to approximate the optimal policy and introduces J4U, a jailbreak technique that recovers these theoretical behavioral characteristics in black-box environments, thereby enabling reliable UQ. Extensive experiments conducted across four models and three datasets demonstrate that our approach reduces expected calibration error (ECE) by up to fivefold and expands statistically significant coverage sixfold, substantially outperforming existing baselines.
📝 Abstract
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
Problem

Research questions and friction points this paper is trying to address.

Black-Box Uncertainty Quantification
Large Reasoning Models
Overconfidence
Alignment
Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box Uncertainty Quantification
Large Reasoning Models
Prompt-level Relaxation Operators
Jailbreak for Uncertainty (J4U)
Calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lucas Biechy
Petscraft, Inria; Université Paris-Saclay; INSA CVL
Cédric Eichler
Cédric Eichler
Associate professor of Computer Science, INSA Val de Loire
privacygraph databasessoftware architecture
A
Adrien Boiret
Petscraft, Inria; INSA CVL; Université d’Orléans; LIFO
N
Nicolas Anciaux
Petscraft, Inria; Université Paris-Saclay; INSA CVL