Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of uncertainty quantification in large language models under semantically equivalent paraphrases. We propose a paraphrase-robust uncertainty quantification framework that leverages a lightweight hidden-state-based surrogate model to aggregate predictions and construct nonconformity scores. By integrating paraphrase-aware scoring with conformal prediction techniques, the framework maintains rigorous theoretical coverage guarantees even when only test-time paraphrasing is applied. Experimental results across seven benchmarks demonstrate that our method consistently produces compact prediction sets with empirical coverage closely approximating the target level. These findings indicate that the proposed approach significantly enhances both the semantic robustness and stability of uncertainty quantification for large language models.
📝 Abstract
Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings. Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores. Under score exchangeability, conformal calibration retains marginal coverage. This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition. We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively. Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting. Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording. Code is available at https://github.com/Raina-Xin/PA_Score.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Quantification
Large Language Models
Conformal Prediction
Paraphrase Robustness
Predictive Confidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conformal Prediction
Uncertainty Quantification
Paraphrase-Aware Scoring
Large Language Models
Proxy Model