Score
Design nonconformity scores: construct and evaluate functions that assign scalar nonconformity or conformity values to data points and candidate predictions for use in conformal inference and calibrated uncertainty methods. This includes integrating Bayesian model outputs, defining robust fallback scoring rules, and calibrating scores to satisfy desired coverage or error-rate guarantees.
This study addresses the lack of systematic evaluation of non-conformity scoring functions in conformal prediction, particularly their inability to simultaneously balance prediction set size and coverage accuracy under class-imbalanced conditions. We conduct a comprehensive comparison of mainstream scoring functions, propose a novel variant, and introduce, for the first time, a unified mechanism for evaluating prediction set size. Within the conformal prediction framework, we empirically analyze the performance of these scoring functions on both standard and imbalanced datasets, elucidating their impact on prediction set characteristics. Experimental results demonstrate that the proposed method significantly reduces prediction set size while maintaining valid coverage, with especially pronounced improvements in class-imbalanced tasks.
Traditional conformal prediction (CP) suffers from poor coverage calibration and excessively wide prediction sets under label scarcity. Method: This paper introduces the first semi-supervised extension of CP, proposing a novel nonconformity measure—Nearest-Neighbor Matching (NNM)—that leverages unlabeled data and pseudo-label similarity to dynamically select reference samples. The approach achieves asymptotically valid marginal and conditional coverage even in low-data regimes, without modifying the underlying predictive model. Contribution/Results: Theoretically, we establish convergence guarantees for coverage validity. Empirically, our method significantly improves coverage stability and prediction set compactness across diverse tasks, while demonstrating robustness to distribution shift. It is compatible with mainstream CP frameworks—including split CP and full CP—and supports both marginal and conditional coverage guarantees.
This work addresses the limitation of conventional softmax-based nonconformity scores, which inadequately capture input sample difficulty and consequently yield conformal prediction sets lacking adaptivity and efficiency. To overcome this, the authors propose leveraging Helmholtz free energy in the pre-softmax logit space as a principled measure of uncertainty and introduce a monotonic transformation to reweight nonconformity scores, rendering prediction sets more sensitive to input difficulty. This approach represents the first application of Helmholtz free energy to calibrate scoring functions in conformal prediction, enhancing adaptivity without requiring complex post-processing. Evaluated across multiple datasets and deep architectures in conjunction with four state-of-the-art scoring functions, the method consistently achieves significant improvements in both the efficiency and adaptivity of prediction sets.
This study addresses the misalignment between conventional prediction-oriented scoring and selection objectives in complex target regions—such as intervals, variance-driven sets, multimodal distributions, or multi-condition scenarios. To resolve this, the authors propose using target membership probability as a nonconformity score to directly rank binary selection events, combined with a null-calibrated conformal selection (NCCS) procedure that leverages non-target calibration samples to produce finite-sample valid p-values. This approach is the first to explicitly distinguish prediction from selection tasks and formalizes a target membership scoring principle. While maintaining comparable performance under mean-monotonic targets, it substantially improves selection efficacy in variance-driven and other complex settings. Moreover, in rare-target regimes, NCCS effectively mitigates the anti-conservatism of empirical FDP thresholds, achieving both high power and rigorous control of the false discovery rate in finite samples.
This work addresses the limitations of traditional anomaly detection methods, which rely on heuristic thresholds and lack statistical interpretability and calibration. We propose a conformal prediction–based framework that transforms anomaly scores into statistically valid p-values, enabling rigorous control of the false discovery rate. To facilitate adoption, we introduce nonconform, an open-source Python toolkit that provides the first unified and user-friendly implementation compatible with both scikit-learn and PyOD, incorporating split-conformal inference and efficient calibration strategies robust to distribution shift. Experimental results demonstrate that our approach maintains high detection performance while delivering reliable probabilistic interpretations and strong statistical guarantees.
This work addresses the challenge of conformal prediction in regression settings where the response variable is subject to two-sided truncation. Existing methods struggle to simultaneously achieve marginal and conditional coverage, often failing to provide valid conditional coverage—particularly on easily predictable instances. To overcome this limitation, the authors introduce a novel nonconformity score tailored to truncated data and propose two calibration strategies: one ensuring tight marginal coverage, and another employing a two-stage mechanism that prioritizes conditional coverage, thereby exposing the inherent limitations of marginal coverage in truncation scenarios. Within the conformal prediction framework, the proposed score leverages the structure of truncated observations to deliver finite-sample theoretical coverage guarantees and, under model consistency, attains oracle-like asymptotic performance, substantially outperforming naive adaptations of existing methods.
This work addresses the challenge of achieving conditional coverage in conformal prediction without relying on strong structural assumptions. To this end, it proposes the PIT-CP method, which post-processes arbitrary nonconformity scores via one-dimensional conditional density estimation—using tools such as mixture density networks or conditional normalizing flows—to map them into approximately feature-independent pivotal scores. The approach requires no stringent modeling assumptions while preserving marginal coverage and the geometric structure of prediction sets, and it substantially improves conditional coverage performance. Theoretical analysis provides both deterministic and high-probability upper bounds on the conditional coverage gap, along with formal guarantees on the volume and symmetric difference of the resulting prediction sets.
This work addresses a previously overlooked issue—“configuration shift”—wherein the coverage validity of conformal predictions in large language models (LLMs) degrades significantly under common configuration changes such as prompt templates, decoding temperatures, or weight quantization. The study formalizes this phenomenon, showing that configuration shifts disrupt the consistency between the nonconformity score distributions used during calibration and testing, thereby undermining coverage guarantees. To mitigate this, the authors derive a theoretical lower bound on coverage and propose two practical solutions: a vulnerability-aware calibration ensemble that requires no test data and a boundary-based recalibration method. Extensive experiments across nine LLMs, four datasets, and four scoring mechanisms demonstrate that the proposed approaches effectively restore target coverage, particularly when test samples are scarce or unavailable, without compromising prediction efficiency.