Score
Designs and implements rescoring or reranking methods that adjust scores for candidate evidence spans, labels, or entailment outputs conditioned on the full claim; this includes systems that take initial extraction outputs and recompute confidence or ranking (sift) to correct unsupported 'supports' verdicts. The work analyzes scoring functions, conditioning strategies, and how to integrate rescoring with decomposition or pipeline stages to recover accuracy lost by earlier decomposition or extraction steps.
This work addresses the “warrant gap” in fact-checking with large language models—where supportive verdicts are often issued without sufficient evidential grounding. To bridge this gap, the authors propose SIFT, a method that preserves full contextual information through claim-conditional re-ranking and automatically validates whether retrieved evidence genuinely entails the claim using a natural language inference (NLI)-based Warrant Support Precision (WSP) metric. By integrating structured evidence decomposition, conditional re-scoring, and automated warrant plausibility assessment, SIFT substantially outperforms standard prompting approaches across four benchmarks, including FEVER and SciFact, achieving accuracy gains of up to 27.6 points and attaining a WSP AUC of 0.92 with a precision of 0.98.
This work investigates the interplay between retrieval and comprehension of in-document extractive evidence by large language models (LLMs) in few-shot learning, specifically examining whether prediction errors stem from retrieval failures and their underlying causes. We conduct error attribution analysis and two-stage ablation studies across five datasets using two representative closed-source LLMs, with human-annotated gold-standard evidence as ground truth. Our key finding—novel and empirically validated—is that prediction errors are strongly coupled with retrieval errors; however, retrieval failure is primarily attributable not to model comprehension deficits, but rather to suboptimal evidence quality (e.g., low clarity or incompleteness). Crucially, improving retrieval accuracy yields significant gains in final prediction performance. These results provide foundational theoretical support and actionable optimization directions for evidence-retrieval–based downstream tasks.
This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.
This work addresses the tendency of models in long-tailed classification to favor frequent classes during inference, which degrades ranking performance on rare classes. From a Bayes-optimal re-ranking perspective, the authors propose a residual decomposition theory that decouples the correction term into a class-specific offset and an input-dependent pairwise interaction term. They theoretically reveal the limitations of the former and establish verifiable conditions under which the latter is effective. Building on these insights, they design REPAIR—a lightweight post-hoc re-ranker that combines a shrinkage-stabilized class term with a linear pairwise term based on competitive features. Experiments across five benchmarks, including image, species, scene, and rare disease diagnosis datasets, demonstrate the method’s efficacy and its ability to accurately identify scenarios where pairwise correction is essential.
This study addresses the evaluation and selection of source-level likelihood ratio (LR) systems for forensic evidence-to-reference comparison tasks by proposing an integrated analytical framework that balances performance and practical feasibility. The authors employ strictly proper scoring rules to quantify how effectively each system updates Bayesian prior odds and present the first systematic comparison among specific-source feature-based, common-source anchored, and unanchored score-based LR approaches. Their findings reveal that specific-source feature-based LRs achieve the highest performance but incur substantial experimental costs, whereas common-source feature-based methods offer strong discriminative power with significantly reduced implementation complexity. All LR systems substantially outperform a baseline relying solely on prior odds. This work thus provides both theoretical grounding and practical guidance for selecting LR systems in forensic practice.
This work addresses the challenges in scientific claim verification posed by highly similar non-evidence paragraphs and distributional shifts between training and inference evidence distributions. To tackle these issues, the authors propose HNR-DAC, a two-stage framework that first employs a Hard-Negative Reranking mechanism to contrast the most confusable non-evidence passages against genuine evidence, followed by Distribution-Aligned Classification, which trains a classifier on reranked results aligned with the inference-time evidence distribution. By innovatively integrating hard-negative reranking with distribution alignment, the approach significantly enhances joint performance in evidence retrieval and claim verification. On NLPCC 2026 Task 10 Track 2, the model achieves a leading average score of 95.13% (Hit@3: 97.21%, Macro-F1: 95.79%, Joint@3: 94.47%), with an official test Macro-F1 of 93.05%, ranking third overall.
This study systematically investigates the effectiveness boundaries of anchor-based pointwise large language model (LLM) reranking methods, with a focus on their dependence on retriever quality, statistical scope, and anchor design. Through replication of GCCP/PAGC and controlled component-level experiments, the authors find that performance gains primarily stem from the contrastive scoring mechanism rather than sophisticated anchor construction or fusion with traditional relevance scores. They propose a simpler sentence-interleaving anchor strategy that matches or surpasses the original methods. Empirical results demonstrate that this approach yields substantial improvements over weak retrievers such as BM25, but offers limited gains when applied to strong dense retrievers like E5. Moreover, the simplified design exhibits greater robustness and generalizability across diverse retrieval settings.
Existing post-hoc calibration methods still exhibit label-dependent reliability discrepancies at the same confidence level, often leading systems to overtrust incorrect predictions or erroneously discard correct ones. This work proposes Label-level Monotonic Reliability Projection (MRP), which learns label-specific monotonic functions to map calibrated confidences to more reliable correctness indicators while preserving the original predicted labels and class probabilities. By leveraging these refined reliability estimates, MRP re-ranks fixed predictions according to residual risk. Introducing, for the first time, label-level modeling of residual reliability, MRP achieves superior performance compared to global confidence remapping through fine-grained monotonic projections. Evaluated on six information access relevance datasets, MRP significantly improves reranking effectiveness and fallback utility while maintaining full-coverage accuracy and expected calibration error (ECE) unchanged.