Score
Designs and implements reranking systems that take multiple candidate outputs and score or reorder them according to factuality and evidence signals, including combining verification scores and retrieved evidence into a final selection. Builds the scoring, aggregation, and selection components that promote candidates supported by evidence and demote or exclude candidates that contain unsupported or hallucinated claims.
This work addresses the entanglement of reranking behavior with retrieval quality in existing evaluation paradigms, which hinders isolated analysis of reranking strategies themselves. To disentangle these factors, the authors propose a model-agnostic, controlled diagnostic framework that constructs fixed evidence pools—each strictly containing eight documents—via clustering on the Multi-News dataset, thereby standardizing inputs to isolate the reranking process. Using BM25 and MMR as interpretable baselines, they systematically evaluate diverse rerankers across 345 clusters. Under conditions that eliminate retrieval variance, the study reveals intrinsic differences among large language model (LLM)-based rerankers in terms of diversity and lexical coverage: some implicitly enhance diversity under high-budget settings while others introduce redundancy; under low-budget constraints, most LLM rerankers consistently underperform baselines and significantly deviate from ideal coverage patterns.
Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.
To address insufficient diversity and incomplete coverage of user preferences in multi-generator re-ranking, this paper proposes a comprehensiveness-driven collaborative re-ranking framework. Methodologically, we first formally define and quantify “list comprehensiveness,” then formulate a joint optimization objective balancing preference alignment and comprehensiveness maximization; we further design a learnable complementarity assessment module to enable automatic generator discovery and collaborative scheduling. Our contributions are threefold: (1) the first comprehensiveness metric and optimization paradigm tailored for re-ranking; (2) a learnable modeling mechanism for generator complementarity; and (3) significant improvements in NDCG (+2.1%) and CTR (+1.8%) on two public benchmarks and online A/B tests, empirically validating the framework’s effectiveness in enhancing both recommendation quality and coverage breadth.
This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.
This work addresses the need for verifiable ranking outputs in decision support systems by introducing the Evidence-Certified Candidate Ranking (ECCR) task, which jointly optimizes ranking and evidence generation to ensure that cited text segments are sufficient to reproduce the final decision. To this end, the authors propose ECPO, a listwise policy optimization framework that integrates skeleton alignment rewards, argument consistency constraints, graph-based features, and an evidence loop reward mechanism. They also introduce CertNDCG—a novel evaluation metric—and an unsupervised certainty verifier to enforce coherence between decisions and their supporting evidence. Experiments on the MAVEN-ERE and RAMS datasets demonstrate that the proposed approach significantly outperforms zero-shot, supervised fine-tuning (SFT), and GRPO baselines, achieving state-of-the-art CertNDCG performance across diverse candidate configurations.
This work addresses the limitation of existing e-commerce reranking models that oversimplify multi-constraint queries into a single overall relevance score, often yielding recommendations that only partially fulfill users’ explicit requirements and lack evidential support. To overcome this, we propose REAlign, a novel framework that introduces an explicit demand-evidence alignment mechanism. REAlign models demand types and grounds item-level visible evidence to distinguish between satisfied, violated, and unsupported conditions, thereby constructing demand-oriented contrastive samples and a demand-aware groupwise relative policy optimization method. A multidimensional list utility function—incorporating demand satisfaction, evidence support, violation penalties, and output validity—is designed to transcend conventional aggregated relevance paradigms. Experiments on two e-commerce benchmarks demonstrate that REAlign significantly outperforms strong supervised and policy optimization baselines, particularly improving top-rank quality and compliance of top results, with ablation studies confirming the complementary effectiveness of each component.