Score
Designs, implements, and evaluates methods that take an initial set of candidate items (often with scores or features) and produce a final ordered list by rescoring, reweighting, or applying pairwise/listwise ranking models; builds the reranker model, feature pipelines, and decision logic. Analyzes ranking quality, calibration, diversity, and efficiency, and runs experiments to optimize trade-offs between relevance, fairness, and computational cost.
This work addresses three key challenges in information retrieval reranking: weak reasoning capability, poor interpretability, and insufficient out-of-distribution (OOD) generalization. To this end, we propose Rank1—the first lightweight reranker incorporating test-time computation. Methodologically, we distill structured reasoning traces (>600K samples) from reasoning-oriented large language models (e.g., o1, R1), perform supervised training on MS MARCO, and support prompt-driven zero-shot transfer. The model architecture is designed for promptability, explicit interpretability, and inference efficiency; quantization further reduces computational and memory overhead. Key contributions include: (1) the first application of the test-time computation paradigm to reranking; (2) generation of human-readable, step-by-step reasoning chains as explicit outputs; and (3) state-of-the-art performance across multiple benchmarks with strong OOD generalization.
Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.
Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.
This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
This work addresses the susceptibility of large language models (LLMs) to input order in list reranking, which induces inconsistent preferences and undermines recommendation reliability. The study introduces a multi-level consistency evaluation framework—spanning pairwise preferences, global preference structures, and output stability—to systematically assess the impact of position bias on LLM reranking behavior. Comprehensive experiments across multiple models and datasets reveal that merely enhancing relevance or balancing positional exposure is insufficient to ensure preference consistency. These findings expose an intrinsic instability in LLM-based rerankers that conventional evaluation metrics fail to capture, thereby underscoring the necessity of explicitly modeling preference consistency in reranking systems.
This study investigates whether reasoning-based rerankers, while enhancing retrieval relevance, adversely affect search fairness—particularly across demographic attributes. Leveraging the TREC 2022 Fair Ranking dataset, the authors systematically evaluate six reranking models across diverse retrieval scenarios to analyze the trade-offs between fairness and relevance. The work introduces a novel metric, Attention-Weighted Ranking Fairness (AWRF), to quantify the impact of reasoning mechanisms on fairness. Findings reveal that reasoning-based reranking exhibits no significant effect on fairness, with AWRF consistently ranging between 0.33 and 0.35, despite notable fluctuations in relevance performance. Moreover, the study uncovers persistent fairness disparities along geographic attributes, highlighting an ongoing challenge in equitable information access.
This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.