Score
Designs and implements methods and systems that produce ordered lists of items by relevance or score, encompassing both coarse-stage candidate selection (粗排) and fine-stage re-ranking (精排) as aspects of 排序. Work includes feature engineering, training and deploying scoring/ranking models, building ranking pipelines for candidate generation and re-ranking, and analyzing ordering quality with appropriate evaluation metrics.
Existing neural re-rankers achieve strong performance but suffer from high query-time computational overhead, poor generalization to complex queries, and multilingual support requiring task-specific fine-tuning. This paper proposes Rank-K—the first listwise re-ranker enabling test-time reasoning—where computation is dynamically allocated per query to achieve adaptive refinement. Its core innovation lies in natively integrating reasoning-capable large language models into a listwise ranking framework, coupled with multilingual unified representation learning and contrastive alignment, eliminating the need for fine-tuning to achieve cross-lingual re-ranking. Experiments demonstrate that, applied to BM25 initial rankings, Rank-K outperforms the state-of-the-art RankZephyr by 23% in NDCG@10; when initialized from the strong retriever SPLADE-v3, it yields a 19% gain. Crucially, Rank-K maintains monolingual effectiveness while achieving robust multilingual transfer—without any language-specific adaptation.
To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.
Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.
Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.
Pointwise large language model (LLM) rankers suffer from limited adherence to standardized comparative guidelines and insufficient capability in holistically evaluating complex passages. To address this, we propose a dynamic multi-perspective evaluation criterion generation method: leveraging prompt engineering to instantiate interpretable, dimension-specific criteria—covering semantics, relevance, structure, and more—in real time, and jointly aggregating scores across these criteria. This mechanism is the first to achieve decomposability, interpretability, and synergistic enhancement in LLM-based evaluation. Evaluated on the BEIR benchmark across eight diverse datasets, our approach significantly improves ranking performance, yielding an average 3.2% relative gain in NDCG@10. Results demonstrate that dynamic, multi-perspective guidance effectively enhances the ranking capability of pointwise LLM rankers.
This work addresses the limitations of current large language model (LLM)-based rerankers, which suffer from high computational overhead and context-length constraints, while conventional truncation strategies rely on static heuristics that lack dynamic awareness of query relevance. The authors propose a novel approach that leverages an LLM to generate a semantic reference document, serving as a dynamic boundary between relevant and non-relevant documents to guide list truncation. This is combined with either non-overlapping or adaptively stepped overlapping window mechanisms to enable efficient list-wise reranking. Notably, this is the first method to employ LLM-generated reference documents for dynamic truncation, overcoming the constraints of fixed hyperparameters and topic-agnostic heuristics. Evaluated on the TREC Deep Learning benchmark, the approach significantly outperforms existing truncation strategies, achieving up to 66% speedup in both in-domain and out-of-domain settings.
This work addresses the entanglement of reranking behavior with retrieval quality in existing evaluation paradigms, which hinders isolated analysis of reranking strategies themselves. To disentangle these factors, the authors propose a model-agnostic, controlled diagnostic framework that constructs fixed evidence pools—each strictly containing eight documents—via clustering on the Multi-News dataset, thereby standardizing inputs to isolate the reranking process. Using BM25 and MMR as interpretable baselines, they systematically evaluate diverse rerankers across 345 clusters. Under conditions that eliminate retrieval variance, the study reveals intrinsic differences among large language model (LLM)-based rerankers in terms of diversity and lexical coverage: some implicitly enhance diversity under high-budget settings while others introduce redundancy; under low-budget constraints, most LLM rerankers consistently underperform baselines and significantly deviate from ideal coverage patterns.
This work addresses the challenges of personalized ranking when users lack knowledge of data attributes or struggle to articulate their preferences explicitly. Existing approaches are limited by reliance on a single candidate item selection strategy, which constrains flexibility and user control. To overcome this, the authors propose a visual analytics framework that integrates model-driven active learning with human-driven item selection, establishing—for the first time—a unified interactive item selection space. This space supports six complementary strategies for expressing list-level preferences and enables iterative learning to produce interpretable rankings. A formative user study (N=10) demonstrates the approach’s effectiveness and reveals trade-offs among accuracy, diversity, novelty, transparency, perceived control, and user satisfaction across different selection strategies.
This paper addresses the problem of recovering fine-grained item rankings from discrete, coarse-grained user ratings (e.g., 1–5 stars) under unknown, user-specific rating thresholds. We propose a probabilistic ordinal query model that jointly models item scores and personalized user thresholds, with ranking quality measured by Spearman distance. Our key contributions are threefold: First, we establish the necessity—and inherent cost—of modeling inter-user threshold heterogeneity for accurate ranking recovery. Second, we quantify the impact of mismatch between rating and threshold distributions via a quadratic divergence factor. Third, we derive tight Θ(n²) lower bounds on both the required number of users and query complexity to achieve ε-optimal ranking—significantly higher than the O(n log n) bound for pairwise comparison models. Finally, we design an algorithm matching this lower bound up to logarithmic factors.
Existing re-ranking methods struggle to model users’ multidimensional intents and often neglect rich multimodal signals such as visual information, thereby limiting improvements in search satisfaction. To address this, this work proposes a user satisfaction–oriented re-ranking framework powered by large language models (LLMs). The approach employs a Query Planner to parse query evolution and disentangle user intents, integrates visual signals generated by a vision-language model (VLM), and leverages a multi-task reinforcement learning–enhanced LLM re-ranker to enable fine-grained, multidimensional relevance assessment. Extensive offline experiments demonstrate significant performance gains over state-of-the-art baselines, and the method has been successfully deployed in a large-scale industrial search system, yielding measurable improvements in online user engagement and satisfaction.