ranking

Designs and implements methods and systems that produce ordered lists of items by relevance or score, encompassing both coarse-stage candidate selection (粗排) and fine-stage re-ranking (精排) as aspects of 排序. Work includes feature engineering, training and deploying scoring/ranking models, building ranking pipelines for candidate generation and re-ranking, and analyzing ordering quality with appropriate evaluation metrics.

ranking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$227K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Rank-K: Test-Time Reasoning for Listwise Reranking

May 20, 2025
EY
Eugene Yang
🏛️ Johns Hopkins University

Existing neural re-rankers achieve strong performance but suffer from high query-time computational overhead, poor generalization to complex queries, and multilingual support requiring task-specific fine-tuning. This paper proposes Rank-K—the first listwise re-ranker enabling test-time reasoning—where computation is dynamically allocated per query to achieve adaptive refinement. Its core innovation lies in natively integrating reasoning-capable large language models into a listwise ranking framework, coupled with multilingual unified representation learning and contrastive alignment, eliminating the need for fine-tuning to achieve cross-lingual re-ranking. Experiments demonstrate that, applied to BM25 initial rankings, Rank-K outperforms the state-of-the-art RankZephyr by 23% in NDCG@10; when initialized from the strong retriever SPLADE-v3, it yields a 19% gain. Crucially, Rank-K maintains monolingual effectiveness while achieving robust multilingual transfer—without any language-specific adaptation.

Enhancing multilingual retrieval effectiveness with Rank-KImproving efficiency of resource-intensive neural rerankersScaling test-time reasoning for hard query handling

RankLLM: A Python Package for Reranking with LLMs

May 25, 2025
SS
Sahel Sharifymoghaddam
🏛️ University of Waterloo

To address the lack of modularity in LLM-based re-ranking within multi-stage retrieval, poor API reliability, and non-determinism in Mixture-of-Experts (MoE) models, this paper introduces the first open-source Python toolkit specifically designed for re-ranking tasks. Its core is a modular re-ranking framework that integrates prompt analysis, response reliability diagnostics, and MoE behavior tracing, while enabling seamless coupling with Pyserini. The toolkit provides a unified abstraction for interfacing with diverse LLMs (10+ open- and closed-source), and embeds multi-granularity evaluation protocols. Experimental reproduction of state-of-the-art methods—including RankGPT, LRL, and RankVicuna—demonstrates consistent SOTA performance on BEIR and MSMARCO benchmarks. The toolkit significantly enhances configurability, robustness, and reproducibility of re-ranking systems.

Addresses reliability concerns in LLM APIs and MoE modelsDevelops modular Python package for LLM-based document rerankingEnables quick reproduction of results for research and applications

Towards Principled Learning for Re-ranking in Recommender Systems

Apr 05, 2025
QL
Qunwei Li
🏛️ Hechun Medical Technology Co. | Ant Group

Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.

Lack of principles for re-ranker learning processMissing quality measurement for re-ranker outputNeed for consistent principles to improve re-ranker performance

Foundations of the Theory of Performance-Based Ranking

Dec 05, 2024
SP
Sébastien Piérard
🏛️ University of Liège

Existing performance ranking methods in entity evaluation struggle to simultaneously satisfy application-specific preferences and theoretical rigor. Method: This paper establishes the first axiomatic, verifiable general theory framework for performance ranking. Grounded in probability theory and order theory, it formally defines core concepts—including performance objects, satisfaction, and importance—and introduces a performance order satisfying axioms such as ranking consistency, along with constructive procedures for deriving such orders. It further proposes a novel parameterized family of universal ranking scores that unifies classical metrics (e.g., accuracy, recall, F1-score) and rigorously proves that several widely used metrics—including precision—violate the ranking consistency axiom. Contribution/Results: The framework provides the first mathematically rigorous yet practically flexible foundation for performance evaluation in computer vision and machine learning, explicitly characterizing the validity boundaries and intrinsic limitations of reliable ranking metrics.

Establish universal theory for performance-based rankingIntroduce axiomatic definition of performance orderingsPropose parametric family of ranking scores

Generating Diverse Criteria On-the-Fly to Improve Point-wise LLM Rankers

Apr 18, 2024
FG
Fang Guo
🏛️ Westlake University | South China University of Technology | Google

Pointwise large language model (LLM) rankers suffer from limited adherence to standardized comparative guidelines and insufficient capability in holistically evaluating complex passages. To address this, we propose a dynamic multi-perspective evaluation criterion generation method: leveraging prompt engineering to instantiate interpretable, dimension-specific criteria—covering semantics, relevance, structure, and more—in real time, and jointly aggregating scores across these criteria. This mechanism is the first to achieve decomposability, interpretability, and synergistic enhancement in LLM-based evaluation. Evaluated on the BEIR benchmark across eight diverse datasets, our approach significantly improves ranking performance, yielding an average 3.2% relative gain in NDCG@10. Results demonstrate that dynamic, multi-perspective guidance effectively enhances the ranking capability of pointwise LLM rankers.

Inadequate comprehensive analysis for complex passagesNeed multi-perspective criteria to enhance ranking performanceStandardized comparison guidance lacking in LLM rankers

Latest Papers

What's happening recently
View more

This work addresses the limitations of current large language model (LLM)-based rerankers, which suffer from high computational overhead and context-length constraints, while conventional truncation strategies rely on static heuristics that lack dynamic awareness of query relevance. The authors propose a novel approach that leverages an LLM to generate a semantic reference document, serving as a dynamic boundary between relevant and non-relevant documents to guide list truncation. This is combined with either non-overlapping or adaptively stepped overlapping window mechanisms to enable efficient list-wise reranking. Notably, this is the first method to employ LLM-generated reference documents for dynamic truncation, overcoming the constraints of fixed hyperparameters and topic-agnostic heuristics. Evaluated on the TREC Deep Learning benchmark, the approach significantly outperforms existing truncation strategies, achieving up to 66% speedup in both in-domain and out-of-domain settings.

computational overheadcontext lengthlarge language models

This work addresses the entanglement of reranking behavior with retrieval quality in existing evaluation paradigms, which hinders isolated analysis of reranking strategies themselves. To disentangle these factors, the authors propose a model-agnostic, controlled diagnostic framework that constructs fixed evidence pools—each strictly containing eight documents—via clustering on the Multi-News dataset, thereby standardizing inputs to isolate the reranking process. Using BM25 and MMR as interpretable baselines, they systematically evaluate diverse rerankers across 345 clusters. Under conditions that eliminate retrieval variance, the study reveals intrinsic differences among large language model (LLM)-based rerankers in terms of diversity and lexical coverage: some implicitly enhance diversity under high-budget settings while others introduce redundancy; under low-budget constraints, most LLM rerankers consistently underperform baselines and significantly deviate from ideal coverage patterns.

evaluationevidence poolsLLM

This work addresses the challenges of personalized ranking when users lack knowledge of data attributes or struggle to articulate their preferences explicitly. Existing approaches are limited by reliance on a single candidate item selection strategy, which constrains flexibility and user control. To overcome this, the authors propose a visual analytics framework that integrates model-driven active learning with human-driven item selection, establishing—for the first time—a unified interactive item selection space. This space supports six complementary strategies for expressing list-level preferences and enables iterative learning to produce interpretable rankings. A formative user study (N=10) demonstrates the approach’s effectiveness and reveals trade-offs among accuracy, diversity, novelty, transparency, perceived control, and user satisfaction across different selection strategies.

candidate item selectionitem-based rankingpersonalized ranking

This paper addresses the problem of recovering fine-grained item rankings from discrete, coarse-grained user ratings (e.g., 1–5 stars) under unknown, user-specific rating thresholds. We propose a probabilistic ordinal query model that jointly models item scores and personalized user thresholds, with ranking quality measured by Spearman distance. Our key contributions are threefold: First, we establish the necessity—and inherent cost—of modeling inter-user threshold heterogeneity for accurate ranking recovery. Second, we quantify the impact of mismatch between rating and threshold distributions via a quadratic divergence factor. Third, we derive tight Θ(n²) lower bounds on both the required number of users and query complexity to achieve ε-optimal ranking—significantly higher than the O(n log n) bound for pairwise comparison models. Finally, we design an algorithm matching this lower bound up to logarithmic factors.

Addressing the challenge of unknown user thresholds in rating aggregationQuantifying query complexity for achieving near-perfect ranking accuracyRecovering fine-grained item rankings from coarse-grained discrete user ratings

Existing re-ranking methods struggle to model users’ multidimensional intents and often neglect rich multimodal signals such as visual information, thereby limiting improvements in search satisfaction. To address this, this work proposes a user satisfaction–oriented re-ranking framework powered by large language models (LLMs). The approach employs a Query Planner to parse query evolution and disentangle user intents, integrates visual signals generated by a vision-language model (VLM), and leverages a multi-task reinforcement learning–enhanced LLM re-ranker to enable fine-grained, multidimensional relevance assessment. Extensive offline experiments demonstrate significant performance gains over state-of-the-art baselines, and the method has been successfully deployed in a large-scale industrial search system, yielding measurable improvements in online user engagement and satisfaction.

re-rankingrich-media searchuser intent