Score
Designs and implements post‑retrieval re‑ranking methods that update pairwise similarity matrices and item scores by identifying k‑reciprocal nearest neighbors and expanding those neighbor sets to produce refined feature representations or adjusted similarities; builds algorithms that exploit structured neighbor information to suppress retrieval false positives and improve ranking precision.
Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.
Traditional retrieval-then-reranking pipelines suffer from two key limitations: dependency on the quality of initial retrieval and the high computational cost of large language model (LLM)-based rerankers. To address these, we propose Reranker-Guided Search (RGS), the first approach to explicitly incorporate reranker preferences into the retrieval process. RGS constructs a proximity graph over approximate nearest neighbors, then performs greedy path search guided jointly by embedding similarity and gradients of reranker scores—dynamically prioritizing high-potential documents. This breaks the rigid “retrieve-then-rerank” sequential paradigm and enables end-to-end optimization under a fixed reranking budget (100 documents). Evaluated on BRIGHT, FollowIR, and M-BEIR benchmarks, RGS improves Recall@100 by 3.5, 2.9, and 5.1 percentage points, respectively, significantly alleviating the precision-efficiency trade-off.
This paper addresses the limited robustness of feature representations and suboptimal cross-camera retrieval accuracy in person re-identification (re-ID). To this end, we propose a dual-module framework comprising Dynamic Multi-Order Neighborhood modeling (DMON) and Asymmetric Relation Optimization (ARO). DMON employs a graph neural network to adaptively aggregate multi-order neighborhood contextual information, thereby enhancing feature discriminability. ARO introduces asymmetric metric learning to refine the query-gallery distance matrix with fine-grained relational supervision. The two modules are jointly optimized in an end-to-end manner, enabling synergistic improvement of both feature representation quality and indexing performance. Extensive experiments on three major benchmarks—Market-1501, DukeMTMC-reID, and MSMT17—demonstrate consistent and significant improvements over strong baselines: Rank-1 accuracy and mAP increase by 2.3%–3.1% on average. Moreover, the framework exhibits strong generalizability to other re-ID tasks.
This study identifies a performance breakpoint and semantic failure in cross-encoder re-rankers (e.g., ColBERTv2, RankT5) for large-scale document re-ranking: retrieval quality degrades significantly when the candidate set exceeds ~1,000 documents—MRR@10 drops by 12.7% on average, and 38% of top-scoring results exhibit neither lexical overlap nor semantic similarity with the query. Through systematic ablation and scaling experiments, augmented with semantic similarity and lexical matching analyses, we empirically challenge the widely held assumption that re-rankers universally outperform first-stage retrievers. Our key contributions are: (1) establishing the effective scale boundary for cross-encoder re-rankers; (2) revealing their propensity for relevance misjudgment under ultra-large candidate lists; and (3) providing theoretical grounding and practical guidance—along with critical deployment warnings—for integrating re-ranking modules into large-scale retrieval systems.
To address the degradation of retrieval performance caused by noisy edges—i.e., erroneous connections between dissimilar images—in neighborhood graph re-ranking, this paper proposes a graph denoising method based on Continuous Conditional Random Fields (C-CRF). It is the first to introduce C-CRF into visual re-ranking, automatically assessing edge reliability via statistical distance modeling of similarity distributions—eliminating reliance on hand-crafted thresholds. A clique potential function is designed to incorporate local structural constraints, and probabilistic inference enables end-to-end, fine-tuning-free graph optimization. Theoretical analysis and extensive experiments demonstrate strong complementarity with mainstream re-ranking techniques. When integrated with three baseline methods, the approach consistently improves mAP and Recall@k on both landmark retrieval and person re-identification benchmarks, validating its effectiveness and generalizability.
This work addresses the tendency of models in long-tailed classification to favor frequent classes during inference, which degrades ranking performance on rare classes. From a Bayes-optimal re-ranking perspective, the authors propose a residual decomposition theory that decouples the correction term into a class-specific offset and an input-dependent pairwise interaction term. They theoretically reveal the limitations of the former and establish verifiable conditions under which the latter is effective. Building on these insights, they design REPAIR—a lightweight post-hoc re-ranker that combines a shrinkage-stabilized class term with a linear pairwise term based on competitive features. Experiments across five benchmarks, including image, species, scene, and rare disease diagnosis datasets, demonstrate the method’s efficacy and its ability to accurately identify scenarios where pairwise correction is essential.
This work addresses the inefficiency of existing similarity search indexes in metric spaces, which struggle to adapt to the local clustering characteristics of queries, leading to excessive redundant distance computations. The authors propose a region-aware adaptive indexing mechanism that introduces, for the first time, the concept of regional scope. By dynamically maintaining query regions and reusing previously computed distances, the method enables precise pruning. Furthermore, it intelligently detects shifts in query distribution by monitoring pruning effectiveness and recursively schedules searches into subregions to enhance efficiency. Experimental evaluation across five real-world datasets and four diverse workloads demonstrates that the proposed approach reduces distance computations by up to 64% and decreases query latency by as much as 46% compared to baseline methods such as the AV-tree.
This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.
This work addresses the challenge that while extensive query reformulations in retrieval systems can improve recall, they often induce query drift and incur high reranking costs, making efficient utilization under limited inference budgets difficult. The authors propose ReformIR, a framework that treats query reformulation as a first-class feature and jointly optimizes reformulation selection and document filtering within a fixed reranking budget. By leveraging a strong neural reranker as a teacher model to provide online relevance estimates, ReformIR trains a lightweight proxy model to adaptively select high-value reformulations and documents, effectively mitigating drift while enhancing recall. Experiments demonstrate that ReformIR significantly outperforms existing methods on MS MARCO and TREC DL19–22 benchmarks, maintaining performance gains even as the number of reformulations increases, thereby validating the efficacy of feedback-driven reformulation optimization.
Efficient in-place deletion of outdated vectors in dynamic graph indexes is highly challenging, as conventional approaches either degrade search performance or incur costly global reconstructions. To address this, this work proposes MERIT, the first framework enabling in-place deletions with overheads approaching those of insertions. MERIT integrates outgoing and searchable incoming neighbors, employs k_r-MST-based local repair, and introduces a versioned edge invalidation mechanism to preserve local connectivity and multi-path searchability. It supports diverse graph structures, including HNSW and Vamana. Experiments on multiple real-world datasets demonstrate that MERIT achieves 3.02–18.87× faster deletion than state-of-the-art methods while maintaining stable or even incrementally improving recall as deletions accumulate.