Score
Designs, implements, and evaluates machine‑learning models and pipelines that produce ordered lists of items for retrieval or search-style systems. This includes learning‑to‑rank algorithms, neural and multi‑stage (coarse‑to‑fine, recall + rerank) architectures, personalization and reranking methods, and work on ranking algorithm design, optimization and large‑scale deployment.
Prior surveys on information retrieval (IR) models conflate architectural design with training methodologies, obscuring the intrinsic evolution of structural innovations in relevance modeling. Method: We systematically trace the architectural progression of IR models—spanning backbone feature extractors and end-to-end relevance modeling—from classical BM25 through CNN/RNN-based rankers to modern BERT dual-encoder and interaction-based architectures, ColBERT, Cross-Encoders, and LLM-based retrievers—explicitly decoupling architecture from training strategy. Contribution/Results: We propose the first longitudinal IR-specific architectural taxonomy, explicitly addressing scalability and adaptability challenges in multimodal, multilingual, and emerging application scenarios. Our framework provides an actionable technology roadmap for industrial system selection and rigorously identifies open research questions and future directions for the academic community.
This work addresses three key challenges in information retrieval reranking: weak reasoning capability, poor interpretability, and insufficient out-of-distribution (OOD) generalization. To this end, we propose Rank1—the first lightweight reranker incorporating test-time computation. Methodologically, we distill structured reasoning traces (>600K samples) from reasoning-oriented large language models (e.g., o1, R1), perform supervised training on MS MARCO, and support prompt-driven zero-shot transfer. The model architecture is designed for promptability, explicit interpretability, and inference efficiency; quantization further reduces computational and memory overhead. Key contributions include: (1) the first application of the test-time computation paradigm to reranking; (2) generation of human-readable, step-by-step reasoning chains as explicit outputs; and (3) state-of-the-art performance across multiple benchmarks with strong OOD generalization.
This study addresses the lack of understanding regarding how reranking performance in multi-stage retrieval systems scales with model size and data volume. We systematically investigate pointwise, pairwise, and listwise reranking paradigms across varying model scales and data budgets, revealing for the first time that metrics such as NDCG and MAP follow predictable power-law scaling behaviors. Through experiments employing cross-encoder architectures and diverse loss functions in both in-domain and out-of-domain settings, we demonstrate that models with fewer than 400M parameters can accurately predict the performance of 1B-parameter models, substantially reducing computational costs. In contrast, metrics like MRR exhibit unreliable scaling properties and do not conform to consistent predictive patterns.
This work addresses the challenge of balancing efficiency and expressiveness in large-scale recommender systems. The authors propose a learnable hierarchical indexing structure that jointly optimizes cross-attention mechanisms and residual quantization, achieving high retrieval accuracy while substantially reducing computational overhead. The approach reveals that intermediate index nodes correspond to high-quality data subsets, thereby enabling practical deployment of test-time training in recommendation scenarios. Experimental results demonstrate consistent superiority over strong baselines on both public and proprietary datasets. Notably, the method has been deployed in the advertising systems of Facebook and Instagram, serving billions of users daily and delivering significant gains in both inference efficiency and recommendation performance.
Re-ranking in recommender systems has long suffered from a lack of theoretical foundations and verifiable quality evaluation criteria. To address this, we propose two principled learning principles—convergence consistency and adversarial consistency—establishing, for the first time, an interpretable and generalizable theoretical basis for re-ranking modeling. Building upon these principles, we design a generic consistency regularization training framework that seamlessly integrates with mainstream listwise models (e.g., BERT4Rec, SetRank) without modifying their backbone architectures. Extensive experiments on multiple public benchmarks demonstrate consistent improvements in NDCG@10 by 1.2–3.7%, validating both the universality and effectiveness of our principles. This work fills a critical gap in the field by introducing formal, learnable constraints for re-ranking optimization.
Under information overload, the retrieval stage in recommender systems has long been underappreciated and lacks systematic investigation. This paper presents the first comprehensive survey of retrieval in industrial multi-stage recommendation pipelines, focusing on three core aspects: user-item similarity modeling, efficient indexing mechanisms (e.g., vector search and inverted indices), and training optimization techniques—including dual-tower architectures, contrastive learning, and negative sampling. We introduce a unified evaluation benchmark spanning three public datasets and integrate insights from leading industry practitioners to holistically characterize deployment practices, performance bottlenecks, and engineering challenges. Our work fills a critical gap in the systematic analysis of retrieval and provides both theoretical foundations and practical paradigms for designing accurate, efficient, and production-ready retrieval components within cascaded recommendation systems.
This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.
This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.
This work addresses the inefficiency in end-to-end evaluation of cascaded information retrieval (IR) pipelines caused by redundant computation. It introduces, for the first time, the Trie data structure into IR experimental design to automatically identify and reuse shared sub-pipelines, thereby constructing highly efficient comparative evaluation plans. Implemented within the PyTerrier framework, the approach supports combined evaluation of diverse models, including BM25, MonoT5, and DuoT5. Experiments on the MSMARCO v2 dataset demonstrate a 26% reduction in runtime compared to conventional linear evaluation plans, while user studies confirm the method’s usability and practical utility for IR researchers.
This work addresses the suboptimal recall and high online inference cost of traditional approximate nearest neighbor (ANN) retrieval, which stems from the decoupled training of embeddings and index structures. The authors propose Multi-Facet Learnable Indexing (MFLI), a novel framework that, for the first time, enables end-to-end joint optimization of item embeddings and a hierarchical index structure within a unified architecture. MFLI constructs multi-facet hierarchical codebooks via residual quantization, supports real-time updates, and performs retrieval directly using the learned index—eliminating the need for ANN search during serving. Evaluated on billion-scale real-world data, MFLI achieves significant improvements over state-of-the-art methods: +11.8% recall on interactive tasks, +57.29% in cold-start content delivery, and +13.5% in semantic relevance. Online deployment further demonstrates enhanced user engagement, reduced popularity bias, and improved serving efficiency.