Score
Designs and implements end-to-end recommendation pipelines that generate and retrieve candidate items or entities via candidate generation, retrieval and recall methods, and that apply ranking and reranking models and sourcing strategies. Builds and analyzes candidate evaluation and assessment frameworks (including skills or technical assessments where relevant) to measure recall, relevance and downstream quality and to optimize candidate-generation and ranking strategies.
This work addresses the challenge of deploying a shared retrieval backbone in industrial systems, where balancing performance and deployment flexibility across multiple downstream tasks remains difficult. To overcome the limitations of conventional approaches that rely on a single optimal checkpoint, the authors propose a multi-stage optimization framework that tailors component-level and hybrid-stage configuration strategies to the distinct performance characteristics of dense retrievers and rerankers throughout training. This approach significantly enhances the adaptability of the shared backbone and improves overall retrieval effectiveness. End-to-end evaluation demonstrates that the resulting shared retrieval service has been successfully deployed across multiple industrial applications, delivering substantial gains in both system performance and scalability.
To address insufficient diversity and incomplete coverage of user preferences in multi-generator re-ranking, this paper proposes a comprehensiveness-driven collaborative re-ranking framework. Methodologically, we first formally define and quantify “list comprehensiveness,” then formulate a joint optimization objective balancing preference alignment and comprehensiveness maximization; we further design a learnable complementarity assessment module to enable automatic generator discovery and collaborative scheduling. Our contributions are threefold: (1) the first comprehensiveness metric and optimization paradigm tailored for re-ranking; (2) a learnable modeling mechanism for generator complementarity; and (3) significant improvements in NDCG (+2.1%) and CTR (+1.8%) on two public benchmarks and online A/B tests, empirically validating the framework’s effectiveness in enhancing both recommendation quality and coverage breadth.
To address low matching accuracy and inefficiency in multi-position concurrent recruitment, this paper proposes an end-to-end intelligent recommendation method that integrates large language model (LLM)-driven semantic understanding with graph-structured similarity computation. We innovatively construct a dual-perspective, multimodal embedding representation—jointly modeling candidates and positions—to unify resumes and job descriptions into a shared semantic space; further, we employ graph neural networks to capture cross-entity relational dependencies, enabling dynamic and interpretable multi-vacancy collaborative matching. Our approach is the first to deeply fuse LLM-powered fine-grained semantic parsing with graph-structural similarity measurement. Evaluated on a real-world recruitment dataset, it achieves an average 32.7% improvement in matching precision and recall, while reducing initial screening time by over 60%.
Under information overload, the retrieval stage in recommender systems has long been underappreciated and lacks systematic investigation. This paper presents the first comprehensive survey of retrieval in industrial multi-stage recommendation pipelines, focusing on three core aspects: user-item similarity modeling, efficient indexing mechanisms (e.g., vector search and inverted indices), and training optimization techniques—including dual-tower architectures, contrastive learning, and negative sampling. We introduce a unified evaluation benchmark spanning three public datasets and integrate insights from leading industry practitioners to holistically characterize deployment practices, performance bottlenecks, and engineering challenges. Our work fills a critical gap in the systematic analysis of retrieval and provides both theoretical foundations and practical paradigms for designing accurate, efficient, and production-ready retrieval components within cascaded recommendation systems.
This work addresses the limitations of traditional multi-stage retrieval systems, which suffer from error propagation due to misaligned stage-wise objectives, and end-to-end generative models, whose effectiveness is hindered by the inefficiency of autoregressive decoding. To bridge this gap, the authors propose DaV-Gen, a novel framework that introduces speculative decoding to information retrieval for the first time. DaV-Gen employs a unified “draft-and-verify” mechanism that jointly performs non-autoregressive candidate drafting and generative fine-grained verification. The model is trained with a combined objective integrating contrastive and fusion losses, effectively merging vector similarity and generative likelihood scores while leveraging a structured vector space for enhanced efficiency. This approach preserves the expressive power of generative models while significantly accelerating inference, achieving both the efficiency of sparse retrieval and the accuracy of generative ranking.
This work addresses the limitations of traditional industrial recommendation systems, which rely on multi-stage cascaded architectures suffering from redundant feature processing, pipeline complexity, and excessively long service chains. The authors propose a unified generate-then-rank framework that encodes user history once and autoregressively generates semantic ID candidates, followed by item-level ranking based on shared encoder states. Innovatively, fine-grained preferences from a high-capacity teacher ranker are distilled into the single model via Rollout distillation, enhancing ranking quality without increasing online inference cost. Evaluated on Yandex Music at scale, the proposed model replaces a legacy cascade of over 15 components with comparable latency while significantly boosting active user count by 1.41%.
This work addresses the challenge in LongEval-RAG tasks where responses must be strictly grounded in a given set of candidate documents. To this end, the authors propose a candidate-constrained retrieval-augmented generation (RAG) system that integrates rule-based chunking, query expansion, pseudo-relevance feedback, reciprocal rank fusion, MiniLM sentence-level reranking, and citation-aware evidence aggregation, complemented by deterministic provenance tracing and a neural sentence selection mechanism. Experimental results demonstrate that the proposed rule-MiniLM variant significantly outperforms baselines across multiple metrics—including BERTScore, retrieval precision, information point coverage, and human evaluation—thereby validating the effectiveness of combining rule-based chunking with neural sentence selection. The study further underscores the critical role of multi-metric evaluation in diagnosing and advancing RAG system performance.