implement collaborative filtering

Design and implement item-item and co-occurrence–based collaborative filtering systems that compute similarity or co-occurrence scores between items (item-KNN) and use those scores to recommend, cluster, or rank items based on observed user behavior. Build practical pipelines that incorporate vote- or count-based features, handle sparse and noisy feedback, and support next-item or session-based prediction scenarios where explicit user vectors may be unavailable.

implementcollaborativefiltering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Comprehensive Survey on Retrieval Methods in Recommender Systems

Jul 11, 2024
JH
Junjie Huang
🏛️ Shanghai Jiao Tong University | China Merchants Bank Credit Card Center

Under information overload, the retrieval stage in recommender systems has long been underappreciated and lacks systematic investigation. This paper presents the first comprehensive survey of retrieval in industrial multi-stage recommendation pipelines, focusing on three core aspects: user-item similarity modeling, efficient indexing mechanisms (e.g., vector search and inverted indices), and training optimization techniques—including dual-tower architectures, contrastive learning, and negative sampling. We introduce a unified evaluation benchmark spanning three public datasets and integrate insights from leading industry practitioners to holistically characterize deployment practices, performance bottlenecks, and engineering challenges. Our work fills a critical gap in the systematic analysis of retrieval and provides both theoretical foundations and practical paradigms for designing accurate, efficient, and production-ready retrieval components within cascaded recommendation systems.

Enhances indexing mechanisms for efficient retrievalImproves similarity computation between users and itemsSurvey explores retrieval methods in recommender systems

Review-based Recommender Systems: A Survey of Approaches, Challenges and Future Perspectives

May 09, 2024
EH
Emrul Hasan
🏛️ Toronto Metropolitan University | Vector Institute | York University | Royal Bank of Canada

This paper addresses core challenges in leveraging user reviews for recommendation systems—namely, inadequate review text modeling, weak integration with explicit ratings, and insufficient interpretability—and proposes the first unified taxonomy for review-enhanced recommendation, systematically surveying representative works from 2014 to 2024. Methodologically, it integrates BERT/LSTM encoders, graph neural networks, attention mechanisms, and multi-task learning to jointly model review semantics, user-item interactions, and fine-grained features (e.g., attribute-level preferences). Its contributions are threefold: (1) it formally defines three emerging research frontiers—multimodal fusion, multi-criteria rating modeling, and ethics-aligned recommendation; (2) it identifies critical bottlenecks in model generalizability, robustness to sparse reviews, and attribution-based interpretability; and (3) it outlines theoretically grounded yet practically feasible future research directions.

Analyzing textual reviews to enhance recommendation performanceExploring future directions like multimodal data integrationSurveying review-based recommender systems approaches and challenges

Flexible Generation of Preference Data for Recommendation Analysis

Jul 23, 2024
SM
Simone Mungari
🏛️ University of Calabria | ICAR-CNR | Revelis SRL | University of Udine

Realistic simulation of recommender systems requires synthetic data generation methods that faithfully reproduce both user behavioral heterogeneity and social influence effects. To address this, we propose HYDRA—a novel probabilistic generative model that jointly captures (i) user community structure (reflecting similar adoption patterns), (ii) multimodal item popularity distributions, and (iii) user engagement levels. HYDRA employs a probabilistic graphical framework to parameterize user–item interaction intensities, enabling controllable, diverse, and high-fidelity simulation of their synergistic effects. Compared to existing approaches, HYDRA significantly improves the reproducibility of social influence propagation and behavioral heterogeneity. Extensive experiments across multiple benchmark datasets demonstrate that HYDRA-generated data closely matches real-world statistics—including distributional properties, long-tail item popularity, and fine-grained interaction patterns. This makes HYDRA a reliable foundational tool for controlled analytical studies, robustness evaluation, and stress testing of recommender systems.

Generating flexible synthetic preference data for recommendation systemsModeling item popularity and user engagement via probability distributionsSimulating user communities with similar item adoption patterns

SC-Rec: Enhancing Generative Retrieval with Self-Consistent Reranking for Sequential Recommendation

Aug 16, 2024
TK
Tongyoung Kim
🏛️ Yonsei University | University of Illinois at Urbana-Champaign

In generative recommendation, inconsistent outputs for identical user histories arise from discrepancies between prompt templates and item indexing schemes, limiting sequential recommendation performance. To address this, we propose a generative retrieval framework that jointly incorporates heterogeneous item indexing and multi-template prompting to leverage large language models (LLMs) for candidate generation. We further introduce the first self-consistency–based re-ranking mechanism for generative recommendation, which jointly models dual-path preferences—textual semantics and collaborative signals—via voting and confidence-weighted aggregation over multi-source LLM generations. Evaluated on three real-world datasets, our method significantly outperforms state-of-the-art approaches, achieving up to a 12.7% improvement in Recall@10. This work marks the first successful integration and trustworthy ranking of multi-source heterogeneous knowledge—spanning semantic and collaborative modalities—within a generative recommendation paradigm.

Addresses inconsistency in generative recommender outputs from varying promptsEnsures consistent recommendation quality across different template-index combinationsModels diverse prompt templates and indices as complementary knowledge sources

Combining Social Relations and Interaction Data in Recommender System With Graph Convolution Collaborative Filtering

Jun 03, 2025
TT
Tin T. Tran
🏛️ VSB-Technical University of Ostrava | Ton Duc Thang University

To address the challenges of data sparsity, noise interference, and ineffective fusion of social influence with collaborative signals in social recommendation, this paper proposes a Robust Graph Convolutional Collaborative Filtering framework (R-GCCF). R-GCCF jointly models the user-item interaction graph and the social relation graph. It introduces, for the first time, an adaptive input denoising mechanism to suppress noise arising from sparse interactions, and enables dynamic, weighted integration of social influence and collaborative similarity within a unified GCN architecture. Additionally, it incorporates social regularization and confidence-weighted interaction modeling. Extensive experiments on multiple public benchmarks demonstrate that R-GCCF consistently outperforms state-of-the-art baselines—including NGCF, LightGCN, and SocialLGN—achieving absolute improvements of 12.7% in Recall@20 and 9.3% in NDCG@20, thereby validating its effectiveness and robustness.

Addressing noisy data challenges in user similarity and social influenceCombining social relations and interaction data for better recommendationsImproving recommendation accuracy using graph convolution collaborative filtering

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively integrating user–item and item–item collaborative filtering to enhance Top-N recommendation performance while maintaining computational efficiency. The authors propose a weighted similarity ensemble method based on shared embeddings, which, for the first time, unifies both recommendation pathways within a single framework. By sharing user and item embeddings across strategies, the approach simplifies model architecture and eliminates the need for separate hyperparameter tuning for each pathway, thereby substantially reducing deployment complexity. Experimental results demonstrate that the proposed method achieves competitive recommendation accuracy across multiple datasets and exhibits robust performance in scenarios favoring different collaborative filtering paradigms.

Collaborative FilteringItem-Item SimilarityRecommender Systems

Existing vector quantization–based item indexing methods struggle to handle the highly skewed and non-stationary item distributions prevalent in streaming recommendation systems, resulting in low assignment accuracy, imbalanced clustering, and insufficient inter-cluster separation. To address these challenges, this work proposes MERGE, an adaptive hierarchical item indexing paradigm that dynamically constructs clusters from scratch, continuously monitors cluster occupancy in real time, and employs a fine-to-coarse merging strategy to build a hierarchical structure. MERGE substantially improves assignment accuracy, cluster uniformity, and inter-cluster separation. Online A/B tests further demonstrate its significant gains on key business metrics, validating its practical effectiveness in real-world streaming recommendation scenarios.

cluster imbalanceitem indexingnon-stationary distribution

This work addresses the limitations of traditional item-based collaborative filtering and two-tower models, which suffer from rigid truncation strategies and weak interaction modeling, hindering fine-grained user interest capture. The authors propose PI2I, a two-stage retrieval framework: in the first stage, a relaxed truncation threshold expands the candidate set to improve recall; in the second stage, an interactive scoring model replaces inner product computation, and negative samples are constructed from trigger–target item pairs to align training with online inference. By integrating flexible index construction with personalized interaction modeling, PI2I significantly enhances recommendation accuracy. Offline experiments demonstrate superior performance over classical collaborative filtering and parity with two-tower models. Deployed on Taobao’s “Guess You Like” feed, it achieves a 1.05% increase in transaction conversion rate and releases a public dataset containing 130 million interactions.

item-to-item collaborative filteringpersonalizationrecommender systems

A Reproducible and Fair Evaluation of Partition-aware Collaborative Filtering

Dec 18, 2025
DD
Domenico De Gioia
🏛️ Politecnico di Bari | University of Cagliari

This paper addresses inconsistencies, data opacity, and missing baselines in evaluating partition-aware collaborative filtering (CF) models (e.g., FPSR/FPSR+), establishing the first fair and reproducible unified benchmark. Methodologically, it introduces a standardized data partitioning protocol, integrates subgraph-local similarity modeling with fine-grained similarity refinement, and open-sources a full-stack evaluation framework. Contributions include: (1) the first fair and reproducible evaluation of partition-aware CF; (2) empirical identification of an inherent accuracy–coverage trade-off induced by partitioning strategies; and (3) demonstration that FPSR variants significantly outperform mainstream baselines in long-tail scenarios, with their global component design proving critical to performance. Results provide empirical foundations and scalable guidance for partition-aware recommendation systems.

Clarifying accuracy-coverage trade-offs in scalable recommender systemsFair comparison of FPSR/FPSR+ with similarity-based baselinesReproducible evaluation of partition-aware collaborative filtering models

Hot Scholars

JZ

Jizhi Zhang

USTC
RecommendationTrustworthy AILarge Personalized Model
CS

Cem Subakan

Assistant Prof. at Laval University, Computer Science Dept. / Mila, Associate Academic Member
Machine LearningLearning AlgorithmsMachine Learning for Speech and Audio
JZ

Jialu Zhang

Assistant Professor at University of Waterloo
Programming LanguagesSoftware EngineeringLLM for EducationAI for Education
FN

Fatemeh Nazary

Polytechnic University of Bari
Generative AIExplainable AIAdversarial MLTrustworthy AI
JL

Jundong Li

Associate Professor, University of Virginia
AIMachine LearningData MiningGraph Learning