Score
Designs, implements, and evaluates systems and algorithms that generate and rank personalized item suggestions for users by modeling user preferences, item attributes, and contextual signals; work includes candidate generation, scoring/ranking, diversity and calibration, online/offline evaluation, and handling scalability, cold-start, and feedback dynamics.
This paper addresses scalability, real-time responsiveness, and trustworthiness challenges hindering the industrial deployment of recommender systems in e-commerce, healthcare, and finance. It systematically surveys technical advances from 2017 to 2024, integrating classical paradigms—such as collaborative filtering and content-based filtering—with state-of-the-art approaches, including graph neural networks, reinforcement learning, and large language models. Methodologically, it introduces the first “theory–industrial practice” mapping framework and proposes a unified evaluation paradigm incorporating fairness, explainability, and cross-domain transferability. The contributions include a comprehensive taxonomy covering 12 recommendation paradigms across 8 major application domains, an industrial decision-making guide for algorithm selection, and the open-sourcing of multiple toolkits and benchmark datasets. These resources bridge academic research and industrial implementation, significantly advancing interdisciplinary collaboration and reproducible system development.
This work addresses the insufficient collaboration and cumulative bias arising from disjoint modeling of recommender and user agents in recommendation systems. We propose the first collaborative optimization framework explicitly designed for a dual-agent closed-loop feedback paradigm. Methodologically, we establish a bidirectional iterative feedback mechanism: the recommender agent generates recommendations and observes responses from the user agent, while the user agent dynamically refines its preference representation based on feedback; both agents co-evolve via an LLM-driven, interpretable interaction protocol. Our key contribution is the first formalization of the recommender–user dual-agent closed-loop feedback process, jointly optimizing recommendation quality and mitigating bias—without exacerbating popularity or positional biases. On three benchmark datasets, our approach achieves average improvements of 11.52% in recommendation accuracy over a recommender-only baseline and 21.12% over a user-only baseline, significantly enhancing fidelity in user behavior modeling.
This paper investigates the long-term impact of feedback loops in recommender systems on individual behavior and collective market dynamics, focusing on the tension between personalization and diversity in online retail. We propose a reproducible simulation framework grounded in real-world Amazon e-commerce data, wherein multiple recommendation algorithms are periodically retrained to model sustained user–system interaction over time. Our analysis reveals a pervasive paradox: while individual-level diversity—measured by apparent interest divergence—increases, collective-level homogenization intensifies, as actual purchase behavior converges toward popular items, exacerbating item popularity skew and significantly eroding market-level diversity. The core contribution is the first systematic identification and quantification of a structural trade-off between individual and collective diversity, providing both theoretical grounding and empirical benchmarks for designing bias-mitigating, sustainable recommender systems.
Realistic simulation of recommender systems requires synthetic data generation methods that faithfully reproduce both user behavioral heterogeneity and social influence effects. To address this, we propose HYDRA—a novel probabilistic generative model that jointly captures (i) user community structure (reflecting similar adoption patterns), (ii) multimodal item popularity distributions, and (iii) user engagement levels. HYDRA employs a probabilistic graphical framework to parameterize user–item interaction intensities, enabling controllable, diverse, and high-fidelity simulation of their synergistic effects. Compared to existing approaches, HYDRA significantly improves the reproducibility of social influence propagation and behavioral heterogeneity. Extensive experiments across multiple benchmark datasets demonstrate that HYDRA-generated data closely matches real-world statistics—including distributional properties, long-tail item popularity, and fine-grained interaction patterns. This makes HYDRA a reliable foundational tool for controlled analytical studies, robustness evaluation, and stress testing of recommender systems.
This paper addresses core challenges in leveraging user reviews for recommendation systems—namely, inadequate review text modeling, weak integration with explicit ratings, and insufficient interpretability—and proposes the first unified taxonomy for review-enhanced recommendation, systematically surveying representative works from 2014 to 2024. Methodologically, it integrates BERT/LSTM encoders, graph neural networks, attention mechanisms, and multi-task learning to jointly model review semantics, user-item interactions, and fine-grained features (e.g., attribute-level preferences). Its contributions are threefold: (1) it formally defines three emerging research frontiers—multimodal fusion, multi-criteria rating modeling, and ethics-aligned recommendation; (2) it identifies critical bottlenecks in model generalizability, robustness to sparse reviews, and attribution-based interpretability; and (3) it outlines theoretically grounded yet practically feasible future research directions.
Contemporary recommender systems personalize content using user attributes or inferred data but suffer from poor explainability and auditability, hindering users’ ability to make informed privacy decisions and undermining algorithmic accountability. This paper introduces the first lightweight, privacy-preserving, end-user–oriented algorithmic auditing paradigm: an interactive sandbox that enables users to actively formulate hypotheses—via synthetically generated user profiles and behavioral data—and observe system responses (e.g., ad delivery) in real time, transforming black-box attribution into a verifiable hypothesis-testing process. The approach integrates synthetic data modeling, a user-facing sandbox interface, and an A/B-style response observation framework. A user study demonstrates significant improvements in users’ comprehension of recommendation logic, attribution accuracy, and privacy decision-making capability; in advertising scenarios, hypothesis validation success reached 92%.
This study addresses the challenge of information overload in massive product review corpora, which often impedes users from efficiently accessing personalized, salient insights. To tackle this issue, the authors propose an end-to-end framework for personalized review presentation that uniquely integrates aspect-level user preference modeling, multi-source sentiment-aware ranking, and large language model (LLM)-driven customized summarization. By leveraging fine-grained semantic matching and synthesizing sentiment signals, the approach effectively selects and summarizes reviews aligned with individual user interests. Extensive evaluation on an Amazon smartphone review dataset and a user study involving 70 participants demonstrates that the proposed method significantly outperforms existing baselines, yielding substantial improvements in relevance, user satisfaction, decision confidence, and reading efficiency.
This study addresses the limitations of prevailing static assumptions in recommender systems research, which fail to capture the long-term dynamic effects of feedback loops on individual diversity and collective demand distributions. The authors propose a dynamic simulation framework that integrates implicit feedback, periodic model retraining, probabilistic user adoption of recommendations, and system heterogeneity, validated through experiments on real-world retail and music streaming datasets. Their findings reveal that while higher recommendation adoption rates initially boost individual diversity, they ultimately lead to its decline over time—a phenomenon termed the “diversity illusion.” Moreover, the concentration of popularity in collective demand distributions intensifies over time, with the extent varying by model architecture and domain. This work moves beyond static evaluation paradigms to systematically uncover the divergent mechanisms governing individual and collective diversity under feedback loops.
This work addresses the challenges of personalized ranking when users lack knowledge of data attributes or struggle to articulate their preferences explicitly. Existing approaches are limited by reliance on a single candidate item selection strategy, which constrains flexibility and user control. To overcome this, the authors propose a visual analytics framework that integrates model-driven active learning with human-driven item selection, establishing—for the first time—a unified interactive item selection space. This space supports six complementary strategies for expressing list-level preferences and enables iterative learning to produce interpretable rankings. A formative user study (N=10) demonstrates the approach’s effectiveness and reveals trade-offs among accuracy, diversity, novelty, transparency, perceived control, and user satisfaction across different selection strategies.
To address the challenge of modeling personalized preferences for cold-start users in music recommendation—due to severe sparsity of user-item interaction data—this paper proposes an Attribute-aware Pairwise Comparison Decision Tree (APCDT) preference elicitation method. At each decision tree node, APCDT jointly incorporates item attribute preferences and pairwise comparison feedback, while employing a personalized query strategy to dynamically select the most informative comparison pair for efficient preference inference. Its key innovations lie in integrating attribute-based prior knowledge into the pairwise comparison framework and leveraging multi-dimensional semantic information to enhance both user clustering and node-wise decision making. Extensive experiments on multiple real-world music datasets demonstrate that APCDT achieves 12.6%–23.4% improvements in Recall@10 over baseline methods, using only 5–8 average comparisons per user—effectively mitigating the cold-start problem while balancing query efficiency and recommendation accuracy.
This work addresses the challenge of dynamically adapting personalized exercise recommendations to learners’ evolving skill levels in large-scale online education. The authors propose a contextual Thompson sampling-based multi-armed bandit approach that integrates learner characteristics and historical performance to perform real-time Bayesian inference of skill proficiency. By optimizing exercise sequences with the explicit objective of maximizing skill gain, the method tailors recommendations to individual learning trajectories. Experiments on real-world data from a mathematics tutoring platform demonstrate that the proposed approach significantly enhances skill acquisition, effectively accommodates individual differences, and simultaneously identifies high-impact exercises and at-risk learners requiring intervention. The solution exhibits strong scalability and practical utility for real-world educational applications.