What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "valence blindness" phenomenon in the ranking stage of existing sign-aware recommender systems, where conventional evaluation metrics fail to effectively identify negative content despite the utilization of negative feedback. To overcome this limitation, this work pioneers a sign-aware evaluation framework that analyzes embedding spaces via linear probing and introduces a novel family of metrics, including Signed Recall, designed to penalize negatively signed items appearing at top-ranked positions. Furthermore, the effectiveness of these metrics as training signals is empirically validated. The research reveals the inadequacy of mainstream models in shielding users from undesirable content. By incorporating a proposed auxiliary loss function, the models are successfully guided toward genuine valence awareness, thereby redefining performance benchmarks in the field of sign-aware recommendation.
📝 Abstract
Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically valence-blind. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-K recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4
Problem

Research questions and friction points this paper is trying to address.

sign-aware recommendation
negative feedback
evaluation metrics
valence-blind
recommendation evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sign-aware Recommendation
Signed Evaluation Metrics
Valence-blind Paradox
Linear Probing
Auxiliary Loss
M
Minchan Kim
Graduate School of Data Science, Seoul National University, Seoul, Republic of Korea
J
Jungmin Hwang
Department of Data Science, Seoul National University of Science and Technology, Seoul, Republic of Korea
Hyunwoo Park
Hyunwoo Park
Associate Professor, Graduate School of Data Science, Seoul National University
Supply NetworksDigital InnovationPlatform StrategyNetwork VisualizationVisual Analytics