🤖 AI Summary
This study addresses the security vulnerability in multi-agent recommender systems, where collaborative reflection mechanisms are susceptible to hijacking by adversarial evidence. We propose the first black-box targeted attack framework that synergizes dual surfaces of semantic construction and structured layout. By exploiting reflection laundering, reflection persistence, and cross-agent propagation, this approach overcomes the limitations of traditional static pipeline attacks and achieves recursive amplification of system-level risks. Experiments on four real-world datasets demonstrate that our method attains state-of-the-art targeted exposure performance, yielding an average E@20 of 0.384, which significantly surpasses the strongest baseline at 0.185. These findings effectively reveal the latent security threats inherent in multi-agent recommender systems.
📝 Abstract
Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts. We term the local rationalisation process reflection laundering, and its system-wide escalation through collaborative reflection collaborative-reflection hijacking. Existing attacks on recommender systems, whether based on interaction-level data poisoning or text-level adversarial perturbations, assume static pipelines and thus cannot exploit this recurrent, multi-agent amplification pathway. To bridge this gap, we first conduct a controlled vulnerability analysis that establishes two exploitable properties underlying collaborative-reflection hijacking: reflective persistence and cross-agent propagation. Then building on these findings, we propose VirusCascade, the first black-box targeted promotion attack that jointly shapes semantic and structural attack surfaces: the former ensures the target item is naturally rationalised as satisfying broad user preferences, the latter positions it for system-wide propagation. Extensive experiments on four real-world datasets across diverse LLM-ARS architectures demonstrate that VirusCascade consistently achieves state-of-the-art targeted exposure under evaluated stealth constraints, reaching a mean E@20 of 0.384 and surpassing the strongest baseline by an absolute margin of +0.185.