🤖 AI Summary
This study addresses the issue in large language model (LLM)-based recommendation where misleading historical interactions distort current instructions, resulting in ranking errors. To mitigate this, we propose AIMS, a novel method that introduces an intervention-guided asymmetric margin supervision mechanism. Specifically, AIMS reformulates the elimination of misleading historical influences into asymmetric margin signals, training with a combination of cross-entropy and asymmetric auxiliary losses alongside a frozen reference model. This approach optimizes ranking performance without requiring history editing during inference. Extensive experiments demonstrate that AIMS significantly improves Recall and NDCG metrics across six LLM backbones and multiple benchmarks, validating both its effectiveness and generalization capability.
📝 Abstract
In instruction-guided generative recommendation, LLM-based recommenders need to balance two goals: responding to the user's current request and aligning with the preferences in their interaction history. When the two conflict, history events can override the request. We show that turning the effect of individual history events into supervision faces two obstacles. First, the events that most influence a recommendation are not necessarily the ones that support the target item. Second, removing a misleading event can raise the target's score but a competing item's score even more, so a higher target score alone does not guarantee a better ranking. We propose Asymmetric Intervention-Guided Margin Supervision (AIMS), which converts the effect of removing individual history events into ranking supervision. For training requests already ranked correctly, a frozen reference model identifies request-specific deletions that improve both the target's score and its margin over a competitor near the recommendation cutoff. These margins serve as training targets, while the complete history is retained as input. Training combines cross-entropy with an asymmetric auxiliary loss that penalizes margin shortfalls and routes its gradient only through the competitor score. Inference is unchanged, requiring no history editing or deletion search. Across six LLM backbones on an industrial dataset and two public benchmarks, AIMS improves Recall and NDCG over strong baselines. Ablations support request-specific margins and asymmetric supervision, and the selected deletions preferentially remove constraint-violating history.