🤖 AI Summary
Large language models (LLMs) face a fundamental misalignment between supervised fine-tuning (SFT) and the core objective of next-point-of-interest (POI) recommendation: single-label training samples are ill-suited for generating ranked top-k lists. To address this, we propose the first reinforcement learning–based framework for LLM-driven POI recommendation, built upon Proximal Policy Optimization (PPO). Our method reformulates recommendation as sequence generation and introduces a customized reward function that jointly optimizes accuracy, diversity, and position sensitivity—enabling effective ranking optimization from only a single ground-truth POI per instance. Evaluated on multiple real-world datasets, our approach consistently outperforms both prompt-based and SFT-based baselines, achieving state-of-the-art performance across standard top-k recommendation metrics.
📝 Abstract
Large language models (LLMs) have been adopted for next point-of-interest (POI) recommendation tasks. Typical LLM-based recommenders fall into two categories: prompt-based and supervised fine-tuning (SFT)-based models. Prompt-based models generally offer greater output flexibility but deliver lower accuracy, whereas SFT-based models achieve higher performance yet face a fundamental mismatch: next POI recommendation data does not naturally suit supervised fine-tuning. In SFT, the model is trained to reproduce the exact ground truth, but each training example provides only a single target POI, so there is no ground truth for producing a top-k list.
To address this, we propose Refine-POI, a reinforcement fine-tuning framework for next POI recommendation. We introduce recommendation-driven rewards that enable LLMs to learn to generate top-k recommendation lists using only one ground-truth POI per example. Experiments on real-world datasets demonstrate that Refine-POI achieves state-of-the-art top-k recommendation performance.