Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) face a fundamental misalignment between supervised fine-tuning (SFT) and the core objective of next-point-of-interest (POI) recommendation: single-label training samples are ill-suited for generating ranked top-k lists. To address this, we propose the first reinforcement learning–based framework for LLM-driven POI recommendation, built upon Proximal Policy Optimization (PPO). Our method reformulates recommendation as sequence generation and introduces a customized reward function that jointly optimizes accuracy, diversity, and position sensitivity—enabling effective ranking optimization from only a single ground-truth POI per instance. Evaluated on multiple real-world datasets, our approach consistently outperforms both prompt-based and SFT-based baselines, achieving state-of-the-art performance across standard top-k recommendation metrics.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language ModelsSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) have been adopted for next point-of-interest (POI) recommendation tasks. Typical LLM-based recommenders fall into two categories: prompt-based and supervised fine-tuning (SFT)-based models. Prompt-based models generally offer greater output flexibility but deliver lower accuracy, whereas SFT-based models achieve higher performance yet face a fundamental mismatch: next POI recommendation data does not naturally suit supervised fine-tuning. In SFT, the model is trained to reproduce the exact ground truth, but each training example provides only a single target POI, so there is no ground truth for producing a top-k list. To address this, we propose Refine-POI, a reinforcement fine-tuning framework for next POI recommendation. We introduce recommendation-driven rewards that enable LLMs to learn to generate top-k recommendation lists using only one ground-truth POI per example. Experiments on real-world datasets demonstrate that Refine-POI achieves state-of-the-art top-k recommendation performance.
Problem

Research questions and friction points this paper is trying to address.

Addresses mismatch between SFT-based models and POI recommendation data
Improves top-k recommendation accuracy with single ground-truth POI
Enhances LLM flexibility and performance for next POI recommendation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement fine-tuning for POI recommendation
Recommendation-driven rewards for top-k lists
State-of-the-art performance with single POI