AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation

📅 2026-01-02
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the misalignment between offline evaluation metrics and real user behavior in recommender systems, as well as the scarcity of interaction data, by proposing a novel LLM-based agent framework that integrates world models with counterfactual reasoning. The approach learns environmental dynamics from human interaction sequences and leverages counterfactual trajectories to guide agents in reflecting on and refining their decision-making policies, thereby generating high-fidelity synthetic interaction data. By uniquely combining world modeling with counterfactual inference, the framework enables LLM agents to internalize environmental dynamics and proactively align with human decision patterns. Experimental results across multiple datasets demonstrate that the proposed method significantly outperforms existing approaches at both micro-behavioral and macro-metric levels, yielding interactions that more closely resemble authentic user behavior.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionMultiagent Systems: Adversarial AgentsCognitive Modeling & Cognitive Systems: Simulating Human Behavior

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
📝 Abstract
Evaluating recommender systems remains challenging due to the gap between offline metrics and real user behavior, as well as the scarcity of interaction data. Recent work explores large language model (LLM) agents as synthetic users, yet they typically rely on few-shot prompting, which yields a shallow understanding of the environment and limits their ability to faithfully reproduce user actions. We introduce AlignUSER, a framework that learns world-model-driven agents from human interactions. Given rollout sequences of actions and states, we formalize world modeling as a next state prediction task that helps the agent internalize the environment. To align actions with human personas, we generate counterfactual trajectories around demonstrations and prompt the LLM to compare its decisions with human choices, identify suboptimal actions, and extract lessons. The learned policy is then used to drive agent interactions with the recommender system. We evaluate AlignUSER across multiple datasets and demonstrate closer alignment with genuine humans than prior work, both at the micro and macro levels.
Problem

Research questions and friction points this paper is trying to address.

recommender system evaluation
human behavior alignment
synthetic users
interaction data scarcity
offline-online gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
LLM agents
counterfactual trajectories
human alignment
recommender system evaluation
🔎 Similar Papers
No similar papers found.
N
Nicolas Bougie
Woven by Toyota
G
Gian Maria Marconi
Woven by Toyota
T
Tony Yip
Woven by Toyota
Narimasa Watanabe
Narimasa Watanabe
Woven by Toyota