Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

πŸ“… 2026-03-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work proposes a Dynamic Preference Inference (DPI) framework to address the challenge of unobservable and dynamically shifting preference weights in multi-objective reinforcement learning. DPI introduces, for the first time, a cognitively inspired approach to dynamic preference modeling by maintaining a Bayesian belief over preference weights through a variational preference inference module, which is updated online using vector-valued returns as evidence of trade-offs. The framework integrates this inference mechanism with a preference-conditioned Actor-Critic architecture, enabling joint policy training. Evaluated on tasks including queue scheduling, maze navigation, and continuous control under event-driven objective switches, DPI demonstrates significantly superior performance compared to fixed-weight and heuristic envelope baselines, highlighting its enhanced capability for adaptive decision-making in non-stationary preference environments.

Technology Category

Knowledge Representation and Reasoning: PreferencesMachine Learning: Learning Preferences or RankingsHumans and AI: Learning Human Values and Preferences

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingEconomics, Online Markets and Human Computation: Fairness, privacy, and diversity in economic environmentsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
πŸ“ Abstract
Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.
Problem

Research questions and friction points this paper is trying to address.

dynamic preference
contextual shifts
sequential decision-making
latent preferences
multi-objective reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Preference Inference
Multi-objective Reinforcement Learning
Contextual Shifts
Latent Preference Modeling
Preference-conditioned Policy
πŸ”Ž Similar Papers
No similar papers found.