On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究发现后见标签在多目标强化学习中会导致偏好覆盖崩溃,提出her_mix方法以恢复受影响的设置并减少偏好丢失。
📝 Abstract
Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned multi-objective RL (MORL) relabels transitions with the preference direction the agent achieved rather than the one asked for. We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degrades 19 of 36 algorithm-environment settings by as much as four standard deviations, improves only one, and leaves the rest unaffected. The harm is not a symptom of noisy relabels; denoising the target recovers almost nothing, and neither prioritized sampling nor any buffer-structural choice reproduces it. Instead, repeated relabeling collapses the critic's coverage onto whatever narrow region of the preference space the agent happened to visit. We name this failure mode \emph{Preference Coverage Collapse}, and quantify it with abandoned preference mass (APM), a value-aware statistic that tracks the harm ($ρ= -0.73$) where a purely structural coverage count does not. We then introduce \texttt{her\_mix}, a single-parameter convex combination pulling the achieved direction back towards the requested preference. At one fixed value across every algorithm and environment, it returns 16 of the 19 harmed settings to baseline, preserves and even improves the one setting in which relabeling helps, and cuts abandoned preference mass from $69\%$ to $6\%$. Protecting coverage over the preference simplex, not filtering noisy relabels, is what makes hindsight relabeling safe for MORL.
Problem

Research questions and friction points this paper is trying to address.

Hindsight Relabeling
Preference-Conditioned MORL
Coverage Collapse
Sample-Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

her_mix
Preference Coverage Collapse
multi-objective reinforcement learning
hindsight relabeling
💼 Related Jobs
No related jobs found.
B
Baptiste Bonin
Mila (Quebec AI Institute), Université Laval
C
Caro Strickland
Mila (Quebec AI Institute), Université Laval
Audrey Durand
Audrey Durand
Assistant Professor, Université Laval, Canada
banditsreinforcement learninghealth informatics