PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In visual reinforcement learning, existing methods struggle to accurately capture behavioral similarity induced by rewards and transition dynamics due to their reliance on fixed or unconstrained latent distances, thereby limiting representation learning efficacy. To address this, this work proposes the Pairwise Adaptive Mahalanobis Distance (PAMD)—a positive-definite, pairwise-conditioned metric structure that can be seamlessly integrated as a plug-and-play module into existing multi-step bisimulation algorithms. By parameterizing a Mahalanobis distance, PAMD introduces structural constraints that preserve expressiveness while overcoming the limitations of fixed norms and the degenerate solutions often arising from unconstrained metrics. Empirical results demonstrate that PAMD significantly enhances the final performance of multiple bisimulation-based RL algorithms on visual MuJoCo continuous control tasks.
📝 Abstract
Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by reward and transition similarity. In practice, the choice of the latent distance can strongly affect performance: using a fixed, pre-specified global norms (e.g., $\ell_p$ norms or other hand-designed metrics) may be overly restrictive to capture the behavioral distance. In contrast, unconstrained pairwise distances may admit degenerate solutions that drive the metric loss down without improving the representation. To address this gap, we introduce **PAMD: Pairwise Adaptive Mahalanobis Distance**, which parameterizes a positive-definite, pair-conditioned metric for measuring latent state similarity. PAMD is a simple plug-in for existing bisimulation-based methods, offering a more expressive yet structured alternative to fixed, pre-specified latent distances. We empirically validate our method on visual MuJoCo continuous-control tasks, where final performance of several recent bisimulation-based RL algorithms is substantially improved when equipped with the distance we propose.
Problem

Research questions and friction points this paper is trying to address.

visual reinforcement learning
bisimulation
latent distance
behavioral distance
representation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

PAMD
adaptive metric
bisimulation
visual reinforcement learning
Mahalanobis distance