🤖 AI Summary
This study addresses the suboptimal policies and successor measure biases arising from least-squares inference of task vectors in behavioral foundation models. To overcome these limitations, we propose BLS, a method that formulates a convex optimization objective jointly minimizing reward reconstruction error while aligning successor measures. This approach enables efficient zero-shot policy inference at test time, accompanied by a theoretical analysis establishing an upper bound on the suboptimality gap. Extensive experiments demonstrate that BLS significantly outperforms existing baselines across locomotion, manipulation, and humanoid control benchmarks, with negligible computational overhead.
📝 Abstract
Behavioral Foundation Models (BFMs) aim to solve a wide range of downstream tasks without test-time policy learning by inferring a task vector from the reward function. While efficient, the retrieved policies are often suboptimal because of how this task vector is inferred, typically with ordinary least squares (OLS). OLS minimizes reward reconstruction error but leaves the ordering of rewards unconstrained, which can bias the successor measure of the retrieved zero-shot policy away from that of the optimal policy. In this work, we propose BLS, an efficient test-time inference method that balances minimizing reward reconstruction error with reducing successor-measure mismatch. Theoretically, we provide a suboptimality gap upper bound characterized by both successor-measure and reward-function residuals. Empirically, we evaluate BLS on top of state-of-the-art BFMs across benchmarks for locomotion, manipulation, and humanoid control. BLS outperforms existing task inference baselines with negligible computational overhead. Project page: https://embodiedai-ntu.github.io/BLS