Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of zero-shot transfer for deep policies in novel environments with significantly divergent observation spaces. Departing from conventional reliance on observational inputs, this work introduces a pioneering training algorithm for purely reward-action-conditioned policies. By conditioning solely on reward signals and actions, the proposed policy executes decision-making without accessing raw observations, leveraging source-domain experience to guide training in target environments. The effectiveness of this approach is validated across diverse simulated and real-world robotic scenarios, demonstrating successful zero-shot adaptation and performance improvements across heterogeneous observations, such as varying rendering styles. Ultimately, this research establishes a new paradigm for cross-domain transfer in reinforcement learning by eliminating the dependency on consistent observation representations.
📝 Abstract
This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.
Problem

Research questions and friction points this paper is trying to address.

zero-shot transfer
reward-based policy
observation space
policy adaptation
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward-Based Policy
Zero-Shot Transfer
Rapid Adaptation
Observation Space
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
Morgan Byrd
Morgan Byrd
PhD Student, Georgia Institute of Technology
M
Maks Sorokin
Georgia Institute of Technology, Atlanta, GA, 30308, USA
R
Robert Wright
Georgia Tech Research Institute, Atlanta, GA, 30308, USA
Sehoon Ha
Sehoon Ha
Georgia Institute of Technology
roboticscomputer graphicsmachine learning