Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation at skill transitions during long-horizon robotic manipulation, where chained skill execution induces observation space shifts. Identifying scene state displacement as the core failure source, we propose a diagnostic mechanism based on feature weight resetting and construct a closed-loop "detect-recover-resume" system. This framework integrates task progress monitoring, learning-based scene recovery, and transition-robust fine-tuning, establishing a new paradigm in which scene recovery outperforms blind retries. Evaluated on the BOSS-44 benchmark, our approach improves full-chain success rates from 7.6% to 26.5%, representing a 3.5-fold increase that achieves 51% of the privileged recovery oracle's performance while significantly surpassing baseline methods such as diffusion policies.
📝 Abstract
Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned $π_{0.5}$ policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon robotic manipulation
Observation-Space Shift
Skill composition
Skill seams
Innovation

Methods, ideas, or system contributions that make the work stand out.

Observation-Space Shift
Long-Horizon Robotic Manipulation
Skill Composition
Detect-Restore-Resume System
Scene Restoration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pranav Wagh
Department of Computer Science, University of North Carolina at Chapel Hill, NC, USA
Yu Fang
Yu Fang
Honda Research Institute Japan Co., Ltd.
Human-Robot InteractionEye-head coordinationEye MovementVisual Perception/Cognition
Y
Yue Yang
Department of Computer Science, University of North Carolina at Chapel Hill, NC, USA
Mingyu Ding
Mingyu Ding
Assistant Professor, UNC Chapel Hill
RoboticsEmbodied AIComputer Vision