Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of capturing long-range dependencies via behavior cloning in non-Markovian environments, where recurrent models and attention mechanisms are constrained by state collapse and limited context lengths. To overcome these limitations, this work proposes Keyframe Mnemonics, a method that innovatively constructs a self-supervised objective through randomly sampled historical observations to distill keyframes as memory, subsequently training a conditioned policy network with theoretically guaranteed context retention over infinite horizons. Experimental results demonstrate that the proposed approach achieves a 100% success rate in synthetic domains while generalizing to extremely long horizons. Furthermore, on robotic manipulation benchmarks, it outperforms the strongest baseline by 13.9%, and maintains an 80% success rate on real-world robots even under twenty-fold longer horizons.
📝 Abstract
Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a $13.9$% average absolute SR improvement over the strongest baseline across $23$ tasks and retaining $80$% SR at $20\times$ longer horizons on a real robot. Code and videos are available at https://keyframe-mnemonics.github.io.
Problem

Research questions and friction points this paper is trying to address.

Behavior Cloning
Non-Markovian Environments
Long-horizon Dependencies
Hidden-state Collapse
Context Length Limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Behavior Cloning
Keyframe Discovery
Non-Markovian Environments
Horizon Generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Prabin Kumar Rath
Arizona State University
O
Omkar Patil
Arizona State University
Nakul Gopalan
Nakul Gopalan
Assistant Professor Arizona State University
RoboticsNatural LanguageReinforcement Learning