ECHO: Event-Augmented Context with Hindsight and Outlook for Wrist-Only Manipulation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of static information and out-of-view cues caused by the limited field of view in wrist-mounted event cameras, proposing the ECHO model. ECHO introduces the first latent world action model relying solely on a wrist-mounted camera. It constructs compact motion representations via a pretrained encoder and incorporates Hindsight and Outlook modules to provide spatiotemporal context. Combined with learnable queries, the model enables future event prediction and enhances policy reasoning capabilities. Validated on the RLBench simulation benchmark and real-world robots, ECHO surpasses RGB baselines by 20.6% and 14.6% in success rate under normal and extreme lighting conditions, respectively, demonstrating robust manipulation across high dynamic ranges.
📝 Abstract
Learning-based manipulation policies relying on RGB cameras often suffer from degraded observations under extreme exposure. Event cameras mitigate this degradation by asynchronously detecting pixel-level intensity changes to offer a high dynamic range. However, their observations heavily depend on camera placement, as fixed cameras miss static scene content while wrist-mounted camera motion causes previously visited regions to leave the field of view. To address these spatial-temporal limitations, we present ECHO (Event-augmented Context with Hindsight and Outlook), a wrist-only latent world action model that encodes wrist events into compact motion representations to provide temporal and spatial context for policy reasoning. Specifically, ECHO utilizes a pretrained event encoder to explain visual-feature changes between frames. Its hindsight module preserves the gripper trajectory with past event stream as addressable off-camera context. Concurrently, the outlook module introduces learnable event foresight queries supervised to anticipate the event window for future actions, enabling the policy to predict upcoming scene changes. Evaluated on wrist-only RLBench tasks, ECHO outperforms RGB and RGB+event baselines by 20.6 and 12.0 percentage points under normal lighting, and by 14.6 and 11.3 points under severe exposure drops, respectively, while also surpassing RGB references using a third-person camera. Real-world experiments with a wrist-mounted event camera validate that ECHO outperforms RGB-only and RGB+event baselines across multiple tasks under both nominal and severely dark lighting. Project page is at https://echo-wam.github.io/.
Problem

Research questions and friction points this paper is trying to address.

event camera
wrist-mounted manipulation
extreme exposure
spatio-temporal limitations
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Event Camera
Latent World Action Model
Wrist-Only Manipulation
Hindsight and Outlook Modules
High Dynamic Range
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xinyue Wang
The Hong Kong University of Science and Technology
Y
Yicheng Jiang
The Hong Kong University of Science and Technology
Z
Zesen Gan
The Hong Kong University of Science and Technology
J
Junhao He
The University of British Columbia
J
Jiaxu Wang
MMLab, The Chinese University of Hong Kong
Junhao Li
Junhao Li
Assistant Project Scientist, Cognitive Science, University of California, San Diego
Non-coding RNAsDNA methylationEpigeneticsBioinformatics
J
Jingtao Zhang
The Hong Kong University of Science and Technology
T
Tianlun He
MMLab, The Chinese University of Hong Kong
Jianan Wang
Jianan Wang
Astribot / IDEA / Deepmind / Oxford
Computer VisionGenerative AIRoboticsLearning Theory
I
Isabel Guan
The Hong Kong University of Science and Technology; ZENBOT
Qiming Shao
Qiming Shao
HKUST / UCLA / Tsinghua University
Topological spintronicsSpin-orbitronicsMagnetic insulatorsQuantum devicesEfficient Learning