PMTRM: Pseudo-Memory Temporal Re-encoding Module for Embodied Policy Learning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the phase ambiguity problem in robotic manipulation, where locally similar observations cause policies to repeat actions or switch erroneously. To mitigate this, we propose PMTRM, a lightweight plug-in module that introduces a pseudo-memory temporal re-encoding mechanism and a temporal heterogeneity loss. It encodes execution history into latent sequences to augment existing policies while preserving their original action heads. The representation is optimized through progressive sim-to-real joint training, temporal masking, and an anchor reconstruction loss, with reconstruction performed exclusively during training to retain critical information. With minimal additional parameters, PMTRM significantly enhances historical awareness across diverse backbone networks. Both simulation and real-world experiments demonstrate substantial improvements in success rates for tasks involving phase ambiguity, achieved with negligible computational overhead.
📝 Abstract
Robotic manipulation often contains repeated motions whose local observations look similar at different phases. When these phases require different actions, a policy that relies mainly on the current observation may repeat completed motions or switch phases at the wrong time. To address this phase ambiguity, we present the Pseudo-Memory Temporal Re-encoding Module (PMTRM), a lightweight plug-in module with only 7.61M parameters that encodes a bounded history of executed states and actions into a latent sequence for existing policies. To help distinguish phases, a temporal heterogeneity objective penalizes positive similarity between distant positions in this sequence, while anchor and reconstruction losses preserve information needed for action prediction. The reconstruction decoder is used only during training, leaving the temporal re-encoder to supply history to the policy at inference. We train the module progressively on synthetic sequences and robot data, then jointly with the policy, using temporal masking to accommodate partial histories. This integration retains the original action head and action space and adds auxiliary losses to the original policy loss. Experiments with multiple policy backbones in simulation and on a real robot show improved task success on tasks with phase ambiguity, with little additional computation.
Problem

Research questions and friction points this paper is trying to address.

phase ambiguity
robotic manipulation
embodied policy learning
repeated motions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pseudo-Memory Temporal Re-encoding
Phase Ambiguity
Temporal Heterogeneity Objective
Plug-in Module
Embodied Policy Learning
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
💼 Related Jobs
No related jobs found.
C
Changchuan Yang
Zhejiang University
Haoxuan Xu
Haoxuan Xu
Beihang University
computer vision
W
Wenbo Chen
The Hong Kong University of Science and Technology (Guangzhou)
S
Shuai Ren
vivo Robotics Lab
J
Jianlong Zheng
South China Normal University
H
Huarui Zhang
The Hong Kong University of Science and Technology (Guangzhou)
T
Tianfu Li
The Hong Kong University of Science and Technology (Guangzhou)
Guanzhong Tian
Guanzhong Tian
Ningbo Research Institute, Zhejiang University
Computer VisionModel CompressionPattern Recognition