Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings
This study addresses the challenge of learning control policies in the absence of action records by proposing a hierarchical model-based reinforcement learning framework. Rather than relying on reconstruction dynamics errors, the method extracts low-dimensional dynamics embeddings via piecewise linear recurrent neural networks and explicit intervention models, which are subsequently reused to parameterize a shared policy network. Experimental evaluations demonstrate that this framework significantly improves transfer performance and cumulative rewards on benchmark systems such as Lorenz attractors, achieving cross-system generalization and interpretable intervention effects. Furthermore, the proposed approach is successfully applied to the suppression of neurally predicted movements, highlighting its practical utility in real-world control scenarios.