π€ AI Summary
This study addresses the challenge of learning control policies in the absence of action records by proposing a hierarchical model-based reinforcement learning framework. Rather than relying on reconstruction dynamics errors, the method extracts low-dimensional dynamics embeddings via piecewise linear recurrent neural networks and explicit intervention models, which are subsequently reused to parameterize a shared policy network. Experimental evaluations demonstrate that this framework significantly improves transfer performance and cumulative rewards on benchmark systems such as Lorenz attractors, achieving cross-system generalization and interpretable intervention effects. Furthermore, the proposed approach is successfully applied to the suppression of neurally predicted movements, highlighting its practical utility in real-world control scenarios.
π Abstract
Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.