🤖 AI Summary
This study addresses the challenge of maintaining effective control in robotic air hockey defense when targets are temporarily occluded, specifically investigating how memory mechanisms can compensate for missing visual input. The proposed approach leverages a DreamerV3 world model combined with knowledge distillation to compress a complex nonlinear policy into a diagonal linear recurrent model. This work demonstrates that a simple linear memory mechanism, when paired with a nonlinear encoder, can effectively substitute for intricate nonlinear recurrent dynamics. The resulting linear model achieves performance comparable to both gated recurrent units (GRUs) and the teacher model while requiring significantly fewer parameters and lower computational overhead. These findings provide empirical evidence that nonlinear dynamics are not strictly necessary for memory-dependent control tasks, offering a more efficient architectural paradigm for partially observable robotic environments.
📝 Abstract
Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.