When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear practical implications of policy memory backpropagation truncation, as employed in Transformer-XL, under the interplay between physical credit assignment and optimizers. By fixing the forward computation while exclusively altering gradient edges, this work proposes evaluating performance based on the optimizer’s actual parameter updates rather than raw gradients. Comparative experiments are conducted using the AdamW optimizer on ship trajectory prediction and quadrotor tracking tasks. The findings reveal the underlying mechanisms governing memory truncation costs under aggressive clipping, demonstrating that truncation substantially increases tracking errors. Furthermore, this work uncovers that optimizer dynamics can significantly amplify the impact of minute gradient discrepancies on the final parameter updates.
πŸ“ Abstract
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forward computation fixed and varying only derivative edges, we measure parameter gradients, the updates the optimizer applies, and continued training in a Transformer vessel-trajectory model and a quadrotor tracking policy. In the vessel model, detaching the key-value cache shrank the gradient to about a tenth of its norm, with little rotation, when gradients flowed through all earlier physical states, but barely changed it under one-step physical credit. In this strongly clipped regime the optimizer, not the gradient, set how far updates differed: global-norm clipping removed most of the gradient difference between memory-cut graphs, whereas AdamW turned a 2% gradient difference between two placements of the cut into update differences of up to 31% at the step where the placement was switched. In a quadrotor trained from initialization with 0.20 m/s velocity noise, removing memory raised tracking error by 43% and cutting memory gradients raised it by 32%; at low noise the cut's mean cost exceeded the value of memory. Two-step truncation segments gave no measurable gain, although with hidden velocity a two-step window captured most of the value of memory; eight-step segments removed half to three quarters of the cost. Switching the cut on only for the last fifth of training understated its cost about threefold at 0.20-0.30 m/s, but not at low noise or with hidden velocity. These results suggest measuring the cost of a memory cut by training with it from initialization, and comparing backward graphs by the updates the optimizer applies rather than by raw gradients.
Problem

Research questions and friction points this paper is trying to address.

policy memory
truncated backpropagation through time
physical credit assignment
optimizer updates
partial observability
Innovation

Methods, ideas, or system contributions that make the work stand out.

truncated backpropagation through time
policy memory
optimizer updates
physical credit
computational graph
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xingjian Li
School of Data Science, The Chinese University of Hong Kong, Shenzhen
Y
Yi Han
COSCO Shipping Technology Co., Ltd.
Jianhua Z. Huang
Jianhua Z. Huang
School of Data Science, Chinese University of Hong Kong, Shenzhen
Applied StatisticsFunctional Data AnalysisMultivariate AnalysisStatistical Machine Learning and Data Mining