PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether pretrained Transformers can function as closed-loop controllers to implement Policy Mirror Descent (PMD). To this end, the project establishes a fixed protocol that leverages causal softmax attention, negative-entropy PMD, and LayerNorm architectures to recover target computations, while evaluating closed-loop control performance through auditing models. This work provides the first empirical evidence that pretrained Transformers can approximate PMD updates, thereby establishing a novel paradigm for closed-loop control. Experimental results demonstrate that the actor loss is merely 1.05 times that of exact PMD, significantly outperforming existing algorithm distillation approaches.
📝 Abstract
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_η(π,Q)=\operatorname{softmax}(\logπ+ηQ)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transformers recover the target computation empirically. A frozen one-step audit model is closest to PMD among the tested fixed rules; in a preregistered five-run $S=4$ repeated-control test, the learned actor with an exact one-step critic reaches median returned-policy loss $1.052\times$ the Exact PMD oracle and retains the criterion across four no-retraining shifts. The same checkpoints with their learned critic give descriptive median $1.050\times$ the oracle (no registered margin). At $S=8$, replacing the exact critic by the learned critic raises median $T=20$ loss to $0.0225$ yet leaves the Liang--Lai and Algorithm Distillation adaptations $20.2$--$24.2\times$ higher-loss; this is a one-sided sampled-critic bound because PolicyAttention consumes 144 generative transitions per round versus 20 on-policy transitions for the adaptations. The strict 20-transition comparison remains open. At $S=8,16$, the exact-critic common-harness comparison remains $17.7$--$28.2\times$ lower-loss than those adaptations, with the information asymmetry stated locally.
Problem

Research questions and friction points this paper is trying to address.

Softmax Attention
Policy Mirror Descent
Closed-Loop Control
Transformer
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Mirror Descent
Causal Softmax Attention
Closed-Loop Control
Transformer
One-Step Critic
🔎 Similar Papers
Y
Yuhe Sui
Quantitative Research Society, Singapore
Yingzhi Tang
Yingzhi Tang
City University of Hong Kong
computer visiondeep learningperson reid3d point cloud
S
Shufang Chen
The University of Hong Kong, Hong Kong SAR, China