🤖 AI Summary
This study addresses the loss of positional information during token merging caused by the singular coordinates of rotary position embeddings in token compression. To overcome this, we propose Aperture, a framework that represents the positional state of merged tokens via weighted Fourier moments. We demonstrate that this state exhibits minimal real dimensionality in continuous space, enabling additive updates and decoupling attention normalization. Furthermore, unified interval-center rotation and sinc gain techniques are introduced. Numerical experiments validate the effectiveness of our approach. On video question answering tasks, Aperture achieves 65.63% accuracy, closely approaching the 67.12% obtained with explicit merging rules, thereby confirming the value of precise positional preservation for downstream applications.
📝 Abstract
Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token's weighted support at the model's rotary frequencies. We prove that these moments have minimal real dimension among continuous states sufficient for the selected expected rotary interactions. Represented mass makes updates additive; attention normalisation remains a separate readout choice. Uniform intervals give a centre rotation times a sinc gain. We characterise when centres determine interval widths and construct matched examples where they do not. Numerical checks verify the weighted-support implementation. In trained temporal readers, compression transfer varies with gain calibration and feature placement. In a prespecified native video question-answering comparison, stored support reaches $65.63\%$ accuracy versus $67.12\%$ for the deployed merging rule. These results separate exact positional preservation under compression from downstream benefit.