Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the substantial parameter count and high KV cache memory overhead inherent in Transformer architectures by proposing a dual-head attention Transformer. The core innovation lies in replacing conventional value projection matrices with structured orthogonal transforms, pioneering the integration of spatial patch mechanisms with Walsh-Hadamard, DCT, Fourier, and Shearlet transforms to construct a streamlined dual-head attention module. This approach significantly reduces model parameter scale and GPU memory consumption while achieving superior performance compared to standard three-head Transformers on ImageNet classification tasks. These results thoroughly validate the proposed architectureโ€™s efficiency and generalization potential in resource-constrained scenarios.
๐Ÿ“ Abstract
To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.
Problem

Research questions and friction points this paper is trying to address.

Transformers
Parameter reduction
Cache memory
Attention mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-headed Transformer
Structured Transforms
Walsh-Hadamard Transform
Parameter Reduction
Attention Mechanism
๐Ÿ”Ž Similar Papers
No similar papers found.