๐ค AI Summary
This study addresses the substantial parameter count and high KV cache memory overhead inherent in Transformer architectures by proposing a dual-head attention Transformer. The core innovation lies in replacing conventional value projection matrices with structured orthogonal transforms, pioneering the integration of spatial patch mechanisms with Walsh-Hadamard, DCT, Fourier, and Shearlet transforms to construct a streamlined dual-head attention module. This approach significantly reduces model parameter scale and GPU memory consumption while achieving superior performance compared to standard three-head Transformers on ImageNet classification tasks. These results thoroughly validate the proposed architectureโs efficiency and generalization potential in resource-constrained scenarios.
๐ Abstract
To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.