Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead of matrix multiplications in Transformers by proposing an alternative based on combining algebras. Methodologically, it leverages bilinear rank optimization and rectangular projections to construct provably optimal sparse interaction tables, achieving quadratic arithmetic complexity under fixed weight block sizes while accommodating GPU execution constraints, causal masking, and KV-cache decoding. Experiments on a 110M-parameter model demonstrate that, without altering the parameter count, the proposed approach increases generation throughput by 6.2%–7.8%. Although downstream evaluation metrics exhibit a slight degradation, these results validate the feasibility of efficient inference at small scales.
📝 Abstract
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.
Problem

Research questions and friction points this paper is trying to address.

Transformer
matrix multiplication
associative algebra
computational efficiency
throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

Associative Algebra Layers
Fast Matrix Multiplication
Transformer
Bilinear Rank
Throughput Optimization
🔎 Similar Papers