DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of dense matrix multiplications in large model inference and the accuracy degradation caused by error amplification in existing weight approximation methods. To this end, it proposes DyRA, an input-adaptive dynamic residual approximation method that innovatively shifts the optimization objective from weight approximation to output activation optimization. By integrating low-rank factorization with an input-dependent dynamic error correction mechanism, DyRA achieves more accurate matrix multiplication approximations under equivalent computational budgets. Experimental results demonstrate that DyRA significantly improves the accuracy-efficiency trade-off across vision, speech, and language models. Notably, on DINOv3, it delivers a 1.5× GPU speedup while reducing accuracy loss by over threefold compared to baselines.
📝 Abstract
Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5$\times$ end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3$\times$ relative to weight-only baselines.
Problem

Research questions and friction points this paper is trying to address.

matrix multiplication
inference efficiency
structured approximation
output error
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Residual Approximation
Structured Matrix Multiplication
Low-rank Factorization
Input-adaptive Correction
Inference Efficiency
🔎 Similar Papers
No similar papers found.