EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration

📅 2025-06-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Collective communication operations—particularly AllReduce—impose significant overhead in large language model (LLM) distributed training, while quantization often degrades numerical stability. Method: This paper proposes the first TPU-optimized, dynamically blocked quantized AllReduce scheme integrated within the XLA compiler. It synergistically combines int8 quantization, compiler-level optimizations, TPU-friendly blocking strategies, and deep computation-communication overlap to jointly improve efficiency and maintain accuracy. Contribution/Results: We present the first dynamic blocking quantized AllReduce implementation for TPUs, achieving substantial communication reduction without compromising numerical stability. Experiments show an 1.8× throughput improvement in AllReduce over BF16 baselines; prefill latency is reduced by 1.25× for Gemma-3 27B and 1.1× for Gemma-3 12B, with negligible quality degradation.

Technology Category

Machine Learning: Hardware-aware MLNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Distributed Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
While Large Language Models (LLMs) have become highly influential, their enormous scale presents significant deployment challenges. Efficiently serving these models typically requires distributing them across numerous accelerator devices, which introduces substantial performance overhead from inter-device communication (collectives). While model quantization has been widely adopted to reduce the memory and compute requirements of LLM weights and activations with minimal quality impact, applying quantization directly to collectives like AllReduce is inherently difficult due to the inter-device summation involved, which can lead to numerical instability or significant error accumulation. In this work, we present a native dynamic block-wise efficient quantized AllReduce within the XLA compiler for TPUs (EQuARX). By using TPU-friendly quantization and deep pipelining of communication and compute, EQuARX with int8 precision achieves a 1.8X speedup over baseline BF16 AllReduce across various network topologies. Furthermore, EQuARX accelerates the prefill stage of Gemma 3 27B by 1.25X and Gemma 3 12B by 1.1X, respectively, with small to negligible impact on quality.
Problem

Research questions and friction points this paper is trying to address.

Reducing communication overhead in distributed LLM deployment
Overcoming quantization challenges in AllReduce operations
Improving efficiency of inter-device communication for TPUs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native dynamic block-wise quantized AllReduce
TPU-friendly quantization for stability
Deep pipelining of communication and compute
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Ibrahim Ahmed
Google
C
Clemens Schaefer
Google
Gil Tabak
Gil Tabak
PhD Student in Applied Physics, Stanford University
D
Denis Vnukov
Google
Zenong Zhang
Zenong Zhang
University of Texas at Dallas
Fuzz TestingSoftware EngineeringSecurity
F
Felix chern
Google
A
Anatoliy Yevtushenko
Google
A
Andy Davis
Google