Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads

📅 2026-09-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决HPC和AI负载中通信效率问题,提出CC-Bench基准套件,通过压缩通信并评估不同库、数据集和硬件上的性能与准确性。
📝 Abstract
Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. Communication compression can reduce this overhead and complement execution-level optimizations, but its benefits remain difficult to assess because existing benchmarks lack support for diverse backends, realistic datasets, application-specific accuracy metrics, and overlap-induced resource contention. We present CC-Bench, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions. CC-Bench uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation. It further combines function-level interception and hardware counter monitoring to characterize per-phase latency, hardware utilization, numerical fidelity, and computation interference. With representative datasets from HPC and LLM workloads, CC-Bench evaluates three compression-enabled communication libraries on CPU and GPU clusters, revealing accuracy-performance trade-offs and bottlenecks to guide practical deployment and optimization.
Problem

Research questions and friction points this paper is trying to address.

Distributed HPC
Communication Compression
Performance Bottleneck
Benchmark
Resource Contention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Communication Compression
Benchmark Suite
Application-Environment Modeling
Cross-Library Evaluation
Hardware Counter Monitoring
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haozhe Fan
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
W
Wei Wang
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
X
Xingchen Liu
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
M
Man Liu
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
X
Xingjian Tian
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
H
Haoquan Long
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Z
Zedong Liu
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
D
Daran Sun
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
J
Jinwu Yang
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
B
Bo Yang
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
J
Jie Liu
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Y
Yonggang Che
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
H
Hairui Zhao
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
G
Guangming Tan
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Dingwen Tao
Dingwen Tao
Chinese Academy of Sciences, IEEE/ACM Senior Member
High Performance ComputingData ReductionDeep LearningSystems for MLGPU