HPC-MQBench: Qualification-First Benchmarking on Slurm with a Single-Broker Kafka Evaluation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.
📝 Abstract
Messaging experiments on high-performance computing clusters must coordinate services, clients, and measurement within scheduler allocations. We present HPC-MQBench, a Slurm-orchestrated benchmark that links resumable experiment control to record checks, delivery accounting, and qualification before rate ranking. Its Kafka evaluation placed producers/controller, one broker, consumers, and monitoring on four nodes. With memory-backed logs, 120 workload configurations yielded 99 qualified observations, 14 outside the producer-delivery policy, and seven with invalid evidence. Two validation stages each repeated ten workloads in five blocks. In the final stage, the selected workload qualified in all five observations, with medians of 2,927 mebibytes per second balanced endpoint rate, 1.24 percent pending deliveries, and 2.56 seconds for the 99th-percentile latency. It qualified in only two observations in the earlier stage, where qualification improved after an allocation boundary without establishing a causal allocation effect. Selection is conditional on the stage and delivery policy. Resource measurements did not isolate a unique bottleneck. The contribution is a framework for distributed experiment control and auditable configuration selection at a fixed broker count; multi-broker support and scaling experiments remain future work.
Problem

Research questions and friction points this paper is trying to address.

High-Performance Computing
Message Queue Benchmarking
Slurm Orchestration
Kafka Evaluation
Experiment Qualification
Innovation

Methods, ideas, or system contributions that make the work stand out.

HPC-MQBench
Qualification-First Benchmarking
Slurm Orchestration
Kafka Evaluation
Auditable Configuration Selection
🔎 Similar Papers
No similar papers found.
S
Sepehr Mahmoodian
University of Göttingen, Göttingen, Germany; Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG), Göttingen, Germany
J
Julian Kunkel
University of Göttingen, Göttingen, Germany; Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG), Göttingen, Germany