Trace-based, time-resolved analysis of MPI application performance using standard metrics

📅 2025-12-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing MPI performance analysis tools rely on time-aggregated metrics, which obscure transient bottlenecks. To address this, we propose a fine-grained, time-windowed trace analysis method that partitions execution traces into fixed or adaptive temporal windows and computes time-resolved metrics—including communication efficiency, load balance, and serialization overhead. Our approach integrates Paraver-based post-processing, critical path reconstruction, and event anomaly correction (e.g., clock skew compensation and unmatched MPI event reconciliation) to enable high-precision localization of transient bottlenecks. Evaluation on real-world applications (LaMEM, ls1-MarDyn) and synthetic benchmarks demonstrates that our method significantly improves both accuracy and scalability in identifying transient performance issues within large-scale traces, thereby overcoming the inherent limitations of global aggregation-based analysis.

Technology Category

Machine Learning: Evaluation and AnalysisPlanning, Routing, and Scheduling: Optimization of Spatio-temporal SystemsData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationSecurity and Privacy: Large-scale security measurementsWeb Mining and Content Analysis: Web measurements
📝 Abstract
Detailed trace analysis of MPI applications is essential for performance engineering, but growing trace sizes and complex communication behaviour often render comprehensive visual inspection impractical. This work presents a trace-based calculation of time-resolved values of standard MPI performance metrics, load balance, serialisation, and transfer efficiency, by discretising execution traces into fixed or adaptive time segments. The implementation processes Paraver traces postmortem, reconstructing critical execution paths and handling common event anomalies, such as clock inconsistencies and unmatched MPI events, to robustly calculate metrics for each segment. The calculated per-window metric values expose transient performance bottlenecks that the timeaggregated metrics from existing tools may conceal. Evaluations on a synthetic benchmark and real-world applications (LaMEM and ls1-MarDyn) demonstrate how time-resolved metrics reveal localised performance bottlenecks obscured by global aggregates, offering a lightweight and scalable alternative even when trace visualisation is impractical.
Problem

Research questions and friction points this paper is trying to address.

Analyzes MPI application performance using time-resolved standard metrics
Identifies transient bottlenecks hidden by time-aggregated metrics in tools
Processes execution traces robustly despite anomalies and large sizes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-resolved MPI metrics via trace discretization
Postmortem Paraver trace processing with anomaly handling
Reveals transient bottlenecks hidden by global aggregates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kingshuk Haldar
High Performance Computing Center Stuttgart (HLRS), University of Stuttgart