Score
Designs, builds, and configures low‑overhead runtime profiling infrastructures and tools that collect lightweight execution statistics, traces, and instrumentation for embedded, real‑time, and system‑level environments. Uses these measurements to analyze algorithmic and system performance — including memory usage, runtime hotspots, cardinality and distribution estimates — to support selection, tuning, and deployment decisions.
Resource management across the IoT/Edge/Cloud continuum faces a fundamental trade-off among prediction accuracy, real-time responsiveness, and deployment lightweightness; existing approaches rely either on runtime sampling or static rules, failing to reconcile these requirements. This paper proposes a metadata-driven runtime workload profiling framework. First, it formally defines the mapping between static metadata and dynamic runtime behavior. Second, it introduces a clustering-based mechanism for extracting semantically significant metadata features—without requiring execution traces or online profiling. Third, it integrates lightweight feature engineering with regression modeling to achieve low-overhead, high-accuracy resource demand prediction. Evaluated across Alibaba’s ML workloads and Google’s cluster traces, the framework maintains high prediction accuracy even under partial data anonymization, substantially outperforming conventional static prediction and online profiling baselines. It achieves superior real-time performance, prediction fidelity, and practical deployability.
This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.
Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.
Modeling complex inter-parameter dependencies in configurable software systems remains challenging for performance tuning. Method: This paper pioneers a fitness landscape (FL) perspective to reconstruct performance analysis, modeling the high-dimensional configuration space as a structured terrain—departing from conventional isolated-point evaluation. It integrates graph-based data mining, fitness landscape analysis (FLA), and large-scale sampling (86 million configurations) across three real-world systems and 32 workload types. Contribution/Results: The study uncovers six universal landscape patterns, enabling robust identification of local optima and precise topological characterization of configuration terrain. These findings substantially deepen insights into black-box system performance, providing both a novel theoretical foundation and reusable practical guidelines for automated configuration tuning and performance modeling.
Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.
This work addresses the limitations of existing CPU benchmarks in accurately evaluating the performance of modern heterogeneous, multithreaded processors under diverse workloads. To this end, the authors present the SPEC CPU 2026 benchmark suite, developed through community collaboration and principled methodology, which introduces the Rolling-Round-Robin Rate approach to standardize the execution of heterogeneous multiprogrammed workloads. The suite incorporates newly designed multithreaded benchmarks exhibiting varied microarchitectural characteristics, selected and hardened through an open-source application curation process. Emphasizing workload diversity, portability, and long-term viability, SPEC CPU 2026 establishes a robust, representative, and authoritative standard for performance evaluation, thereby supporting next-generation computer architecture research.
This study addresses the misleading nature of Java microbenchmarks conducted in isolation, which often yield distorted performance profiles due to the JVM’s dynamic compilation mechanisms—particularly inaccurate runtime profiling data such as branch probabilities and call-site types. The paper presents the first systematic investigation into profiling biases introduced by the absence of realistic contextual information in microbenchmarks, demonstrating that such distortions persist even when benchmarks strictly adhere to JMH best practices. By integrating JMH with JVM dynamic compilation internals and runtime profiling techniques, the authors empirically analyze representative cases of benchmark misinterpretation and propose an enhanced set of practical guidelines. These recommendations substantially improve the representativeness and reliability of microbenchmark results with respect to real-world application performance.
Existing approaches struggle to accurately capture the fine-grained temporal behavior of tasks under varying resource contexts, limiting resource utilization efficiency in soft real-time systems. This work proposes a generative analytical method based on nonparametric conditional multi-marginal Schrödinger bridges (MSB), introducing this framework for the first time to real-time system modeling. The method synthesizes high-fidelity task execution traces under unobserved resource configurations while providing maximum likelihood guarantees. By incorporating hardware resource context for adaptive inference, it significantly improves resource utilization efficiency in multicore real-time systems, as demonstrated through extensive evaluations on real-world benchmarks, thereby validating its effectiveness and practicality.
This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.