lightweight runtime profiling

Designs, builds, and configures low‑overhead runtime profiling infrastructures and tools that collect lightweight execution statistics, traces, and instrumentation for embedded, real‑time, and system‑level environments. Uses these measurements to analyze algorithmic and system performance — including memory usage, runtime hotspots, cardinality and distribution estimates — to support selection, tuning, and deployment decisions.

lightweightruntimeprofiling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Formal and Empirical Study of Metadata-Based Profiling for Resource Management in the Computing Continuum

Apr 29, 2025
AM
Andrea Morichetta
🏛️ Vienna University of Technology | Universitat Pompeu Fabra

Resource management across the IoT/Edge/Cloud continuum faces a fundamental trade-off among prediction accuracy, real-time responsiveness, and deployment lightweightness; existing approaches rely either on runtime sampling or static rules, failing to reconcile these requirements. This paper proposes a metadata-driven runtime workload profiling framework. First, it formally defines the mapping between static metadata and dynamic runtime behavior. Second, it introduces a clustering-based mechanism for extracting semantically significant metadata features—without requiring execution traces or online profiling. Third, it integrates lightweight feature engineering with regression modeling to achieve low-overhead, high-accuracy resource demand prediction. Evaluated across Alibaba’s ML workloads and Google’s cluster traces, the framework maintains high prediction accuracy even under partial data anonymization, substantially outperforming conventional static prediction and online profiling baselines. It achieves superior real-time performance, prediction fidelity, and practical deployability.

Combining past execution traces with metadata for fast workload associationOptimizing resource distribution in IoT-Edge-Cloud computing continuumProfiling workload using static metadata for resource allocation

From Profiling to Optimization: Unveiling the Profile Guided Optimization

Jul 22, 2025
BL
Bingxin Liu
🏛️ Beijing Normal University | Phytium Technology Co., Ltd.

This work addresses key challenges in Profile-Guided Optimization (PGO): high sampling overhead, poor adaptability to dynamic inputs, and weak cross-architecture portability. We systematically survey and restructure the PGO technical landscape, proposing the first multi-dimensional classification framework for PGO—explicitly identifying three core research directions: low-overhead profiling, dynamic workload adaptation, and cross-architecture profile migration. Our approach unifies instrumentation- and sampling-based analysis, enabling compiler- and linker-time collaborative optimization in GCC and LLVM across heterogeneous targets including x86 and ARM. Empirical evaluation on standard benchmarks demonstrates an average performance improvement of 12.3%, while reducing profiling overhead to under 0.8%. These results significantly enhance the industrial deployability and generalization capability of PGO.

Addressing challenges like sampling overhead and cross-architecture portability.Enhancing performance through Profile Guided Optimization (PGO).Systematically categorizing PGO research by profiling methods and optimizations.

gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks

Jul 08, 2025
YH
Yigong Hu
🏛️ Boston University | University of Washington | University of California, Los Angeles

Modern software systems frequently exhibit application-level resource contention bottlenecks—such as blocking on custom application events—that evade detection by conventional performance profilers due to complex dependencies and bespoke resource management. To address this, we propose OmniResource Profiling, the first method to jointly leverage system-level metrics and application-level event-waiting relationships. It employs a lightweight LLM-assisted static analysis to automatically identify custom resources and cross-execution-trace runtime variable comparison for precise root-cause localization. Evaluated on 12 known performance issues across five real-world applications, OmniResource achieves 100% diagnostic accuracy and uncovers two previously undetected bottlenecks. Crucially, it requires no intrusive instrumentation, balancing high precision with practical deployability. This work delivers the first end-to-end solution for application-level resource contention analysis.

Diagnosing performance bottlenecks in complex applicationsIdentifying application-level resource contention issuesUncovering root causes of performance problems

Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective

Dec 22, 2024
MH
Mingyu Huang
🏛️ University of Electronic Science and Technology of China | University of Exeter

Modeling complex inter-parameter dependencies in configurable software systems remains challenging for performance tuning. Method: This paper pioneers a fitness landscape (FL) perspective to reconstruct performance analysis, modeling the high-dimensional configuration space as a structured terrain—departing from conventional isolated-point evaluation. It integrates graph-based data mining, fitness landscape analysis (FLA), and large-scale sampling (86 million configurations) across three real-world systems and 32 workload types. Contribution/Results: The study uncovers six universal landscape patterns, enabling robust identification of local optima and precise topological characterization of configuration terrain. These findings substantially deepen insights into black-box system performance, providing both a novel theoretical foundation and reusable practical guidelines for automated configuration tuning and performance modeling.

Complex InterrelationshipsOptimizationSoftware Performance Tuning

Latest Papers

What's happening recently
View more

Traditional approaches struggle to provide deep visibility into the internal behavior of the gem5 simulator. This work proposes a non-intrusive, lightweight runtime call-stack analysis framework that, for the first time, treats the simulator’s own execution path as a novel lens for understanding simulated system behavior. Built upon the Linux perf_event interface, the framework enables parallel sampling, real-time symbol resolution, and hierarchical call-tree aggregation, with support for component-level customizable analysis. Experimental results demonstrate its effectiveness in uncovering performance bottlenecks in TimingSimpleCPU and identifying deadlock and livelock issues within the Ruby memory system—capturing critical behavioral characteristics that conventional statistical methods fail to detect.

cache coherencecall-stack profilinggem5

This work addresses the limitations of existing CPU benchmarks in accurately evaluating the performance of modern heterogeneous, multithreaded processors under diverse workloads. To this end, the authors present the SPEC CPU 2026 benchmark suite, developed through community collaboration and principled methodology, which introduces the Rolling-Round-Robin Rate approach to standardize the execution of heterogeneous multiprogrammed workloads. The suite incorporates newly designed multithreaded benchmarks exhibiting varied microarchitectural characteristics, selected and hardened through an open-source application curation process. Emphasizing workload diversity, portability, and long-term viability, SPEC CPU 2026 establishes a robust, representative, and authoritative standard for performance evaluation, thereby supporting next-generation computer architecture research.

benchmark longevityCPU benchmarkingheterogeneous workloads

This study addresses the misleading nature of Java microbenchmarks conducted in isolation, which often yield distorted performance profiles due to the JVM’s dynamic compilation mechanisms—particularly inaccurate runtime profiling data such as branch probabilities and call-site types. The paper presents the first systematic investigation into profiling biases introduced by the absence of realistic contextual information in microbenchmarks, demonstrating that such distortions persist even when benchmarks strictly adhere to JMH best practices. By integrating JMH with JVM dynamic compilation internals and runtime profiling techniques, the authors empirically analyze representative cases of benchmark misinterpretation and propose an enhanced set of practical guidelines. These recommendations substantially improve the representativeness and reliability of microbenchmark results with respect to real-world application performance.

dynamic compilationJava Virtual Machinemicrobenchmarks

Existing approaches struggle to accurately capture the fine-grained temporal behavior of tasks under varying resource contexts, limiting resource utilization efficiency in soft real-time systems. This work proposes a generative analytical method based on nonparametric conditional multi-marginal Schrödinger bridges (MSB), introducing this framework for the first time to real-time system modeling. The method synthesizes high-fidelity task execution traces under unobserved resource configurations while providing maximum likelihood guarantees. By incorporating hardware resource context for adaptive inference, it significantly improves resource utilization efficiency in multicore real-time systems, as demonstrated through extensive evaluations on real-world benchmarks, thereby validating its effectiveness and practicality.

context-dependent profilingexecution time analysisreal-time systems

This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.

Embedded SystemsEvent CorrelationHardware Performance Counters

Hot Scholars

ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor
ZW

Zhibin Wang

Zhejiang University
new particle formationaerosolshygroscopicityblack carbon
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
GL

Ge Li

Full Professor of Computer Science, Peking University
Program AnalysisProgram GenerationDeep Learning
KT

Kewei Tu

School of Information Science and Technology, ShanghaiTech University, China
Natural Language ProcessingMachine Learning