design microbenchmarks

Design and implement small, focused benchmarking programs and test harnesses that isolate and exercise specific timing‑critical code paths to measure execution latency and overhead precisely. Build measurement procedures that control and vary environmental and system variables, replicate target device conditions, and analyze results to attribute observed overhead to hardware or software components.

designmicrobenchmarks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

$mu$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics

Jan 22, 2025
NJ
Nils Japke
🏛️ TU Berlin | Simula Research Laboratory

To address the low efficiency of performance monitoring in CI/CD pipelines caused by excessive repetition in microbenchmarking, this paper proposes a static, data-driven method for determining the minimal reliable repetition count per microbenchmark. Leveraging historical execution data and five statistical stability metrics—including coefficient of variation (CV), relative standard deviation (RSD), and interquartile range (IQR)—our approach statically infers the minimum repetitions required to achieve measurement reliability. It explicitly models JVM warm-up effects and supports cross-language adaptation for both Java and Go. Furthermore, the method is tightly integrated with industrial-grade performance regression detection frameworks. Experimental evaluation across 14 open-source projects demonstrates that, while preserving regression detection accuracy, our technique reduces microbenchmark measurement time by 95.83% (Go) and 94.17% (Java) on average.

Efficiency in TestingPerformance MonitoringSoftware Development

This work addresses the lack of reproducible and fairly comparable benchmarks in quantum software testing, which has largely relied on small, hard-coded circuits that poorly reflect real-world development practices. To bridge this gap, the authors introduce Qolumbina—the first scalable benchmark suite for quantum software testing—systematically curated from 40 representative programs sourced from open-source repositories and rigorously refactored with standardized interfaces, test cases, and formal specifications. The study further proposes a novel taxonomy of testing characteristics specific to quantum programs and employs program complexity modeling to enable scalability analysis and systematic evaluation. Covering a diverse range of testing attributes, Qolumbina has already facilitated empirical studies on execution overhead and fault detection capability, revealing the critical influence of backend dependencies on the interpretation of testing outcomes.

benchmarkempirical evaluationquantum software testing

This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.

application benchmarkscloud service benchmarkingmicrobenchmarks

Detecting Performance-Relevant Changes in Configurable Software Systems

Nov 21, 2025
SB
Sebastian Böhm
🏛️ Saarland University | Leipzig University

Detecting performance regressions in configurable software is costly, and configuration sampling often misses localized performance degradation. Method: This paper proposes ConfFLARE, a technique that combines data-flow dependency analysis with change-impact propagation tracking to identify code changes interacting—via data flow—with performance-sensitive code. It further integrates configuration-feature identification to automatically select the subset of performance-sensitive configurations most likely affected by each change. Contribution/Results: ConfFLARE eliminates the need for exhaustive configuration-based performance testing. In evaluations on synthetic and real-world systems, it reduces the number of required test configurations by 79% and 70%, respectively, while achieving near-complete coverage of performance regression cases. It precisely pinpoints relevant features and significantly improves both the efficiency and completeness of performance regression detection.

Detecting performance-relevant changes in configurable software systems efficientlyIdentifying performance regressions through data-flow interactions with critical codeReducing measurement costs by selecting relevant configurations for testing

Latest Papers

What's happening recently
View more

Existing benchmarks for execution-time optimization patches primarily target Python, C++, or .NET, lacking configurable, reproducible solutions tailored to Java. This work proposes JETO-Mine, the first framework for automatically mining and validating Java performance patches with customizable filtering and statistically rigorous validation. JETO-Mine employs a three-stage pipeline integrating static analysis, LLM-driven issue categorization, Docker-based dynamic testing, and significance testing to construct JETO-Bench—a benchmark comprising 660 candidate patches and 91 manually verified effective ones. Experimental evaluation demonstrates that JETO-Bench effectively assesses patch generation tools (e.g., OpenHands achieves a 14.3% repair success rate) and reveals a widespread absence of performance validation tests in Java projects.

benchmarkexecution time improvementJava

This study addresses a critical gap in quantum software research: the absence of a systematic auditing mechanism for empirically grounded comparative claims, which has led to a pervasive “instantiation gap” characterized by insufficient evidentiary support. To bridge this gap, the authors propose CLAIMSTAB-QC, the first source-bound auditing framework tailored to empirical comparisons in quantum software. By integrating claim modeling, audit scope delimitation, evidence boundary identification, and directional classification, the framework enables precise validation of comparative assertions against original source materials. An evaluation across 455 claims from 119 papers reveals that only eight claims possessed sufficient matched evidence for direct auditing; among these, two were confirmed, four lacked adequate support, and two were contradicted—highlighting substantial deficiencies in the empirical rigor of current quantum software studies.

benchmarkingempirical comparisonevidence auditing

This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.

compiler optimizationembedded softwareintegration testing

This work addresses the inefficiencies and semantic inconsistencies arising from separately implementing driver and monitor programs in traditional hardware module testing. To overcome this, the authors propose a domain-specific language (DSL) tailored to hardware communication protocols, which enables the unified specification of both driver and monitor logic through an imperative syntax, thereby ensuring their semantic consistency for the first time. Building upon this DSL, they develop a prototype tool that leverages waveform parsing and transaction-level trace inference techniques to accurately reconstruct protocol-compliant transaction sequences from raw signal waveforms. Experimental results demonstrate that the approach significantly improves development efficiency, with further validation planned on real-world interconnect protocols such as Wishbone and AXI-Stream.

driverhardware communicationmonitor

This work addresses the challenge of obtaining trustworthy worst-case execution time (WCET) bounds, as existing analysis tools are often closed-source or lack support for commercial microcontrollers. We present an open-source static WCET analyzer designed for commercial off-the-shelf (COTS) microcontrollers. Adopting a simplicity-first design philosophy, the tool directly parses architecture-specific binaries to estimate WCET, natively supports mainstream chips such as the ESP32-C6, and explicitly reports unsound results arising from hardware or software limitations to ensure transparency. Experimental evaluations successfully derive reliable WCET bounds for both the ESP32-C6 and MSP430 platforms. These results validate the tool’s usability and its advantage of eliminating licensing barriers, demonstrating its suitability for real-time systems education and engineering practice.

Embedded HardwareMicrocontrollersOpen-Source

Hot Scholars

SA

Shaizeen Aga

AMD Research
Near-data processingSecure hardwareParallel Computer Architecture
MS

Mohammad Sadrosadati

Senior Researcher and Lecturer, ETH Zürich
Heterogeneous ComputingProcessing-In-MemoryMemory SystemsInterconnection Networks
SC

Sunita Chandrasekaran

Associate Professor, Dept. of CIS, University of Delaware
High Performance ComputingParallel ProgrammingOpenMPOpenACC
GP

Gennady Pekhimenko

University of Toronto
Computer ArchitectureSystemsSystems for MLMachine Learning
CG

Christina Giannoula

Postdoctoral Researcher, University of Toronto
Computer ArchitectureComputer SystemsProcessing-In-MemoryMachine Learning