evaluate middleware implementations

Designs and runs empirical studies and benchmark suites to compare middleware implementations, building measurement rigs and analysis pipelines that quantify discovery, data exchange, latency, throughput, reliability, and resource usage. Analyzes collected metrics to classify implementations, identify performance bottlenecks and trade-offs, and evaluate behavior across different network and deployment conditions.

evaluatemiddlewareimplementations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications

Apr 16, 2025
SH
Soren Henning
🏛️ Dynatrace Research | Johannes Kepler University Linz

This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.

Assess temporal effects on stream processing applicationsEvaluate benchmark result accuracy across cloud regionsQuantify cloud performance variability impact on benchmarks

How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis

Mar 14, 2024
BT
B. Turkkan
🏛️ IBM Research | University at Buffalo | University of New Hampshire | Microsoft | MongoDB

Existing distributed coordination services lack standardized testing methodologies and tools, resulting in incomplete evaluations and non-comparable results. Method: We conduct a systematic survey of evaluation practices across mainstream coordination services, identify critical gaps in benchmarking consistency, fault tolerance, and scalability, and distill six core evaluation requirements. Leveraging literature analysis and cross-tool comparison, we identify 12 key evaluation parameters and categorize five typical defects. Contribution/Results: We propose a standardized evaluation framework incorporating consistency models, fault injection, and distributed topology configurations. Furthermore, we introduce a reproducible, comparable, and scenario-driven benchmarking guideline—specifically designed for coordination services—that fills a critical gap in the domain’s dedicated evaluation ecosystem.

Coordination ServicesDistributed SystemsTesting Frameworks

This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.

application benchmarkscloud service benchmarkingmicrobenchmarks

Towards an Optimized Benchmarking Platform for CI/CD Pipelines

Oct 21, 2025
NJ
Nils Japke
🏛️ Technische Universität Berlin | DATEV eG

Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.

Detecting performance regressions in CI/CD pipelines earlyIntegrating benchmark optimizations into practical CI/CD systemsOptimizing resource-intensive benchmarks for efficient execution

An HPC Benchmark Survey and Taxonomy for Characterization

Sep 10, 2025
AH
Andreas Herten
🏛️ Forschungszentrum Jülich | Lawrence Livermore National Laboratory | Texas A&M University

The high-performance computing (HPC) domain suffers from an abundance of benchmarking tools and the absence of a standardized, unified classification framework. Method: This paper proposes the first standardized benchmark taxonomy for HPC, derived from a systematic literature review and multi-dimensional feature analysis across hardware, software, and algorithmic layers. A structured classification model is constructed, with key attributes—including target workload, portability, scalability, and measurement granularity—concisely tabulated. An interactive web-based platform is further developed to enable dynamic, dimension-driven querying, cross-benchmark comparison, and visual analytics. Contribution/Results: The taxonomy systematically organizes over 100 mainstream HPC benchmarks, significantly enhancing efficiency and consistency for architects, researchers, and scientific users in system evaluation, benchmark selection, and performance optimization. It establishes a foundational framework for standardizing HPC performance assessment and facilitates reproducible, comparable, and interpretable benchmarking practices.

Developing taxonomy to categorize HPC benchmarks systematicallyProviding structured comparison of hardware and software evaluation toolsSurveying existing HPC benchmarks for comprehensive characterization

Latest Papers

What's happening recently
View more

This study addresses the lack of empirical evidence regarding performance disparities between microservice and monolithic architectures in e-commerce scenarios. Utilizing k6, we conducted quantitative load testing on both architectural paradigms sharing identical application logic and database backends. Experimental results demonstrate that under a 100-concurrent-user workload, the microservice architecture achieves a 5.4% increase in throughput and a 39% reduction in p95 tail latency, alongside lower error rates and superior fault isolation. These findings bridge the empirical gap in runtime performance comparison, validating the scalability advantages and specific failure modes of microservices under high load. Consequently, this research provides a reliable quantitative basis for informed architectural decision-making in e-commerce systems.

E-Commerce ApplicationMicroservices ArchitectureMonolithic Architecture

This study addresses the infrastructure complexity of cloud-edge-end协同 architectures, which has emerged as a major bottleneck hindering developer productivity and innovation. Through 101 semi-structured interviews across 86 organizations, this work empirically identifies deployment complexity and onboarding difficulty as core challenges. It proposes four architectural directions to mitigate these issues: Object-as-a-Service (unified object abstraction), internal developer platforms, declarative AI/ML pipelines, and lightweight edge runtimes. Findings indicate that high-level abstractions and automation significantly enhance developer experience—outweighing the impact of execution performance optimizations—and thereby establish a new paradigm for platform engineering and distributed system design.

cloud-edge infrastructuredeveloper productivitydistributed computing

This study addresses the longstanding fragmentation in microservice energy efficiency research, which has been siloed across runtime, infrastructure, and architectural layers, lacking a unified lifecycle perspective and consistent measurement methodology. Employing Kitchenham’s systematic literature review approach—augmented by searches across four major databases and snowballing techniques—the authors analyze 40 core studies to integrate multidimensional viewpoints for the first time. Their synthesis reveals an overwhelming emphasis on runtime optimizations, such as scheduling and resource management, typically relying on coarse-grained monitoring and model-based estimations, while largely neglecting energy-aware integration during architectural design and fine-grained measurement practices. The work establishes energy efficiency as a critical cross-cutting architectural attribute throughout the microservice lifecycle and underscores the urgent need for early-design support and a unified measurement framework.

architectural concernenergy efficiencymicroservice architectures

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

Hot Scholars

SF

Stefano Forti

Department of Computer Science, University of Pisa
cloud-edge continuumdistributed systemsgreen computingautomated reasoning
IL

Isaac Lera

Universitat de les Illes Balears
PerformanceCloud ComputingSemantic Web
DH

David Hästbacka

Associate Professor (tenure track), Tampere University
Software EngineeringSoftware ArchitectureIndustrial InformaticsSmart Energy Systems