Score
Designs and conducts quantitative analyses and comparisons of system or model performance metrics, producing comparative tables, curves, counters, and statistical summaries that quantify differences and trade-offs. This includes per-class and stratified evaluations, cost–performance and performance–accuracy trade-offs, variability and trend analyses, performance decomposition and gap analysis, and other empirical methods to identify trends and support design decisions.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.
Configuration space explosion complicates performance impact modeling, while gray-box approaches rely on structural knowledge (e.g., module execution graphs) to improve model accuracy—yet the mechanisms by which structural features (e.g., number of modules or configuration options) and structural knowledge influence modeling difficulty and optimization potential remain unclear. Method: We formally define “modeling hardness” and “improvement opportunity,” establishing an analytical framework and matrix to quantify the interplay among system structural complexity, structural knowledge level, and modeling benefit. Controlled experiments on synthetic systems integrate module execution graph analysis with gray-box modeling. Contribution/Results: We identify module count and configuration option count as dominant determinants of modeling hardness. Under high hardness, strong structural knowledge significantly increases improvement opportunity. Structural knowledge primarily enhances ranking accuracy, whereas hardness predominantly degrades prediction accuracy. Our findings provide theoretical foundations and strategic guidance for allocating structural knowledge investment according to specific modeling objectives.
This study systematically evaluates the impact of model quantization on the correctness and resource efficiency of deep learning systems, while also exploring methodologies for cross-study evidence aggregation in data-driven empirical research. Methodologically, it innovatively applies Structured Synthesis Methods (SSM) for the first time in this domain, integrating findings from six empirical studies covering 19 models through a qualitative-quantitative mixed analysis. Results demonstrate that quantization yields substantial resource gains—average storage compression of ×3.2, inference latency reduction of −41%, and GPU energy consumption decrease of −38%—with only a marginal correctness degradation (−1.7% on average), representing a well-controlled trade-off. The study identifies both consistent patterns and fragmentation bottlenecks in quantization effects, and proposes a refined empirical research framework and methodological guidelines tailored to quantization techniques. These contributions provide foundational methodological support and practical guidance for optimizing trustworthy AI systems.
Detecting performance regressions in configurable software is costly, and configuration sampling often misses localized performance degradation. Method: This paper proposes ConfFLARE, a technique that combines data-flow dependency analysis with change-impact propagation tracking to identify code changes interacting—via data flow—with performance-sensitive code. It further integrates configuration-feature identification to automatically select the subset of performance-sensitive configurations most likely affected by each change. Contribution/Results: ConfFLARE eliminates the need for exhaustive configuration-based performance testing. In evaluations on synthetic and real-world systems, it reduces the number of required test configurations by 79% and 70%, respectively, while achieving near-complete coverage of performance regression cases. It precisely pinpoints relevant features and significantly improves both the efficiency and completeness of performance regression detection.
Cloud performance is influenced by multi-scale, time-varying factors, and existing decomposition methods struggle to effectively capture their intermittent behavior and complex periodic patterns. This work proposes both a hybrid (expert-informed) and a fully automated time series decomposition approach that, for the first time, stably extracts multi-scale trend and seasonal components from a single performance trace. These components are directly leveraged for Serverless function performance prediction and AWS resource scheduling. The proposed method significantly outperforms baseline approaches, achieving prediction MAPE as low as 1.8% (hybrid) and 2.1% (fully automated), reducing latency variability on AWS by over 60%, and decreasing peak latency by up to 10%, thereby offering high-precision decision support for cloud resource provisioning.
This work addresses the challenge that state-dependent behaviors in modern computing environments—such as those introduced by adaptive system mechanisms—induce time-dependent biases in traditional software benchmarking, undermining reliable performance comparisons. The paper reframes benchmarking as a decision problem aimed at identifying the fastest program and introduces an experimental paradigm centered on pairwise performance comparisons, thereby avoiding strong assumptions about modeling system dynamics. By leveraging contrastive estimators, consistent statistical inference, and test strategies under finite evaluation budgets, the approach eliminates program-specific biases without relying on absolute performance metrics. The method provides asymptotic guarantees for correct decisions in stateful environments, offering a robust and reliable benchmarking framework for performance-sensitive software development.
This work addresses the high false positive rate (12.5%) and false negative rate (6.8%) of Mozilla’s existing T-test–based performance anomaly detection system, which hampers continuous integration efficiency. The authors introduce the first benchmark dataset comprising 174 engineer-annotated performance time series and conduct a systematic evaluation of 25 change-point detection algorithms combined with 15 ensemble strategies. They propose an ensemble voting mechanism that integrates offline, online, and hybrid methods to effectively mitigate the precision–recall trade-off. Experimental results and engineer feedback demonstrate that the proposed approach improves the F1-score by 11% over the original system and has been successfully integrated into Mozilla’s performance engineering infrastructure.
This work addresses the challenge of automatically translating natural language descriptions of software performance requirements into precise mathematical formulations, a task often hindered by linguistic ambiguity and cognitive uncertainty. The authors propose an interactive, retrieval-augmented preference elicitation method that uniquely integrates domain-specific knowledge into both preference inference and dialogue guidance. By leveraging this knowledge to steer conversational interactions, the approach incrementally refines user intent into accurate mathematical functions. Evaluated on four real-world datasets, the method substantially outperforms ten state-of-the-art baselines, achieving up to a 40-fold improvement in performance with only five rounds of interaction. This significant gain demonstrates its effectiveness in reducing users’ cognitive load while enhancing the efficiency and precision of requirements engineering.