cross-platform evaluation

Designs and implements methods, datasets, instrumentation, and analysis pipelines to collect, annotate, test, and compare system behavior and performance across heterogeneous platforms, sites, or systems. Includes establishing provenance and annotator-tracking, building cross-platform test harnesses and evaluation protocols, and producing reproducible metrics and analyses that enable valid comparison and diagnosis across environments.

cross-platformevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories

Jan 25, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa | Siemens AG

Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.

Data Analysis VariabilityResearch ReliabilitySoftware Engineering

Automatic Metadata Capture and Processing for High-Performance Workflows

Jun 18, 2025
PS
Polina Shpilker
🏛️ Tufts University | Sandia National Laboratories

In heterogeneous high-performance computing (HPC) environments, workflow metadata collection remains challenging due to fragmentation, poor reusability, and lack of standardization—hindering FAIR (Findable, Accessible, Interoperable, Reusable) compliance and impeding efficient performance analysis. To address this, we propose an automated metadata management framework. It features a lightweight runtime collection mechanism supporting fine-grained metadata capture across heterogeneous workflow systems (e.g., Snakemake, Nextflow); a dual-format unified storage scheme combining JSON Schema (for semantic expressiveness and human readability) and SQLite (for efficient querying); and a performance-aware metadata schema redesign that natively models task dependencies, resource consumption, and temporal behavior. Experimental evaluation demonstrates significant improvements in metadata findability, interoperability, and reusability—enabling reproducible, data-driven workflow performance analysis.

Apply FAIR principles to enhance research reproducibilityCapture metadata for workflows on heterogeneous architecturesStandardize and reorganize metadata for performance analysis

To address the challenges of fragmented engineering data and inefficient, inconsistent manual reporting in large-scale software development, this paper proposes and implements a centralized engineering productivity analytics framework supporting near-real-time aggregation and visualization. Methodologically, the framework introduces a dual-mode data storage architecture coupled with a precomputation engine, integrated with scheduled data ingestion (cron), proactive alerting, and role-based access control (RBAC). It unifies heterogeneous data from multiple source systems via a dual-schema design and leverages Metabase to enable cross-dimensional visual analytics—spanning development efficiency, software quality, and operational effectiveness. Empirical evaluation demonstrates that the deployed system reduces average weekly manual reporting effort by 20 person-hours, significantly accelerates bottleneck identification latency, and concurrently improves engineering decision responsiveness and platform scalability.

Aggregating siloed DORA and KPI metrics from multiple engineering platformsProviding near-real-time visibility into developer experience and system healthReducing time-consuming manual reporting prone to inconsistencies

Latest Papers

What's happening recently
View more

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

This work addresses the challenges of fragmented and heterogeneous performance monitoring units in modern heterogeneous multi-core SoCs, which complicate data collection, synchronization, and correlation. The paper proposes a centralized performance monitoring architecture—demonstrated for the first time in a RISC-V SoC—that enables unified cross-component monitoring. Hardware microarchitectural events are collected by Event Monitoring Units (EVUs) and routed via the AXI4 bus to an Advanced Performance Monitoring Unit (APMU) for consolidated processing. By integrating programmable counters and a dedicated processing unit, the architecture simplifies interface design, enhances event correlation capabilities, and supports event-driven software mechanisms. Experimental results validate its effectiveness in real-time resource management, application profiling, and counter attribution scenarios.

Embedded SystemsEvent CorrelationHardware Performance Counters

This study addresses the widespread neglect in machine learning research of when validation occurs during data annotation—a critical factor influencing both label quality and cost—despite overreliance on post-hoc quality control. Drawing inspiration from the “shift-left” principle in software engineering, this work proposes a tripartite classification of quality checkpoints across early, intermediate, and late stages of the annotation pipeline and introduces a parameterized error propagation model that, for the first time, treats validation timing as a quantifiable design variable. Through error propagation modeling, process decomposition, and literature analysis, the authors find that only 4% of recent studies report validation timing. Their analysis demonstrates that early-stage quality checks can reduce error correction costs by up to two orders of magnitude. The paper calls for standardized reporting of timing configurations, platform support for tunable timing parameters, and empirical studies on stage-specific detection rates.

annotation pipelinesdata qualityerror propagation

This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.

compiler optimizationembedded softwareintegration testing

This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.

application benchmarkscloud service benchmarkingmicrobenchmarks

Hot Scholars

TZ

Tianwei Zhang

Nanyang Technological University
Computer System Security
SI

Sotiris Ioannidis

Technical University of Crete and Foundation for Research and Technology - Hellas
securitysystemshardwaremetasurfaces
YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
KI

Kevin I-Kai Wang

Department of Electrical, Computer, and Software Engineering, The University of Auckland
Wireless Sensor NetworkUbiquitous ComputingPervasive HealthcareMachine Learning
YD

Yuexi Du

PhD candidate @ Yale University
Computer VisionMedical Image AnalysisMulti-modal Learning