shadow execution

Designs and implements systems that run an alternative or instrumented version of software alongside the primary implementation (or on recorded runtime traces) to log, replay, and compare execution behavior. Builds replay/simulation pipelines and analysis tools that quantify implementation mismatches and measure their downstream impact.

shadowexecution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

Current approaches to software evolution analysis and continuous integration rely heavily on test pass/fail outcomes, often overlooking fine-grained runtime behavior, which limits their ability to detect partial oracles, flaky failures, and silent performance or output drifts. This work proposes a novel paradigm—behavioral co-versioning—that jointly manages Git code history with queryable archives of runtime behavior. During test execution, method-level inputs, outputs, and performance signals are captured and stored append-only, indexed by commit and test context. By treating runtime behavior as a first-class artifact, the approach enables semantic-level differencing, behavior-aware regression localization, and historical auditing. A Python-based prototype demonstrates feasibility, successfully uncovering behavioral evolutions invisible to conventional textual diffing techniques.

Behavioral Co-Versioningcontinuous integrationexecution history

This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.

compiler optimizationembedded softwareintegration testing

Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.

debuggingexecution tracehypothesis generation

Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories

Jan 25, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa | Siemens AG

Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.

Data Analysis VariabilityResearch ReliabilitySoftware Engineering

Latest Papers

What's happening recently
View more

This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.

CI/CD pipelinesDevOpsDigital Twin

This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.

deterministic workflowsLLM agent tracestool-use

Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.

behavioural gapscode coverageexpected behaviour

This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.

AI agentsfoundational assumptionsreplication

Hot Scholars

JS

Jin Song Dong

Professor of Computer Science, National University of Singapore
Formal MethodsTrusted AISafe AIModel Checking
RM

Rui Miao

Meta
NetworkingNetworked SystemsDistributed Systems
WC

Wenyu Chen

Massachusetts Institute of Technology
optimizationstatistical learning
HC

Haibo Chen

FACM & FIEEE, Distinguished Professor, Shanghai Jiao Tong University
Operating SystemsMLSysDistributed Systems