replayability testing

Designs, builds, and evaluates systems and test harnesses that reproduce program executions deterministically across runs and machines, including tools such as deterministic and bitwise-deterministic replayers, reference-patch replay, replayable verifiers, and deterministic baselines. Analyzes and diagnoses non-replayability by measuring and isolating sources of divergence (cross-machine differences, ordering variance, randomness/seed noise, state drift, memory-layout or scheduler effects) and quantifying replay failures.

replayabilitytesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.75
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.

deterministic workflowsLLM agent tracestool-use

This work addresses the challenge of irreproducible decision trajectories in tool-augmented large language model (LLM) agents within financial regulatory audit replay scenarios. To this end, we propose the Determinism-Faithfulness Assurance Harness (DFAH), a framework that systematically quantifies both trajectory determinism and evidence-conditioned faithfulness of LLM agents for the first time, revealing a positive correlation between these two properties. DFAH establishes a replay-capable agent evaluation paradigm tailored for financial compliance, integrating multi-model, multi-configuration benchmarking, deterministic trajectory tracing, and faithfulness assessment, accompanied by an open-sourced stress-testing toolkit. Experimental results across three financial compliance benchmarks demonstrate that Tier-1 models employing a schema-first architecture achieve the determinism required for audit replay, while non-agent configurations of 7–20B parameter models attain 100% determinism.

determinismfinancial servicesLLM agents

This work addresses the inefficiency of traditional fuzzing in black-box or obfuscated binary programs where static instrumentation is infeasible and control-flow feedback is unavailable. The authors propose a dynamic feedback mechanism based on Execution Divergence Graphs (EDGs), which constructs control-flow-like structures at runtime by analyzing execution traces to precisely identify path divergences and avoid redundant exploration of loops. Requiring no static program information, the approach integrates divergence detection with an EDG-guided input mutation strategy. Evaluated on multiple obfuscated targets, it substantially outperforms blind fuzzers, demonstrating its effectiveness in non-instrumented settings. Furthermore, the framework is extensible to multidimensional feedback channels, such as power consumption, broadening its applicability in side-channel-aware fuzzing scenarios.

black-box fuzzingcontrol-flow discoveryexecution traces

Traditional redundancy mechanisms are vulnerable to common-mode failures because replicated program instances share identical memory layouts and code. To address this limitation, this work proposes a structured address space decorrelation approach that generates multiple semantically equivalent program variants through independent compilation, each exhibiting distinct memory layouts. At runtime, the method extracts normalized instruction traces—comprising opcodes, registers, operands, and results—while eliminating address dependencies, and performs cross-variant comparison to detect faults. This technique effectively identifies common-mode errors induced by arbitrary program counter jumps or data pointer corruptions, thereby significantly enhancing the capability of runtime semantic consistency verification.

Divergent Multi-Version Executionfault detectioninstruction-trace

Latest Papers

What's happening recently
View more

This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.

automated verifiabilityimplementation diversityprogram verification

This work addresses the long-standing lack of systematic validation for processor specifications, which can lead to distorted program behavior and security vulnerabilities. It presents the first automated differential testing framework tailored for open-source SLEIGH specifications, automatically generating decodable instructions and initial execution states by parsing specification structures, and systematically validating them against multiple hardware reference implementations across architectures. Applied to x86-64 and AArch64, the approach uncovered 38,920 semantic discrepancies, identified 125 unique defects—many of which were subsequently fixed—and significantly improved specification fidelity. Furthermore, it exposed inconsistencies across vendor implementations and led to eight concrete recommendations, establishing a new paradigm for ensuring the reliability of instruction set architecture specifications.

disassembleremulatorprocessor specification

Analysis algorithms for quantitative automata are often complex and prone to subtle implementation bugs that existing testing approaches struggle to expose effectively. This work proposes a debugging framework that integrates non-degenerate random generation, property validation, and automated counterexample minimization to efficiently produce small, actionable counterexamples. Applied to implementations of parametric timed automata, the method successfully uncovered five previously unknown bugs in IMITATOR, a mainstream model checking tool. Notably, the simplest counterexample involves only two locations and one transition, demonstrating a significant improvement in both bug detection capability and debugging efficiency.

algorithm debuggingimplementation bugsmodel checking

This study addresses the challenge of reproducing failures in cyber-physical system (CPS) simulation testing, where non-deterministic behaviors often hinder consistent fault manifestation. To tackle this issue, the work extends delta debugging to stochastic CPS scenarios for the first time, introducing three novel delta debugging algorithms tailored for randomized environments. The proposed approach integrates statistical failure analysis, repeated execution, and environment-aware input minimization to identify a minimal triggering input while preserving fault semantics. Empirical evaluation on case studies involving elevator scheduling and autonomous mobile robots demonstrates that the method significantly enhances the stability of failure reproduction and substantially reduces debugging time, effectively mitigating execution flakiness inherent in such systems.

Cyber-Physical Systemsdelta debuggingfailure reproducibility

Hot Scholars

DS

Dawn Song

Professor of Computer Science, UC Berkeley
Computer Security and Privacy
AS

Alireza Shirmarz

Postdoc Reseracher at Federal University of Sao Carlos (UFSCar)
SDNML/DLP4/TofinoCloud Gaming(CG)
YS

Yinzhe Shen

PhD student at KIT
end-to-end autonomous driving
GY

Guangba Yu

Postdoc, The Chinese University of Hong Kong
Cloud ComputingLLMOpsAIOpsDistributed Systems