Score
Design and implement tools, formats, and storage for capturing, serializing, indexing, versioning, and restoring snapshots of a system or program's execution state; build utilities to diff and compare those snapshots to detect behavioral or state changes and to measure resource and storage differences for historical replay and analysis.
Current approaches to software evolution analysis and continuous integration rely heavily on test pass/fail outcomes, often overlooking fine-grained runtime behavior, which limits their ability to detect partial oracles, flaky failures, and silent performance or output drifts. This work proposes a novel paradigm—behavioral co-versioning—that jointly manages Git code history with queryable archives of runtime behavior. During test execution, method-level inputs, outputs, and performance signals are captured and stored append-only, indexed by commit and test context. By treating runtime behavior as a first-class artifact, the approach enables semantic-level differencing, behavior-aware regression localization, and historical auditing. A Python-based prototype demonstrates feasibility, successfully uncovering behavioral evolutions invisible to conventional textual diffing techniques.
This work addresses the challenge of automatically recovering traceability links among software architecture documentation, models, and source code—a longstanding barrier to effective system maintenance and consistency assurance. To bridge this gap, we present the first end-to-end ecosystem for architecture-level traceability recovery, comprising a RESTful API supporting four distinct tracing pipelines, an interactive web-based frontend named TraceView, and TraceViz, an embedded visualization plugin for Visual Studio Code. The system integrates seamlessly into developer workflows through asynchronous task processing and caching optimizations, enabling intuitive exploration of traceability links directly within the IDE. All components are publicly deployed, and preliminary user studies indicate that TraceViz significantly enhances developers’ cognitive efficiency during software comprehension tasks.
This paper addresses the challenge of verifying micropatches in binary code. We propose an automated, binary-level differential analysis method for comparing two program versions. Leveraging angr-based symbolic execution, our approach identifies reachable final states under identical inputs, enabling fine-grained matching and root-cause attribution of differences in register states, memory contents, and side effects. We introduce a novel compatibility-aware state alignment mechanism to overcome mismatches arising from control- and data-flow inconsistencies at the binary level. An interactive web visualization interface—built with React and TypeScript—supports dynamic difference pruning, multi-dimensional filtering, and provenance exploration. Evaluated on real-world firmware and application binaries, our method demonstrates strong capability in verifying functional equivalence and security preservation of micropatches, significantly improving verification efficiency, precision, and interpretability.
This study addresses the lack of systematic educational resources for Mining Software Repositories (MSR) instruction targeting secondary and tertiary (undergraduate, master’s, and doctoral) students. Methodologically, it introduces the first pedagogical MSR framework by decomposing MSR methodology into teachable knowledge modules—integrating version-control analytics, defect-data extraction, communication-log mining, and qualitative-quantitative mixed-method analysis—while embedding ethical guidelines, cross-method validation, and scaffolded learning design. Its contributions include: (1) a tiered hands-on practice system; (2) instructor-facing teaching guides; and (3) ready-to-use, curated educational datasets. The resulting standardized curriculum package spans all academic levels. Empirical evaluation demonstrates that the framework significantly enhances students’ ability to conduct reproducible, ethically grounded empirical studies on real-world software repositories—thereby filling a critical gap in standardized MSR education.
In digital forensics, the atomicity and integrity of storage snapshots lack rigorous definitions that jointly guarantee both instantaneousness and causal ordering—undermining evidentiary admissibility in legal proceedings. To address this, we propose a novel atomicity definition grounded in causal consistency, overcoming the limitation of conventional time-based atomicity models. We further rectify conceptual flaws in existing integrity definitions and introduce a revised, theoretically sound yet engineering-practical integrity criterion—explicitly supporting copy-on-write (CoW) implementations. Our approach integrates causal modeling, formal snapshot semantics, CoW mechanism analysis, and formalization of forensic quality criteria, yielding a verifiable snapshot semantic framework. This work establishes the first theoretical foundation for forensic tool design that unifies causal ordering with instantaneous state capture, thereby significantly enhancing the forensic validity and judicial admissibility of live data acquisition.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.
This work addresses the long-standing challenge in storage research of lacking scalable and reproducible experimental artifacts. To overcome this limitation, the study systematically proposes a reproducibility framework for storage experiments and implements it by leveraging the Chameleon cloud infrastructure and the Trovi platform. By integrating standardized packaging and environment configuration techniques, the authors construct six publicly available, unrestricted, and fully reproducible storage experiment artifacts. These artifacts substantially enhance experimental transparency and continuity, and have already been successfully deployed across multiple research communities and educational initiatives. The resulting infrastructure establishes a robust foundation for reproducible research in the storage domain.