Score
Design and implement instrumentation and telemetry that capture and record fine-grained runtime events, timestamps, and causal links across components, including trace collectors and loggers that synchronize and timestamp events. Build pipelines and analysis tools to reconstruct execution histories as directed execution graphs, compare passing and failing traces, and localize root causes from collected traces.
Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.
WebAssembly debugging and monitoring face three key challenges: tool fragmentation, high overhead from generic instrumentation frameworks, and the labor-intensive development of low-level, high-performance probes. This paper introduces Whamm—the first declarative instrumentation DSL tailored for WebAssembly—unifying bytecode rewriting and engine-internal support to jointly optimize abstraction and performance. Its core contributions are: (1) a programming model based on declarative matching rules, static/dynamic predicates, and automatic state reporting; and (2) deep integration with Wasm runtimes, enabling inline probes and intrinsification optimizations. Evaluated across diverse monitoring tasks, Whamm delivers expressive power comparable to manual implementations while achieving near-native performance, significantly reducing instrumentation overhead. It maintains compatibility with mainstream Wasm runtimes and supports seamless cross-platform deployment.
Existing debugging methods for code agents struggle to trace state transitions and error propagation in complex tasks and lack scalability. This work proposes a traceable architecture that employs an evolutionary extractor to parse heterogeneous runtime logs, constructs a hierarchical trajectory tree integrated with persistent memory, and introduces a fault-origin localization algorithm. For the first time, this approach enables automatic reconstruction of the complete state transition history and precise error tracing for agents executing multi-stage, parallel tool invocations. Evaluated on the CodeTraceBench benchmark, the method significantly outperforms existing baselines and supports accurate replay of original failing execution paths through diagnostic signal regeneration.
This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.
Debugging embedded programs is notoriously challenging due to tight software-hardware coupling, and existing tools often rely on external hardware probes or serial logging, resulting in low efficiency. This work proposes Inline, a novel programming tool that, for the first time, enables real-time inline visualization of hardware logs directly within source code. It introduces a domain-specific expression language to support programmable manipulation of logs, allowing developers to intuitively trace execution flow and precisely localize faults. Seamlessly integrated into standard embedded development environments, Inline significantly lowers the barrier to effective debugging. A user study with twelve participants demonstrates marked improvements in both debugging efficiency and accuracy when using the tool.
This work addresses the limitations of traditional structural coverage metrics in embedded software testing, which are often confined to the unit level and fail to reflect true coverage completeness in integration and system testing. Instrumentation-based approaches risk perturbing runtime behavior, while pure tracing techniques suffer from unreliability under high compiler optimization. To overcome these challenges, the paper proposes an integration-test-driven coverage strategy featuring a novel “integration-first” closed-loop workflow. By synergistically combining embedded tracing with hybrid runtime analysis (hRA) to preserve semantic boundaries, and leveraging source-to-target mapping for evidential traceability alongside Hyper Coverage for cross-variant merging, the approach establishes a unified evidence-integration mechanism. Evaluated on -O3-optimized release binaries, it reliably achieves branch, condition, and MC/DC coverage measurements and precisely identifies source code lines consistently uncovered across all variants, thereby significantly enhancing confidence in the test completeness of embedded systems.
This work addresses the challenge of automatically recovering traceability links among software architecture documentation, models, and source code—a longstanding barrier to effective system maintenance and consistency assurance. To bridge this gap, we present the first end-to-end ecosystem for architecture-level traceability recovery, comprising a RESTful API supporting four distinct tracing pipelines, an interactive web-based frontend named TraceView, and TraceViz, an embedded visualization plugin for Visual Studio Code. The system integrates seamlessly into developer workflows through asynchronous task processing and caching optimizations, enabling intuitive exploration of traceability links directly within the IDE. All components are publicly deployed, and preliminary user studies indicate that TraceViz significantly enhances developers’ cognitive efficiency during software comprehension tasks.