sandbox dynamic analysis

Design and operate isolated execution environments and interactive analysis workflows that run software, documents, or other artifacts to observe and record runtime behavior and content, including OS‑level interactions, interprocess activity, file and network operations, and triggered hidden events. Build instrumentation, traces, and reports that audit and catalog safety‑relevant execution findings — e.g., information‑flow evidence, malicious effects, or policy violations — to support detection, triage, and reproducible analysis.

sandboxdynamicanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.

action tracedata analysisend-to-end workflow

This work addresses the inefficiency in notebook-based distributed workflows, where minor modifications often trigger full re-execution, severely hindering iterative development and reproducibility. To overcome this limitation, the authors propose NBRewind, a system that, for the first time, enables fine-grained incremental execution and cross-platform portability while preserving reproducibility. NBRewind integrates a dual-kernel architecture—comprising auditing and replay components—with cell-level incremental checkpoints and inter-cell dataflow analysis. It further leverages standardized notebook packaging to facilitate efficient partial re-execution. Evaluation in real-world high-performance computing (HPC) scenarios demonstrates that NBRewind incurs minimal overhead for incremental checkpointing and substantially improves both execution efficiency and cross-site reproducibility.

checkpointingdistributed workflowsiterative development

This work addresses semantic isolation issues in persistent AI workflows, where dynamic changes in prompts, model aliases, or tools during execution can cause inconsistencies between runtime state and underlying assumptions. The paper formally characterizes four classes of semantic anomalies and introduces a multi-level semantic isolation model that establishes a complete isolation spectrum—from “semantic read-committed” to “semantic snapshot isolation”—through three guarantees: resource stability, cross-resource compatibility, and continuation inheritance. A lightweight middleware is designed to support semantic context propagation, dynamic binding validation, branch-merge control, and microsecond-scale compatibility checks. Evaluation of the prototype system, SemIso, on LangGraph reveals that 7.4% of persistent workflows exhibit semantic binding risks, and demonstrates its effectiveness in efficiently intercepting incompatible operations.

AI workflowsdurable executionisolation anomalies

Applying Process Mining on Scientific Workflows: a Case Study

Jul 06, 2023
ZS
Zahra Sadeghibogar
🏛️ RWTH Aachen University

SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.

Correlate jobs with explicit or implicit dependencies.Document workflows and identify performance bottlenecks.Extract case IDs from SLURM-based HPC logs.

Traditional testing approaches rely on command-line interfaces, operate at coarse granularity, and lack runtime context—causing isolation from highly interactive development environments (HIDEs) and introducing disruptive context switches that exceed developers’ attention thresholds. This paper proposes a runtime-aware testing library that deeply integrates unit testing into HIDEs via three core mechanisms: (1) runtime code instrumentation, (2) dynamic test registration, and (3) a context-aware execution engine. These enable immediate invocation of interactive debugging capabilities—including value inspection and stack trace visualization—upon test failure. The approach achieves sub-second test re-execution with end-to-end response times under the critical 2-second threshold, substantially reducing context-switching overhead and enhancing development continuity and debugging efficiency. Its principal contribution is the first semantic-level coordination mechanism between testing infrastructure and the HIDE toolchain, enabling bidirectional, context-rich interaction between tests and the live development environment.

Slow test reexecution breaks developer focus thresholdsTests lack runtime context and IDE tool integrationTraditional testing disrupts flow in interactive development environments

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing large language models, which are typically confined to isolated tasks and struggle to integrate into industrial-scale, multi-stage security workflows. To bridge this gap, the authors propose the first role-based multi-agent framework tailored to the entire vulnerability lifecycle, incorporating specialized agents—Planner, Analyzer, Fixer, and Verifier—augmented with CodeQL static analysis for enhanced precision. By introducing a role-oriented multi-agent architecture into end-to-end vulnerability management, this approach effectively aligns the capabilities of large models with real-world security engineering demands. Evaluated on 25 real-world C/C++ vulnerabilities, the system achieves a detection accuracy of 44%—comparable to GPT-5.5—and a repair accuracy of 19%, offering a practical and collaborative paradigm for intelligent security operations.

LLM-based securityrole-based agentic architecturesecure software engineering

Traditional Software Bill of Materials (SBOM) approaches struggle to accurately capture the components dynamically loaded at runtime in languages like Python, thereby limiting supply chain security and incident response capabilities. This work proposes MEM-SBOM, the first memory forensics–based framework for generating runtime SBOMs without requiring prior instrumentation. By analyzing the in-memory structures of the Python interpreter, bytecode, and package version metadata, MEM-SBOM directly reconstructs the true execution state from process memory. This approach overcomes the limitations of methods relying on static metadata or runtime monitoring, enabling both post-incident forensic analysis and deployment in production environments. Evaluated on 51 real-world Python applications, MEM-SBOM achieves 100% accuracy in component extraction, fully recovers runtime dependencies missed by existing tools, and precisely identifies vulnerable function calls.

memory forensicsPython dependenciesruntime SBOM

This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.

debugginglarge language modelsreproducibility

This work addresses the lack of transparency in scientific AI agents when executing complex research workflows, which hinders human understanding, scrutiny, and intervention. To bridge this gap, the paper introduces the first monitoring and visualization framework that dynamically models an agent’s execution trace as a structured directed graph in real time. By capturing fine-grained intermediate events—such as tool invocations and code executions—the framework renders the evolving workflow structure explicitly. This approach substantially enhances the interpretability and controllability of AI-driven scientific processes. Empirical evaluations across AI, neuroscience, and biology demonstrate strong endorsement from domain experts, confirming its effectiveness in supporting result traceability, fault localization, and human–agent collaborative analysis.

execution traceshuman-AI collaborationscientific AI agents

This work addresses the vulnerability of large language model (LLM) agents to privilege escalation when processing attacker-controllable context, stemming from the absence of unified access control over fields, semantic payloads, and invocation events. To mitigate this, the paper introduces Context-to-eXecution Integrity (CXI), a mechanism that enforces fine-grained authorization at execution boundaries. CXI employs policy tags to protect fields, type-directed release for validating propagated values, and opaque data slots to preserve forensic evidence. Crucially, a deterministic gating mechanism binds field access, effect generation, and invocation permissions to a unified action manifest before execution proceeds. This approach achieves, for the first time, coordinated field-, effect-, and call-level permission binding, delivering triple-layered integrity guarantees. Evaluated on AgentDojo (720 tasks) and a code-agent benchmark (400 repository tasks), CXI successfully completed 231 security-sensitive tasks with zero escapes, demonstrating both efficacy and compatibility.

Authority CheckContext-to-Execution IntegrityLLM Agents

Hot Scholars

HZ

Hongyu Zhang

Chongqing University
Software EngineeringMining Software RepositoriesData-driven Software EngineeringSoftware Analytics
AN

Aran Nayebi

Assistant Professor of Machine Learning, Carnegie Mellon University (CMU)
Computational NeuroscienceArtificial IntelligenceDeep LearningMachine Learning
MV

Matteo Varvello

Researcher at Bell Labs
Web performancevideomiddleboxesCCN
QY

Qingqing Ye

Assistant Professor, The Hong Kong Polytechnic University
data privacy and securityadversarial machine learning
VP

Vasileios P. Kemerlis

Associate Professor, Brown University
OS SecuritySoftware HardeningFuzz Testing