AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of runtime risks faced by large language model (LLM)-driven agents throughout their complete execution lifecycle. The authors propose the first end-to-end, four-dimensional runtime safety assessment framework and introduce AgentS4D, a sandbox benchmark comprising 328 risk scenarios generated from six risk entry points, six elicitation strategies, and nine harm categories, with evidence collected at seven critical checkpoints. Through 6,560 experimental runs across four agent frameworks and five LLMs, they find that 68.0% of executions trigger unsafe signals, and notably, 66.22% of these still successfully complete their tasks—demonstrating that task completion rate is an unreliable proxy for safety. The results further confirm that the interplay between risk carriers and elicitation strategies significantly influences system safety.
📝 Abstract
Large language model (LLM)-based workspace agents execute stateful, multi-step workflows across heterogeneous resources, external tools, and persistent state. Their safety must therefore be assessed from actions, side effects, and state changes throughout execution. Although recent benchmarks have advanced executable safety testing and trajectory-aware verification, they rarely provide a unified account of where risks enter, how they elicit unsafe behavior, which harms they target, and where supporting evidence appears during execution. We introduce AgentS4D, a sandboxed benchmark for lifecycle-wide runtime safety evaluation. Its four-dimensional runtime-safety framework uses six risk-entry sources, six induction strategies, and nine target harms to guide case construction, while seven lifecycle checkpoints organize post-run evidence. AgentS4D contains 328 risk-injected cases. We evaluate all 20 combinations of four harnesses (Hermes, OpenClaw, Claude Code, and Codex) and five LLM backends (GPT-5.5, Gemini 3.1 Pro, DeepSeek-V4-Pro, MiniMax-M3, and Qwen3.7-Plus) on these cases, yielding 6,560 runs. Overall, 4,461 runs (68.0%) trigger prespecified unsafe signals. Across the 20 configurations, the observed safety of an agent system varies with both its harness-LLM pairing and how risk is introduced. Agent systems exhibit markedly different safety behavior when the same induction strategy reaches them through different risk carriers. They also respond differently to the same target harm when it is realized through different carriers and strategies. Moreover, 4,344 runs (66.22% overall) are unsafe yet complete. Thus, task completion cannot establish runtime safety, and testing only one form of a risk can conceal important weaknesses. Evaluations should examine complete agent configurations across diverse risk conditions and retain evidence throughout execution.
Problem

Research questions and friction points this paper is trying to address.

runtime safety
LLM-based agents
risk evaluation
execution lifecycle
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

runtime safety
LLM-based agents
risk benchmarking
execution lifecycle
safety evaluation
🔎 Similar Papers
Jiajun Zhou
Jiajun Zhou
Zhejiang University of Technology
Graph Data MiningGraph Data AugmentationBlockchain Data AnalysisGraph for Cybersecurity
Z
Zhaoxuan Ke
College of Information Engineering, Zhejiang University of Technology, Hangzhou 310023, China
J
Jihang Ye
Institute of Cyberspace Security, Zhejiang University of Technology, Hangzhou 310023, China; Binjiang Institute of Artificial Intelligence, ZJUT, Hangzhou 310056, China
Xuanze Chen
Xuanze Chen
Ph.D, HongShan Capital
Biophotonicssuper-resolution microscopy
S
Shanqing Yu
Institute of Cyberspace Security, Zhejiang University of Technology, Hangzhou 310023, China; Binjiang Institute of Artificial Intelligence, ZJUT, Hangzhou 310056, China
Qi Xuan
Qi Xuan
Professor, Zhejiang University of Technology
AI SecuritySocial NetworkDeep LearningData Mining