HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation regarding security mechanisms in coding agent frameworks and their runtime implications. We propose the first comprehensive empirical benchmark, establishing a taxonomy of ten security mechanisms. By integrating LLM-as-a-judge evaluation, deterministic oracles, and network isolation techniques, we quantify the utility and attack resilience of nine mechanisms across six mainstream frameworks using 400 cases spanning 23 tasks. Our findings reveal that automatic approval escalates the attack success rate from 29.2% to 95.6%, demonstrating that single-layer defenses are readily circumvented. Furthermore, this work provides verifiable security configuration recommendations that effectively balance legitimate task execution requirements with robust safety assurance.
📝 Abstract
Coding agent harnesses mediate tool use and authorize actions, yet their security mechanisms and runtime effects remain incompletely characterized. We present HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses. First, we derive a ten-mechanism taxonomy and then assess 400 harness-mechanism cells using independent ratings by researchers and large language model (LLM) judges. We find that about half of confirmed mechanism implementations are opt-in, while closed-source harnesses exhibit substantial evidence gaps. Second, we introduce HarnessSecurity-Bench, a benchmark of 23 tasks across five attack surfaces without sacrificing legitimate task requirements. Using separate deterministic oracles to measure task utility and attack effects with security setting comparisons, we evaluate nine mechanisms across six leading harnesses: Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, and GitHub Copilot. Under a controlled LLM baseline GLM-5.2, we conduct 2,500 trials, recording 81,155 tool calls and over 2.2 billion tokens. Enabling auto-approve increases utility and raises attack success from 29.2% to 95.6%. Network isolation and read-only mode reduce attack effects with substantial utility losses, while command allowlisting and command denylisting reduce attack effects with a small utility loss and a utility gain, respectively. Task-level cases show that restrictions on a shared capability can obstruct both legitimate and malicious operations, and that allowed tools or commands can leave unauthorized operations reachable through alternative execution paths. Harness providers should make security settings verifiable, test alternative execution paths to protected operations, and assess attack effects alongside task utility and execution costs.
Problem

Research questions and friction points this paper is trying to address.

Coding Agent Harnesses
Security Mechanisms
Attack Surfaces
Benchmark
AI Security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coding Agent Harnesses
Security Benchmark
Mechanism Taxonomy
Attack Surface Evaluation
Utility-Security Tradeoff
💼 Related Jobs
No related jobs found.
Z
Zhengyang Zhu
School of Software Engineering, Sun Yat-sen University, Zhuhai, China
L
Liming Huang
School of Software Engineering, Sun Yat-sen University, Zhuhai, China
R
Runmin Ji
East China Normal University, Shanghai, China
Mingxi Ye
Mingxi Ye
Sun Yat-sen University
Fuzz TestingBlockchain Security
Zihan Zhou
Zihan Zhou
South China University of Technology
Computer Vision,Image Processing,Deep Learning
H
Hanyang Guo
School of Software Engineering, Sun Yat-sen University, Zhuhai, China
J
Jingwen Wu
Department of Computer Science, Hong Kong Baptist University, Hong Kong, China
Y
Yuhan Ye
Tsinghua University, Beijing, China
Y
Yuming Feng
Department of New Networks, Peng Cheng Laboratory, Shenzhen, China
Hong-Ning Dai
Hong-Ning Dai
Hong Kong Baptist University
Industrial Internet of ThingsBlockchain TechnologiesExtended RealityBig Data Analytics
Zibin Zheng
Zibin Zheng
IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability