ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified evaluation framework for risk handling in coding agents by constructing the first benchmark grounded in software engineering risk management theory. Methodologically, it proposes the ATMA framework alongside violation rate and evidence responsiveness metrics, employing human-calibrated agent judges and repository-level tasks for systematic evaluation. The research establishes risk handling as an independent capability dimension, revealing over-defensiveness rates ranging from 11.2% to 58.7%. It demonstrates that strong task performance does not equate to appropriate risk handling, and confirms that violations significantly degrade the developer experience.
📝 Abstract
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
Problem

Research questions and friction points this paper is trying to address.

coding agents
risk treatment
benchmarking
unnecessary defensive work
software engineering risk management
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coding Agents
Risk Management
Benchmark
Agentic Judge
ParanoiaEval
🔎 Similar Papers
No similar papers found.
Hanjun Luo
Hanjun Luo
New York University Abu Dhbai
Trustworthy AILarge Language ModelText-to-Image
X
Xiucheng Zhang
New York University
Z
Zhuoning Xu
New York University
Z
Zhimu Huang
New York University Abu Dhabi
Y
Yingbin Jin
The Hong Kong Polytechnic University
X
Xinfeng Li
The Hong Kong Polytechnic University
Hanan Salam
Hanan Salam
SMART lab @NYU Abu Dhabi / Co-founder of Women in AI
Artificial IntelligenceHuman-Machine InteractionHuman-Robot Interaction