Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of Attack Success Rate (ASR) in existing LLM safety evaluations, where ASR fails to distinguish between model refusal, task non-recognition, and intent misinterpretation, rendering a low ASR insufficient to indicate high safety. This work proposes jointly evaluating ASR with an Operational Comprehension Rate, employing intent-obscuring prompts, English reconstruction experiments, and formal logic contrastive analysis to precisely determine whether models genuinely recognize and process harmful tasks. The research reveals, for the first time, disparities in intent recovery masked by ASR: models exhibit significantly different comprehension rates under similar ASRs, and high comprehension may still accompany frequent harmful assistance. These findings demonstrate the inadequacy of single-metric evaluation and underscore the necessity of jointly reporting both metrics.
📝 Abstract
Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at https://github.com/kevinjiang0121-cyber/IRIS.
Problem

Research questions and friction points this paper is trying to address.

LLM safety evaluation
attack success rate
intent recovery
operative understanding rate
intent-obscuring prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent Recovery
Safety Evaluation
Operative Understanding Rate
Attack Success Rate
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haitong Jiang
College of Computer Science and Software Engineering, Shenzhen University
C
Chunlin Liu
College of Computer Science and Software Engineering, Shenzhen University
S
Sihan Tang
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen), Shenzhen, China
C
Chan Wu
College of Computer Science and Software Engineering, Shenzhen University
X
Xiaoqing Su
College of Computer Science and Software Engineering, Shenzhen University
Yuhong Feng
Yuhong Feng
Associate Professor
Workflow ManagementCloud ComputingThe Internet of thingsLinux Operating System