AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

📅 2026-09-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出AURA-Eval框架,通过控制增强和行为细粒度诊断来评估LLM代理在工具使用轨迹中的风险识别和安全任务完成能力。
📝 Abstract
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories. Its pipeline identifies safety-critical decision points, generates controlled variations, and constructs counterparts differing in whether a request has a safe fulfillment path. Using 157 sourced trajectories, we generate 1,249 evaluation items and evaluate 20 frontier and open-weight models. We developed rubrics to classify risk detection, action strategy, and scenario-specific action safety. Our results show that LLM agents engage in unsafe behavior more often when no safe fulfillment path exists. In these cases, frontier proprietary models more often recognize risk and exhibit safer behavior by proposing alternatives, while evaluated open-weight models more often directly execute unsafe requests. Increasing impact or reducing opportunities for oversight before execution also exposes greater vulnerability across models.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
safety evaluations
risk recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

AURA-Eval
risk awareness
safety-critical decision points
controlled augmentation
granular diagnosis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruoxi Shang
University of Washington
C
Christina-Maria Androna
Oumi
Orfeas Menis Mastromichalakis
Orfeas Menis Mastromichalakis
PhD Student, National Technical University of Athens
Explainable AIAI EthicsNLP
Yu Feng
Yu Feng
University of Pennsylvania
Natural Language ProcessingMachine Learning
A
Aniruddhan Ramesh
Oumi, University of Cincinnati
Rico Angell
Rico Angell
Graduate Student, University of Massachusetts Amherst
S
Shang Hong Sim
Oumi
C
Chrysoula Zerva
National Technical University of Athens
E
Emmanouil Koukoumidis
Oumi