ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of perspective awareness in large language model (LLM) agents within industrial scenarios, where they fail to mitigate irreversible risks based on user roles. To this end, we introduce and formally quantify "perspective awareness" as a novel evaluation dimension. Methodologically, we construct an action trajectory benchmark incorporating perspective constraints by leveraging text-based world models to simulate operational environments alongside domain expert-validated anonymized query data. Experimental results demonstrate that even state-of-the-art models achieve only a 69% task completion rate, with over half of the generated trajectories involving unauthorized operations. This work bridges a critical gap in role-adaptability evaluation and reveals significant deficiencies in the perspective awareness capabilities of existing LLM agents.
📝 Abstract
Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial maintenance and equipment fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to user's role: taking actions and providing information that respect the role's knowledge and capability boundaries. Unlike coding, where mistakes are usually recoverable, agent responses in these settings are enacted on physical equipment, and can therefore cause irreversible equipment damage, production loss, or personnel harm. Existing benchmarks, however, largely overlook the need for agents to infer what a role intends and acting only through tools that role may legitimately use, a capability which we term Perspective Awareness. To this end, we introduce ReFract, a benchmark of 150 expert-validated entries in which an agent must act differently in response to the same query depending on user's role. Entries of ReFract are grounded in anonymized queries from domain support conversations, against which we construct Text World Models that simulate the agent's operating environments and assemble perspective-aware action trajectories. State-of-the-art LLMs solve at most 69% of the tasks with more than 50% of their trajectories contain attempts of taking perspective-violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and motivates agents that calibrate not just how to act, but for whom.
Problem

Research questions and friction points this paper is trying to address.

Perspective Awareness
LLM Agents
Benchmarking
High-stakes Settings
Role Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Perspective Awareness
Text World Models
LLM Agents
Benchmark
Role-Calibrated Actions