Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of sub-agent identity spoofing and subsequent privilege hijacking in LLM multi-agent orchestration. We propose TrustFork, a benchmark that leverages adversarial identity substitution and conflicting evidence injection to quantify, for the first time, the influence of "brand trust" on execution privileges. Experimental results demonstrate that in 72% of scenarios, orchestrators adopt high-risk responses while disregarding contradictory evidence, revealing critical authorization blind spots inherent in identity-label-based delegation. Furthermore, this work establishes that concealing identity cues constitutes the optimal defense strategy and argues that privilege allocation should be anchored to objective evidence rather than identity labels. These findings provide essential theoretical and practical foundations for the secure orchestration of multi-agent systems.
📝 Abstract
LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents display, and an attacker can spoof them. Displayed identity thus decides operational authority, meaning who is trusted to check the work and who is allowed to change it. As a result, a risky subagent can keep authority over execution even after other evidence contradicts it. We introduce TrustFork, an LLM agent safety benchmark with 1,890 tasks and 27,826 valid trajectories across 16 agent systems. These systems run eight orchestrators under the OpenCode, OpenClaw, and Pi harnesses. In each task, one subagent carries a risky goal while the other three stay aligned with the user, so contradicting evidence can exist. A task can also change the identity a subagent displays without changing the model behind it, which lets us trace a shift in authority to the label. Our analysis shows that even when another subagent contradicts the risky response, the orchestrator still acts on it in 72.0% of cases on average, most often in the systems with the least terminal harm. Swapping the family labels nearly triples how often the orchestrator obtains the risky response. A safer response is available in 84.0% of tasks, yet it decides the outcome in only 25.0%. The harness also decides which responses reach the orchestrator. Among three runtime defenses, hiding identity cues helps most consistently, while verifying before action helps only when the harness returns enough evidence. TrustFork shows that production agent orchestration must bind authority to evidence before execution causes harm. Our project is in https://henrymao2004.github.io/agent-orchestration-safety/.
Problem

Research questions and friction points this paper is trying to address.

LLM agent orchestration
identity hijacking
trust exploitation
agent safety
subagent spoofing
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM agent orchestration
identity spoofing
TrustFork benchmark
runtime defenses
authority hijacking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xutao Mao
City University of Hong Kong
R
Rui Qian
Fudan University
Linghan Chen
Linghan Chen
Department of Materials Science, Tohoku University
Y
Yudong Gao
The Hong Kong University of Science and Technology
J
Junchi Liao
University of Electronic Science and Technology of China
J
Jiulin Cai
University of Science and Technology of China
J
Jinman Zhao
University of Toronto
Cong Wang
Cong Wang
Department of Computer Science, City University of Hong Kong
cloudsecuritybig datacomputation outsourcingaccess control