CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of LLM-based agents to credential leakage under phishing attacks, highlighting the inherent tension between security and task utility. To investigate this, we construct a comprehensive benchmark encompassing both user-instructed and autonomous monitoring scenarios, employing sandbox isolation to measure agents’ actual submission behaviors rather than relying on self-reports. By jointly quantifying information leakage rates and task completion metrics, our evaluation reveals that all mainstream models exhibit susceptibility to such leakage. Furthermore, we demonstrate that naively suppressing information disclosure significantly degrades performance on legitimate tasks. This work provides an empirical framework for evaluating the security-utility trade-off dilemma and establishes a critical benchmark for designing robust, secure AI agents.
πŸ“ Abstract
Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing. Safe execution requires distinguishing malicious requests from genuine ones without simply refusing to act. Despite its practical importance, this problem remains underexplored and it is unclear whether current agents or existing defenses can achieve it. To study this problem, we first propose CredLeak-Bench, a comprehensive benchmark designed to evaluate how effectively and securely agents automate human workflows when confronted with phishing and identity verification. The benchmark covers both user-directed authentication and autonomous inbox monitoring, where agents are not explicitly instructed to log in. It systematically varies deceptive cues and pairs phishing scenarios with legitimate counterparts, enabling joint evaluation of information leakage and utility on genuine tasks. Within a sandboxed environment, leakage is measured through actual submissions of information rather than agents' self-reported behavior. Our evaluation reveals that all tested models are vulnerable to leakage. Agents also disclose sensitive information during autonomous inbox monitoring, demonstrating that phishing can induce disclosure without a user request to authenticate. Furthermore, most evaluated mitigations that reduce leakage also impair performance on genuine tasks, exposing a security utility trade off in existing defenses. These findings show why reducing leakage alone is insufficient: effective defenses must prevent unauthorized disclosure while preserving legitimate task completion. CredLeak-Bench provides a controlled framework for measuring both objectives and evaluating progress toward secure, useful agents.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
credential leakage
phishing attacks
security-utility trade-off
autonomous task execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Credential Leakage
LLM Agents
Phishing Evaluation
Security-Utility Tradeoff
Benchmark