MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AI代理报告漏洞过快且需大量人力处理的问题,本文提出MobileCybench框架,通过可执行探针评估漏洞报告,提高检测效率。
📝 Abstract
AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application, and require considerable human labor to process. To mitigate this, we introduce a framework for evaluating vulnerability reports via probes, executable checks of security properties. A reported exploit is evaluated by replaying it against the application and running the probes: a triggered probe indicates both that the exploit succeeded and which security property it violated. As a probe encodes a security property rather than a known vulnerability, it can detect vulnerabilities that were not known when the probe was written. We instantiate the framework as MobileCybench, a benchmark for vulnerability discovery by AI agents in 13 Android applications, with 495 probes written and reviewed by the authors. We evaluate 5 coding agents (OpenCode with GPT-5.5, GPT-5.6-Sol, and GLM-5.2; Claude Code with Opus 4.8 and Opus 5) under 4 settings: as a malicious app on the victim's device or as a remote attacker with a low-privilege account, each with either only an obfuscated APK or access to the application's source code. Given only the obfuscated APK, the top agent, OpenCode with GPT-5.6-Sol, triggers probes in 53.8% of applications in the malicious-app setting and 16.7% in the remote-attacker setting. With source code, the trigger rate across all agents and both attack settings increases from 28.8% to 32.8%. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, the majority of which have been confirmed by maintainers.
Problem

Research questions and friction points this paper is trying to address.

vulnerability reports
AI agents
security properties
human labor
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable probes
vulnerability discovery
security properties
AI agents
benchmark
💼 Related Jobs
No related jobs found.
A
Andy K. Zhang
Stanford University
A
Ava Huang
Stanford University
J
Joey Ji
Stanford University
W
Wai Han
Stanford University
T
Thomas Qin
Stanford University
N
Nardos Demilew
Stanford University
M
Michael Tian-Yue Liu
UC Berkeley
B
Brian Song
UC Berkeley
R
Riya Dulepet
Stanford University
B
Brian Wang
UC Berkeley
K
Kyleen Liao
Stanford University
C
Cuiyuanxiu Chen
Stanford University
N
Nishka Kacheria
Stanford University
A
Andrew Wu
Stanford University
P
Pratham Rangwala
UC Berkeley
X
Xinjie Wang
UC Berkeley
L
Laura Gomezjurado Gonzalez
Stanford University
A
Anita Ding
UC Berkeley
B
Benjamin Yi
UC Berkeley
Daniel E. Ho
Daniel E. Ho
Stanford University
Regulatory policyartificial intelligenceadministrative lawantidiscrimination
Dan Boneh
Dan Boneh
Professor of Computer Science, Stanford University
CryptographyComputer SecurityComputer Science Theory
Dawn Song
Dawn Song
Professor of Computer Science, UC Berkeley
Computer Security and Privacy
Ion Stoica
Ion Stoica
Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
Percy Liang
Percy Liang
Associate Professor of Computer Science, Stanford University
machine learningnatural language processing