🤖 AI Summary
This study investigates how privacy guidance influences the information source selection of large language model (LLM) agents. Addressing the insufficient granularity of existing system-level instructions, we propose a hybrid guidance mechanism that integrates system-level and skill-level privacy annotations, alongside an evaluation framework incorporating fine-grained privacy metadata. Through synthetic task evaluations and multi-model comparative experiments, we demonstrate that this hybrid mechanism reduces confidential data access rates by 50%, effectively constraining agents' information-seeking behaviors. Our findings confirm the critical role of skill-level privacy annotation in the safety alignment of LLM agents, offering a novel paradigm for constructing privacy-aware agent systems.
📝 Abstract
While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user. The evaluation framework comprises 55 synthetic tasks spanning 11 categories of personal information, with 169 associated skills that describe the available acquisition pathways. We consider privacy guidance through system-level instructions, skill-level metadata labels, or both. Separately, we vary user availability and urgency framing. With users available and no privacy guidance, agents access confidential sources in 30% of valid runs on average across five open-weight models, despite sufficient alternatives. This rate increases to 45% when users are unavailable, whereas urgency framing has no detectable effect. System-level privacy instructions alone have limited effects on confidential access, while skill-level intrusiveness labels produce a modest reduction (24% on average), but combining the two roughly halves confidential access. Our findings motivate incorporating privacy annotations into skill specifications and evaluating their effectiveness alongside system-level instructions.