🤖 AI Summary
This study addresses the security risk that the legitimate web-scraping capabilities of large language models (LLMs) can be maliciously exploited to covertly exfiltrate sensitive data. To this end, it proposes LLMLeak, a novel attack vector that embeds secret URLs to induce LLMs to invoke built-in tools for information retrieval, subsequently encoding and transmitting the extracted data via DNS or web servers to establish a covert channel that circumvents network restrictions. This mechanism bypasses traditional detection paradigms by abusing legitimate functionalities to achieve data theft without requiring code generation. Experimental evaluations across eleven open-source models demonstrate an attack success rate of 79.7%, while a real-world chatbot case study further substantiates the practical severity of this threat.
📝 Abstract
With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties.
In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM's tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak.