🤖 AI Summary
This work addresses the vulnerability of large language model (LLM)-driven natural language database interfaces to SQL injection attacks, where adversarial prompts can induce the generation of malicious queries. To counter this threat, the paper presents the first systematic defense framework tailored to the LLM-to-SQL translation pipeline, integrating prompt sanitization, semantic and behavioral anomaly detection, attack signature matching, and model fine-tuning into an end-to-end collaborative protection architecture. Evaluated across diverse adversarial scenarios, the proposed approach achieves high detection accuracy with low false-positive rates, substantially enhancing the security of LLM-based database applications. The framework demonstrates both technical innovation and practical applicability in safeguarding natural language interfaces against prompt-based SQL injection exploits.
📝 Abstract
Natural language interfaces to structured databases are becoming increasingly common, largely due to advances in large language models (LLMs) that enable users to query data using conversational input rather than formal query languages such as SQL. While this paradigm significantly improves usability and accessibility, it introduces new security risks, particularly the amplification of SQL injection vulnerabilities through the prompt-to-SQL translation process. Malicious users can exploit these mechanisms by crafting adversarial prompts that manipulate model behavior and generate unsafe queries. In this work, we propose a multi-layered security framework designed to detect and mitigate LLM-mediated SQL injection attacks. The framework integrates a front-end security shield for prompt sanitization, an advanced threat detection model for behavioral and semantic anomaly identification, and a signature-based control layer for known attack patterns. We evaluate the proposed framework under diverse and realistic attack scenarios, including prompt injection, obfuscated SQL payloads, and context-manipulation attacks. To ensure robustness, we generate and curate a comprehensive benchmark dataset of adversarial prompts and assess performance across a fine-tuned LLM configuration. Experimental results demonstrate that the proposed approach achieves high detection accuracy while maintaining low false-positive rates, significantly improving the secure deployment of LLM-powered database applications.