🤖 AI Summary
This study investigates security boundaries of vision-language models (VLMs) during document upload and parsing, exposing a critical gap in their handling of malicious multimodal inputs. Method: We propose a novel penetration testing methodology targeting LLM virtual workspaces by steganographically embedding the EICAR antivirus test signature into JPEG metadata—leveraging multi-layer obfuscation including JPEG steganography, Base64 encoding, string reversal, and in-sandbox Python parsing. We systematically evaluate leading VLMs—including GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet—under this attack vector. Contribution/Results: All evaluated models successfully uploaded, extracted, and potentially executed the embedded signature, demonstrating uncontrolled file parsing and execution capabilities within their virtualized environments. This work extends cloud-based AI security assessment frameworks to containerized file processing scenarios and, for the first time, reveals supply-chain–level execution risks induced by multimodal inputs. It provides reproducible empirical evidence and a methodological foundation for VLM security hardening.
📝 Abstract
This study demonstrates a novel approach to testing the security boundaries of Vision-Large Language Model (VLM/ LLM) using the EICAR test file embedded within JPEG images. We successfully executed four distinct protocols across multiple LLM platforms, including OpenAI GPT-4o, Microsoft Copilot, Google Gemini 1.5 Pro, and Anthropic Claude 3.5 Sonnet. The experiments validated that a modified JPEG containing the EICAR signature could be uploaded, manipulated, and potentially executed within LLM virtual workspaces. Key findings include: 1) consistent ability to mask the EICAR string in image metadata without detection, 2) successful extraction of the test file using Python-based manipulation within LLM environments, and 3) demonstration of multiple obfuscation techniques including base64 encoding and string reversal. This research extends Microsoft Research's"Penetration Testing Rules of Engagement"framework to evaluate cloud-based generative AI and LLM security boundaries, particularly focusing on file handling and execution capabilities within containerized environments.