🤖 AI Summary
This study addresses the limitations of conventional phishing detection methods that rely on easily spoofed URL or HTML features while neglecting visual deception as the core mechanism of such attacks. To overcome this, we propose a purely visual detection paradigm based on YOLOv8 that directly analyzes webpage screenshots—specifically their layout structures and color schemes—to identify phishing sites. Furthermore, targeted image data augmentation techniques are incorporated to enhance model robustness. Experimental results demonstrate that the proposed approach achieves a classification accuracy of 92% with an inference latency of approximately 100 milliseconds per image, while reducing the false positive rate by 11%. This work presents an efficient and resilient alternative for phishing detection that is inherently resistant to adversarial obfuscation at the code level.
📝 Abstract
Phishing remains one of the most common vectors for financial and identity fraud, and most detection systems still rely on inspecting a page's URL, HTML markup, or domain registration history. These signals are easy for an attacker to rotate or obfuscate, and they say very little about what actually convinces a victim to hand over a password or a card number: the way the page looks. This paper describes a visual, image-based approach to phishing detection that treats a rendered webpage the same way a human eye would, as a picture that either matches a trusted brand or doesn't. A YOLOv8 convolutional neural network was trained to classify full-page website screenshots as phishing or legitimate based on layout, logo placement, color scheme, and login-form structure, rather than on text extracted from the page. The system reached 92% classification accuracy on a held-out test set, processed a single screenshot in roughly 100 milliseconds, and, after a round of data augmentation aimed specifically at lighting, compression, and scaling variation, cut the false-positive rate by 11% relative to the pre-augmentation baseline. The paper walks through the dataset construction, the augmentation strategy, the model architecture and training setup, and the resulting performance, and closes with a discussion of where this kind of visual detector fits alongside, rather than instead of, existing URL- and content-based defenses.