Redefining Website Fingerprinting Attacks With Multiagent LLMs

📅 2025-09-15
📈 Citations: 0
Influential: 0
📄 PDF

career value

239K/year
🤖 AI Summary
Existing website fingerprinting (WFP) methods exhibit poor generalization in modern web environments—particularly single-page applications (SPAs)—largely because script-based traffic replay fails to capture the diversity and continuity of real user behavior. Method: We propose a novel, continuous traffic synthesis paradigm powered by multi-agent large language models (LLMs). Departing from page-boundary constraints, our approach employs a role-driven, collaborative multi-agent system that simulates personalized, semantically rich, and temporally coherent browsing behaviors to generate high-fidelity encrypted network traffic. Contribution/Results: This paradigm drastically reduces data synthesis cost and enhances scalability. Evaluations across 20 modern websites and traffic from 30 real users show that WFP models trained on LLM-synthesized data achieve >80% accuracy—surpassing <10% accuracy attained with script-generated data. Results demonstrate that modeling behavioral continuity and user persona is critical for improving WFP generalization.

Technology Category

Application Category

📝 Abstract
Website Fingerprinting (WFP) uses deep learning models to classify encrypted network traffic to infer visited websites. While historically effective, prior methods fail to generalize to modern web environments. Single-page applications (SPAs) eliminate the paradigm of websites as sets of discrete pages, undermining page-based classification, and traffic from scripted browsers lacks the behavioral richness seen in real user sessions. Our study reveals that users exhibit highly diverse behaviors even on the same website, producing traffic patterns that vary significantly across individuals. This behavioral entropy makes WFP a harder problem than previously assumed and highlights the need for larger, more diverse, and representative datasets to achieve robust performance. To address this, we propose a new paradigm: we drop session-boundaries in favor of contiguous traffic segments and develop a scalable data generation pipeline using large language models (LLM) agents. These multi-agent systems coordinate decision-making and browser interaction to simulate realistic, persona-driven browsing behavior at 3--5x lower cost than human collection. We evaluate nine state-of-the-art WFP models on traffic from 20 modern websites browsed by 30 real users, and compare training performance across human, scripted, and LLM-generated datasets. All models achieve under 10% accuracy when trained on scripted traffic and tested on human data. In contrast, LLM-generated traffic boosts accuracy into the 80% range, demonstrating strong generalization to real-world traces. Our findings indicate that for modern WFP, model performance is increasingly bottlenecked by data quality, and that scalable, semantically grounded synthetic traffic is essential for capturing the complexity of real user behavior.
Problem

Research questions and friction points this paper is trying to address.

Addressing website fingerprinting generalization failures in modern web environments
Overcoming behavioral entropy and data scarcity in encrypted traffic classification
Developing scalable synthetic traffic generation using multi-agent LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent LLMs simulate persona-driven browsing behavior
Scalable data generation pipeline with contiguous traffic segments
Synthetic traffic boosts model accuracy to 80% range
🔎 Similar Papers