🤖 AI Summary
To address the challenge of dynamic traffic monitoring and detecting unknown anomalies—such as zero-day attacks and network congestion—in cloud environments, this paper introduces the first approach leveraging pretrained large language models (LLMs) for network traffic time-series modeling. Methodologically, we propose a hybrid architecture integrating Transformer-based attention mechanisms with an autoencoder to jointly capture temporal dependencies and semantic context in traffic sequences. A lightweight transfer learning mechanism is further designed to ensure adaptability to unseen network topologies and adversarial scenarios. Experimental results demonstrate that our system significantly outperforms conventional methods in both accuracy and inference efficiency: zero-day attack detection rate improves by 23.6%, false positive rate decreases by 41.2%, and fine-grained congestion patterns are effectively identified. This work establishes a novel paradigm for LLM-driven network anomaly detection, bridging the gap between natural language processing and network traffic analysis.
📝 Abstract
The rapidly evolving cloud platforms and the escalating complexity of network traffic demand proper network traffic monitoring and anomaly detection to ensure network security and performance. This paper introduces a large language model (LLM)-based network traffic monitoring and anomaly detection system. In addition to existing models such as autoencoders and decision trees, we harness the power of large language models for processing sequence data from network traffic, which allows us a better capture of underlying complex patterns, as well as slight fluctuations in the dataset. We show for a given detection task, the need for a hybrid model that incorporates the attention mechanism of the transformer architecture into a supervised learning framework in order to achieve better accuracy. A pre-trained large language model analyzes and predicts the probable network traffic, and an anomaly detection layer that considers temporality and context is added. Moreover, we present a novel transfer learning-based methodology to enhance the model's effectiveness to quickly adapt to unknown network structures and adversarial conditions without requiring extensive labeled datasets. Actual results show that the designed model outperforms traditional methods in detection accuracy and computational efficiency, effectively identify various network anomalies such as zero-day attacks and traffic congestion pattern, and significantly reduce the false positive rate.