Impact of Latent Space Dimension on IoT Botnet Detection Performance: VAE-Encoder Versus ViT-Encoder

๐Ÿ“… 2024-03-01
๐Ÿ›๏ธ 2024 3rd International Conference for Innovation in Technology (INOCON)
๐Ÿ“ˆ Citations: 1
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study systematically investigates the impact of latent-space dimensionality on IoT botnet detection performance, comparing Vision Transformers (ViTs) and Variational Autoencoder (VAE) encoders for dimensionality reduction of structured network traffic data (CSV format). Addressing the absence of spatial locality and hierarchical structure in IoT trafficโ€”key inductive biases assumed by ViTsโ€”we reveal, for the first time, that ViTs underperform significantly relative to VAEs due to model-structure mismatch. Under a unified framework, we couple ViT and VAE encoders with downstream classifiers (MLP, LSTM) and evaluate them on the N-BaIoT and CICIoT2022 datasets. Results demonstrate that VAE consistently outperforms ViT across all tested latent dimensions, achieving average improvements of 5.2%โ€“13.8% in four core metrics: accuracy, precision, recall, and F1-score. This confirms VAEโ€™s superior suitability for unsupervised representation learning on non-image sequential traffic data.

Technology Category

Computer Vision: Representation Learning for VisionMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Safety and Robustness

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
๐Ÿ“ Abstract
The rapid evolution of Internet of Things (IoT) technology has led to a significant increase in the number of IoT devices, applications, and services. This surge in IoT devices, along with their widespread presence, has made them a prime target for various cyber-attacks, particularly through IoT botnets. As a result, security has become a major concern within the IoT ecosystem. This study focuses on investigating how the latent dimension impacts the performance of different deep learning classifiers when trained on latent vector representations of the train dataset. The primary objective is to compare the outcomes of these models when encoder components from two cutting-edge architectures: the Vision Transformer (ViT) and the Variational Auto-Encoder (VAE) are utilized to project the high dimensional train dataset to the learned low dimensional latent space. The encoder components are employed to project high-dimensional structured .csv IoT botnet traffic datasets to various latent sizes. Evaluated on N-BaIoT and CICIoT2022 datasets, findings reveal that VAE-encoder based dimension reduction outperforms ViT-encoder based dimension reduction for both datasets in terms of four performance metrics including accuracy, precision, recall, and F1-score for all models which can be attributed to absence of spatial patterns in the datasets the ViT model attempts to learn and extract from image instances.
Problem

Research questions and friction points this paper is trying to address.

Investigates latent dimension impact on IoT botnet detection performance
Compares VAE-Encoder and ViT-Encoder for dimension reduction
Evaluates performance metrics on N-BaIoT and CICIoT2022 datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

VAE-encoder reduces IoT botnet traffic dimensions
Compares VAE and ViT encoder performances
VAE outperforms ViT in detection metrics
๐Ÿ”Ž Similar Papers
No similar papers found.