🤖 AI Summary
This study addresses critical challenges in IoT device identification, where poor method selection, inadequate handling of data heterogeneity, suboptimal feature extraction, and misleading evaluation metrics often undermine model generalization and reproducibility. For the first time, this work systematically identifies common pitfalls in the field, elucidates the trade-offs between unique and categorical device identifiers, and exposes misconceptions in practices such as data augmentation and session modeling. Building on these insights, the paper proposes a comprehensive best-practice framework spanning methodology design, data processing, and multi-dimensional evaluation. This framework substantially enhances model robustness, reproducibility, and deployment reliability, offering actionable guidance for developing high-performance, generalizable IoT security identification systems.
📝 Abstract
This paper critically examines the device identification process using machine learning, addressing common pitfalls in existing literature. We analyze the trade-offs between identification methods (unique vs. class based), data heterogeneity, feature extraction challenges, and evaluation metrics. By highlighting specific errors, such as improper data augmentation and misleading session identifiers, we provide a robust guideline for researchers to enhance the reproducibility and generalizability of IoT security models.