Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning

📅 2026-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underexplored application of self-supervised learning (SSL) and tabular representation learning (TRL) in binary program clustering. It presents the first systematic evaluation of multiple SSL methods—including BYOL, SimSiam, Barlow Twins, and VICReg—and TRL approaches such as PCA, Autoencoder, UMAP, and VIME. Building upon this analysis, the authors propose VIME-R, an innovative model that replaces conventional random perturbations with retrieval-augmented augmentation to significantly enhance representation quality. Experimental results demonstrate that VIME-R improves clustering homogeneity by 2.7%–5.8% on the Ember and Bodmas datasets, establishing a new state-of-the-art performance for VIME-based methods in this domain.
📝 Abstract
Malware clustering is a critical task in cybersecurity that helps discover threats and analyze evolving malware families. While self-supervised learning (SSL) and tabular representation learning (TRL) have achieved breakthroughs in other domains, their application to binary program clustering (the task of clustering all incoming samples regardless of label) remains largely unexplored. This study presents the first systematic investigation of SSL and TRL methods for binary program clustering, conducted in two phases on the public Ember and Bodmas datasets. In Phase 1, we establish a performance ceiling by adapting prominent vision-based SSL models (BYOL, SimSiam, Barlow Twins, VICReg) for tabular data with supervised pair generation, finding that BYOL and SimSiam achieve performance comparable to fully supervised models, while Barlow Twins and VICReg significantly underperform. In Phase 2, we evaluate purely unsupervised TRL methods against strong baselines (PCA, Autoencoder, UMAP), demonstrating that VIME establishes a new state of the art for binary program clustering. Informed by these findings, we propose VIME-R, a retrieval-augmented extension of VIME that replaces random marginal-distribution corruption with retrieval-based augmentation to generate more informative training pairs. VIME-R further improves upon VIME, achieving 2.7\%-5.8\% higher Homogeneity on both datasets. Our results highlight retrieval-augmented tabular representation learning as a promising direction for enhancing automated malware analysis. Code will be made available.
Problem

Research questions and friction points this paper is trying to address.

binary program clustering
malware clustering
self-supervised learning
tabular representation learning
cybersecurity
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised learning
tabular representation learning
binary program clustering
retrieval-augmented learning
VIME-R
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Martin Mocko
Faculty of Information Technology, Brno University of Technology, Brno, Czech Republic; Kempelen Institute of Intelligent Technologies (KInIT), Bratislava, Slovakia
Daniela Chudá
Daniela Chudá
Kempelen Institute of Intelligent Technologies
securityuser authenticationuser modellingplagiarism