Binary BPE: A Family of Cross-Platform Tokenizers for Binary Analysis

📅 2025-11-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In binary analysis, byte-level tokenization inefficiently consumes Transformer context capacity, while text-oriented tokenizers fail to handle the full binary byte range (0x00–0xFF). To address this, we propose Binary BPE—the first cross-platform, multi-architecture unified byte-pair encoding tokenizer family designed specifically for executable binaries. Trained on a large-scale corpus encompassing Linux, Windows, macOS, Android, and malware binaries, Binary BPE supports vocabulary sizes from 4K to 64K, achieving an average compression ratio of 3–8 bytes per token and improving context utilization by 2–3×. It is the first tokenizer to automatically discover interpretable structural patterns—such as file headers and instruction sequences—in ELF, PE, and Mach-O binaries without supervision. Fully compatible with neural language models and downstream binary analysis tools, Binary BPE enables efficient binary language modeling and practical static analysis. Our implementation and pre-trained tokenizers are publicly available on Hugging Face.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Hardware-aware MLComputer Vision: Language and Vision

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
Sequence models for binary analysis are bottlenecked by byte-level tokenization: raw bytes waste precious context window capacity for transformers and other neural network architectures, and many existing text-oriented tokenizers fail on arbitrary 0x00--0xFF sequences. To address this issue, we introduce the Binary BPE tokenizer family, a set of cross-platform Byte Pair Encoding (BPE) tokenizers for executables trained on a large corpus of binaries spanning multiple platforms, architectures, and operating systems, including Linux, Windows, macOS, Android, and malware sources. We release trained tokenizers with vocabularies of 4K, 8K, 16K, 32K, and 64K tokens, enabling both systematic scaling studies and practical deployment from resource-constrained edge devices to high-throughput datacenters. These tokenizers discover interpretable patterns (ELF/PE headers, instruction sequences, cross-platform strings) while yielding multi-byte compression per token. On representative uncompressed executables (e.g., ELF/PE/Mach-O rather than compressed APKs), the Binary BPE tokenizers typically allow for roughly 2-3x more binary content per fixed-length transformer context window than raw bytes, enabling more efficient research and practical deployment for content identification, malware detection, reverse engineering, and optimization. We release the trained Binary BPE tokenizers on HuggingFace, providing a drop-in, open-source foundation for binary-focused language models and context-efficient agentic tools.
Problem

Research questions and friction points this paper is trying to address.

Byte-level tokenization wastes context window capacity in binary analysis models
Existing text tokenizers fail to handle arbitrary byte sequences in executables
Lack of cross-platform tokenizers for diverse binary formats and architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Binary BPE tokenizers for executables across platforms
Trained tokenizers with scalable vocabulary sizes
Multi-byte compression enabling longer context windows
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
M. Bommarito