🤖 AI Summary
In binary analysis, byte-level tokenization inefficiently consumes Transformer context capacity, while text-oriented tokenizers fail to handle the full binary byte range (0x00–0xFF). To address this, we propose Binary BPE—the first cross-platform, multi-architecture unified byte-pair encoding tokenizer family designed specifically for executable binaries. Trained on a large-scale corpus encompassing Linux, Windows, macOS, Android, and malware binaries, Binary BPE supports vocabulary sizes from 4K to 64K, achieving an average compression ratio of 3–8 bytes per token and improving context utilization by 2–3×. It is the first tokenizer to automatically discover interpretable structural patterns—such as file headers and instruction sequences—in ELF, PE, and Mach-O binaries without supervision. Fully compatible with neural language models and downstream binary analysis tools, Binary BPE enables efficient binary language modeling and practical static analysis. Our implementation and pre-trained tokenizers are publicly available on Hugging Face.
📝 Abstract
Sequence models for binary analysis are bottlenecked by byte-level tokenization: raw bytes waste precious context window capacity for transformers and other neural network architectures, and many existing text-oriented tokenizers fail on arbitrary 0x00--0xFF sequences. To address this issue, we introduce the Binary BPE tokenizer family, a set of cross-platform Byte Pair Encoding (BPE) tokenizers for executables trained on a large corpus of binaries spanning multiple platforms, architectures, and operating systems, including Linux, Windows, macOS, Android, and malware sources. We release trained tokenizers with vocabularies of 4K, 8K, 16K, 32K, and 64K tokens, enabling both systematic scaling studies and practical deployment from resource-constrained edge devices to high-throughput datacenters. These tokenizers discover interpretable patterns (ELF/PE headers, instruction sequences, cross-platform strings) while yielding multi-byte compression per token. On representative uncompressed executables (e.g., ELF/PE/Mach-O rather than compressed APKs), the Binary BPE tokenizers typically allow for roughly 2-3x more binary content per fixed-length transformer context window than raw bytes, enabling more efficient research and practical deployment for content identification, malware detection, reverse engineering, and optimization. We release the trained Binary BPE tokenizers on HuggingFace, providing a drop-in, open-source foundation for binary-focused language models and context-efficient agentic tools.