parse binary files

Designs and implements parsers and serializers that read, interpret, and write structured binary file formats, including deserializing fields, preserving section ordering and layout, and handling multiple file formats. Builds in-memory and on-disk representations and export pipelines that encode file contents into interoperable or compact feature representations for analysis, transformation, or consumption by other tools.

parsebinaryfiles

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Serializing Java Objects in Plain Code

May 18, 2024
JW
Julian Wachter
🏛️ Karlsruhe Institute of Technology | KTH Royal Institute of Technology

Existing object serialization formats (e.g., Protobuf, JSON, XML) exhibit poor readability and limited auditability in source-embedded contexts such as test cases. This paper introduces ProDJ—the first pure-code serialization technique for Java—that converts runtime objects directly into syntactically valid, executable, and highly readable Java source code expressions. Its core innovation lies in leveraging the target language’s native syntax for serialization, integrating reflective introspection, abstract syntax tree (AST) generation, and cycle-aware object graph traversal to simultaneously ensure readability, executability, and maintainability. Evaluation demonstrates that ProDJ successfully serializes over 174,000 real-world objects with negligible runtime overhead. A user study confirms that developers significantly prefer ProDJ-generated Java code over JSON or XML—particularly for development tasks requiring human involvement, such as test generation.

Enables object reconstruction in source codeImproves serialization readability in JavaReplaces binary formats with plain code

This work addresses the challenge of efficient inter-process communication and serialization of algebraic data in distributed computing environments by proposing and implementing a tunable serialization framework. The framework supports customizable serialization strategies tailored to algebraic data structures and innovatively adapts the mrdi file format for data transmission in distributed settings. By integrating domain-specific serialization mechanisms with the mrdi format, the system substantially enhances communication efficiency and processing performance for algebraic data across distributed systems. This approach provides flexible and high-performance low-level support for applications that rely heavily on structured algebraic representations, offering both adaptability and scalability without compromising on throughput or latency.

Algebraic DataDistributed ComputingInterprocess Communication

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Binsparse: A Specification for Cross-Platform Storage of Sparse Matrices and Tensors

Jun 23, 2025
BB
Benjamin Brock
🏛️ Intel Corporation | Massachusetts Institute of Technology | Quansight | Texas A&M University | Anaconda | Chan Zuckerberg Initiative | NVIDIA

Existing sparse matrix/tensor formats (e.g., Matrix Market, FROSTT) rely on ASCII text, resulting in storage redundancy and inefficient parsing—bottlenecks for sparse computation performance. This paper introduces Binsparse: the first general-purpose, modular, embeddable, cross-platform binary sparse data specification. It synergistically combines JSON metadata (encoding dimensions, data types, and layout semantics) with native binary arrays (supporting CSR, CSC, COO, etc.) and natively interoperates with modern container formats including HDF5, Zarr, and NPZ. On SuiteSparse and FROSTT benchmarks, an HDF5-based CSR implementation achieves, without compression, 2.4× smaller file size, 26.5× faster read throughput, and 31× faster write throughput versus ASCII formats; with compression, it attains 7.5× smaller size, 2.6× faster reads, and 1.4× faster writes. The open-source project provides reference implementations in multiple programming languages, bridging the gap between computational efficiency and cross-platform portability.

High cost of loading/storing matrices exceeds computation costInefficient text storage increases file size and parsing timeLack of cross-platform binary sparse matrix storage format

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic understanding regarding the usability of textual serialization formats such as JSON and XML, particularly concerning the factors that influence cognitive efficiency and user experience. Through a large-scale crowdsourced experiment (N=215) complemented by semi-structured interviews (N=9), the authors conduct a mixed-methods evaluation of multiple formats in realistic editing tasks. While HJSON and YAML demonstrate marginal advantages in specific modification scenarios, these benefits vanish in both simpler and more complex contexts. Crucially, the findings reveal that usability is not primarily determined by syntactic differences but rather by socio-technical ecosystem factors—including tooling support, documentation quality, and community practices. This work provides the first empirical evidence that ecosystem support exerts a more decisive influence on format usability than syntax design alone.

cognitive efficiencydata serializationsociotechnical ecosystems

Existing decompilation evaluations predominantly rely on syntactic similarity or isolated readability metrics, which inadequately capture the practical reusability of recovered code. To address this limitation, this work proposes a three-dimensional evaluation paradigm centered on reusability—encompassing readability, recompilability, and functionality—and introduces DEBENCH, the first automated multidimensional benchmark comprising 240 atomic functions and 640 binary samples. Leveraging LLM-as-judge for readability scoring, URAF fine-grained metrics, 50-round iterative compilation repair, and Frida-driven multilevel dynamic differential tracing, the study systematically uncovers significant discrepancies across evaluation dimensions: only 1.2% of outputs from the best decompiler–LLM combination achieve full functional equivalence; Clang-generated code exhibits 2.6× higher functionality than GCC’s; and functional recovery capability varies by up to 20× across decompilers. The analysis further identifies three dominant failure modes, including type system collapse.

binary decompilationevaluationfunctionality

This work addresses the challenge of reliable source-level binary patching in the absence of original source code and toolchains, where existing decompilers often produce outputs riddled with syntactic and semantic errors. To overcome this limitation, the authors propose a static patching framework that integrates decompilation with binary-aware recompilation. By leveraging information extracted directly from the original binary, the framework corrects semantic distortions in decompiled code and enables automated patch generation. The approach substantially improves recompilation correctness, fixing approximately 81% of erroneous functions produced by Hex-Rays, successfully patching 13 out of 14 real-world CVEs, and increasing user experiment success rates from 3.7% to 100%. Furthermore, it supports large-model-driven fully automated patching, demonstrating robust practical applicability.

binary patchingdecompilationrecompilation

Developing cross-platform graphical user interfaces (GUIs) and plugins for command-line tools in structural bioinformatics is often costly and complex. This work proposes a three-stage automated workflow that leverages a platform-agnostic formal GUI specification, decouples model, view, and presenter components through the Model–View–Presenter (MVP) architectural pattern, and employs a dedicated code generator to automatically produce plugins for target platforms—namely VMD, PyMOL, and the web. To the best of our knowledge, this is the first systematic application of the MVP pattern to the automatic GUI generation for CLI tools, substantially enhancing logic reusability, cross-platform portability, and development efficiency. The framework’s generality, extensibility, and practical utility are demonstrated by successfully generating plugins for multiple tools from the Structural Bioinformatics Library across all three platforms.

CLI automationcross-platformGUI generation

This work addresses the unreliability of disassembly caused by the absence of compiler-intended semantic information in stripped binary executables. To overcome this limitation, the authors propose a novel lightweight metadata embedding mechanism that explicitly encodes critical semantics—such as code regions and memory boundaries—directly into the binary. This approach yields a decidable intermediate representation situated between raw binaries and source code. For the first time, it enables disassembly that is both decidable and recompilable, facilitating precise lifting to high-level intermediate representations. Experimental evaluation demonstrates that the embedded metadata incurs only 17% of the size overhead of DWARF debug information, introduces no runtime performance penalty, and successfully supports behavior-preserving binary lifting, instrumentation, and recompilation across a wide range of real-world C/C++ programs.

binary formatcompilation metadatadisassembly

Hot Scholars

SB

Sebastian Baltes

University of Bayreuth
software engineeringempirical software engineering
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
AR

Andreas Rausch

Full Professor for Software Systems Engineering, Institute for Software & Systems Engineering, TU
Software Systems EngineeringRequirements Engineering and Software ArchitectureDesign and ModelingEngineering Processes
DB

Dominique Briechle

Institute for Software and Systems Engineering (ISSE), Clausthal University of Technology
AC

Andrew Case

Volexity, Volatility Foundation, Louisiana State University, University of New Orleans
Computer ForensicsMemory ForensicsMalware Analysis