Score
Designs and implements parsers and serializers that read, interpret, and write structured binary file formats, including deserializing fields, preserving section ordering and layout, and handling multiple file formats. Builds in-memory and on-disk representations and export pipelines that encode file contents into interoperable or compact feature representations for analysis, transformation, or consumption by other tools.
Existing object serialization formats (e.g., Protobuf, JSON, XML) exhibit poor readability and limited auditability in source-embedded contexts such as test cases. This paper introduces ProDJ—the first pure-code serialization technique for Java—that converts runtime objects directly into syntactically valid, executable, and highly readable Java source code expressions. Its core innovation lies in leveraging the target language’s native syntax for serialization, integrating reflective introspection, abstract syntax tree (AST) generation, and cycle-aware object graph traversal to simultaneously ensure readability, executability, and maintainability. Evaluation demonstrates that ProDJ successfully serializes over 174,000 real-world objects with negligible runtime overhead. A user study confirms that developers significantly prefer ProDJ-generated Java code over JSON or XML—particularly for development tasks requiring human involvement, such as test generation.
This work addresses the challenge of efficient inter-process communication and serialization of algebraic data in distributed computing environments by proposing and implementing a tunable serialization framework. The framework supports customizable serialization strategies tailored to algebraic data structures and innovatively adapts the mrdi file format for data transmission in distributed settings. By integrating domain-specific serialization mechanisms with the mrdi format, the system substantially enhances communication efficiency and processing performance for algebraic data across distributed systems. This approach provides flexible and high-performance low-level support for applications that rely heavily on structured algebraic representations, offering both adaptability and scalability without compromising on throughput or latency.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
Existing sparse matrix/tensor formats (e.g., Matrix Market, FROSTT) rely on ASCII text, resulting in storage redundancy and inefficient parsing—bottlenecks for sparse computation performance. This paper introduces Binsparse: the first general-purpose, modular, embeddable, cross-platform binary sparse data specification. It synergistically combines JSON metadata (encoding dimensions, data types, and layout semantics) with native binary arrays (supporting CSR, CSC, COO, etc.) and natively interoperates with modern container formats including HDF5, Zarr, and NPZ. On SuiteSparse and FROSTT benchmarks, an HDF5-based CSR implementation achieves, without compression, 2.4× smaller file size, 26.5× faster read throughput, and 31× faster write throughput versus ASCII formats; with compression, it attains 7.5× smaller size, 2.6× faster reads, and 1.4× faster writes. The open-source project provides reference implementations in multiple programming languages, bridging the gap between computational efficiency and cross-platform portability.
This study addresses the lack of systematic understanding regarding the usability of textual serialization formats such as JSON and XML, particularly concerning the factors that influence cognitive efficiency and user experience. Through a large-scale crowdsourced experiment (N=215) complemented by semi-structured interviews (N=9), the authors conduct a mixed-methods evaluation of multiple formats in realistic editing tasks. While HJSON and YAML demonstrate marginal advantages in specific modification scenarios, these benefits vanish in both simpler and more complex contexts. Crucially, the findings reveal that usability is not primarily determined by syntactic differences but rather by socio-technical ecosystem factors—including tooling support, documentation quality, and community practices. This work provides the first empirical evidence that ecosystem support exerts a more decisive influence on format usability than syntax design alone.
Existing decompilation evaluations predominantly rely on syntactic similarity or isolated readability metrics, which inadequately capture the practical reusability of recovered code. To address this limitation, this work proposes a three-dimensional evaluation paradigm centered on reusability—encompassing readability, recompilability, and functionality—and introduces DEBENCH, the first automated multidimensional benchmark comprising 240 atomic functions and 640 binary samples. Leveraging LLM-as-judge for readability scoring, URAF fine-grained metrics, 50-round iterative compilation repair, and Frida-driven multilevel dynamic differential tracing, the study systematically uncovers significant discrepancies across evaluation dimensions: only 1.2% of outputs from the best decompiler–LLM combination achieve full functional equivalence; Clang-generated code exhibits 2.6× higher functionality than GCC’s; and functional recovery capability varies by up to 20× across decompilers. The analysis further identifies three dominant failure modes, including type system collapse.
This work addresses the challenge of reliable source-level binary patching in the absence of original source code and toolchains, where existing decompilers often produce outputs riddled with syntactic and semantic errors. To overcome this limitation, the authors propose a static patching framework that integrates decompilation with binary-aware recompilation. By leveraging information extracted directly from the original binary, the framework corrects semantic distortions in decompiled code and enables automated patch generation. The approach substantially improves recompilation correctness, fixing approximately 81% of erroneous functions produced by Hex-Rays, successfully patching 13 out of 14 real-world CVEs, and increasing user experiment success rates from 3.7% to 100%. Furthermore, it supports large-model-driven fully automated patching, demonstrating robust practical applicability.
Developing cross-platform graphical user interfaces (GUIs) and plugins for command-line tools in structural bioinformatics is often costly and complex. This work proposes a three-stage automated workflow that leverages a platform-agnostic formal GUI specification, decouples model, view, and presenter components through the Model–View–Presenter (MVP) architectural pattern, and employs a dedicated code generator to automatically produce plugins for target platforms—namely VMD, PyMOL, and the web. To the best of our knowledge, this is the first systematic application of the MVP pattern to the automatic GUI generation for CLI tools, substantially enhancing logic reusability, cross-platform portability, and development efficiency. The framework’s generality, extensibility, and practical utility are demonstrated by successfully generating plugins for multiple tools from the Structural Bioinformatics Library across all three platforms.
This work addresses the unreliability of disassembly caused by the absence of compiler-intended semantic information in stripped binary executables. To overcome this limitation, the authors propose a novel lightweight metadata embedding mechanism that explicitly encodes critical semantics—such as code regions and memory boundaries—directly into the binary. This approach yields a decidable intermediate representation situated between raw binaries and source code. For the first time, it enables disassembly that is both decidable and recompilable, facilitating precise lifting to high-level intermediate representations. Experimental evaluation demonstrates that the embedded metadata incurs only 17% of the size overhead of DWARF debug information, introduces no runtime performance penalty, and successfully supports behavior-preserving binary lifting, instrumentation, and recompilation across a wide range of real-world C/C++ programs.