Score
Designs and implements parsers and deserializers that read binary or textual bytecode and source code and produce structured representations such as abstract syntax trees, intermediate representations, and control‑flow graphs. Builds translation and frontend components that normalize multilingual code inputs into machine‑readable models for analysis, transformation, or further compilation.
Existing code datasets are predominantly confined to single-language lexical features or isolated parsers, hindering cross-lingual syntactic reasoning and structural analysis. To address this, we propose a language-agnostic, universal Abstract Syntax Tree (AST) abstraction schema that enables structural alignment and semantic normalization across ten mainstream programming languages. We introduce the first large-scale, high-fidelity multilingual code parsing dataset—comprising over 7 million source files—generated via a unified compilation pipeline and stored in Parquet format, accompanied by reproducibility scripts and interactive visualization tools. The dataset is publicly released on Hugging Face and GitHub. Empirical analysis reveals substantial syntactic structural commonalities across languages, providing foundational support for cross-lingual program understanding, pretraining of code models, and static program analysis.
Existing object serialization formats (e.g., Protobuf, JSON, XML) exhibit poor readability and limited auditability in source-embedded contexts such as test cases. This paper introduces ProDJ—the first pure-code serialization technique for Java—that converts runtime objects directly into syntactically valid, executable, and highly readable Java source code expressions. Its core innovation lies in leveraging the target language’s native syntax for serialization, integrating reflective introspection, abstract syntax tree (AST) generation, and cycle-aware object graph traversal to simultaneously ensure readability, executability, and maintainability. Evaluation demonstrates that ProDJ successfully serializes over 174,000 real-world objects with negligible runtime overhead. A user study confirms that developers significantly prefer ProDJ-generated Java code over JSON or XML—particularly for development tasks requiring human involvement, such as test generation.
This work addresses the limitations of current large language models in code translation, which rely heavily on superficial statistical patterns and lack deep program semantic understanding, particularly in real-world scenarios where high-quality semantic supervision is often unavailable. To overcome this, the authors propose Multisage, a novel framework that automatically constructs multi-dimensional structured semantic representations—such as data-flow graphs, type constraints, and API usage—from source code and generates diverse semantic augmentation signals, including natural language summaries, test cases, and API descriptions. The framework incorporates a self-calibration mechanism through semantic-preserving mutations and cross-semantic consistency verification, eliminating the need for external annotations. Evaluated on the HumanEval-X benchmark, Multisage improves translation success rates by up to 2.22× over state-of-the-art prompting, fine-tuning, and chain-of-thought approaches, with especially pronounced gains on smaller models.
In binary analysis, byte-level tokenization inefficiently consumes Transformer context capacity, while text-oriented tokenizers fail to handle the full binary byte range (0x00–0xFF). To address this, we propose Binary BPE—the first cross-platform, multi-architecture unified byte-pair encoding tokenizer family designed specifically for executable binaries. Trained on a large-scale corpus encompassing Linux, Windows, macOS, Android, and malware binaries, Binary BPE supports vocabulary sizes from 4K to 64K, achieving an average compression ratio of 3–8 bytes per token and improving context utilization by 2–3×. It is the first tokenizer to automatically discover interpretable structural patterns—such as file headers and instruction sequences—in ELF, PE, and Mach-O binaries without supervision. Fully compatible with neural language models and downstream binary analysis tools, Binary BPE enables efficient binary language modeling and practical static analysis. Our implementation and pre-trained tokenizers are publicly available on Hugging Face.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This work addresses the challenge of automatically translating APL code into C#, a task hindered by APL’s sparse syntax, scarcity of parallel corpora, and high domain-specific barriers. To overcome these limitations, the authors propose a large language model–based neural code translation framework that integrates natural language–mediated guidance, retrieval-augmented generation, and iterative refinement, complemented by a dual verification mechanism based on compilation and execution. The study introduces the first multi-level APL-to-C# equivalent code dataset and an automated functional validation evaluation pipeline, moving beyond conventional direct translation approaches. Experimental results demonstrate that the proposed method substantially improves both translation quality and functional correctness, successfully enabling accurate conversion of APL programs of varying complexity into idiomatic C#.
This study investigates whether pretrained code models encode cross-lingually consistent formal type semantics in their hidden representations. To this end, the authors construct a parallel Java–Python code dataset and employ linear probing on residual stream activations to analyze how type information is represented. They further design cross-lingual transfer experiments to assess whether models can recover formal type annotations from untyped code. This work presents the first direct interpretability analysis targeting formal type semantics and cross-lingual representation alignment in pretrained models. The results demonstrate that such models indeed learn transferable, cross-lingually aligned type structures, and that these representations exhibit robustness to lexical perturbations and syntactic discrepancies between languages.
This work addresses the challenge posed by syntactic disparities across programming languages in cross-lingual code understanding tasks. The authors propose a novel approach that integrates standardized abstract syntax tree (AST) node labels with a Graph Matching Network (GMN) to enable precise alignment of functionally equivalent code fragments within a shared semantic space. By unifying AST label normalization and GMN for the first time, the method effectively bridges the syntactic gap between languages. Experimental results demonstrate significant improvements over state-of-the-art techniques: on cross-language code clone detection, the model achieves an F1 score of 99.93%, representing a nearly 3% absolute gain, while cross-language code retrieval performance improves by 13% in Mean Reciprocal Rank (MRR).
Traditional large language models struggle to directly process raw bytes of executable files, limiting their applicability to binary understanding tasks such as malware analysis. This work proposes the first large language model natively designed for byte-level input, integrating a custom byte tokenizer, byte-level language modeling, and injection of binary-domain knowledge to enable semantic understanding and question answering over compiled code. Experimental results demonstrate that the proposed approach achieves 69% accuracy in malware family classification and 98% accuracy in architecture classification, substantially outperforming general-purpose large language models. These findings underscore the effectiveness and necessity of native byte-level modeling combined with domain-specific knowledge for advancing binary analysis capabilities.