Score
Designs and implements data serialization schemes and remote procedure call endpoints, including defining message schemas, serializers/deserializers, wire-format compatibility, and code generation for efficient binary encoding. Builds or integrates Protocol Buffers schemas and generated types and implements gRPC services and clients (stubs, servers, transport mapping, streaming, error handling, and performance/versioning considerations).
This work addresses the challenge of efficient inter-process communication and serialization of algebraic data in distributed computing environments by proposing and implementing a tunable serialization framework. The framework supports customizable serialization strategies tailored to algebraic data structures and innovatively adapts the mrdi file format for data transmission in distributed settings. By integrating domain-specific serialization mechanisms with the mrdi format, the system substantially enhances communication efficiency and processing performance for algebraic data across distributed systems. This approach provides flexible and high-performance low-level support for applications that rely heavily on structured algebraic representations, offering both adaptability and scalability without compromising on throughput or latency.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
Existing object serialization formats (e.g., Protobuf, JSON, XML) exhibit poor readability and limited auditability in source-embedded contexts such as test cases. This paper introduces ProDJ—the first pure-code serialization technique for Java—that converts runtime objects directly into syntactically valid, executable, and highly readable Java source code expressions. Its core innovation lies in leveraging the target language’s native syntax for serialization, integrating reflective introspection, abstract syntax tree (AST) generation, and cycle-aware object graph traversal to simultaneously ensure readability, executability, and maintainability. Evaluation demonstrates that ProDJ successfully serializes over 174,000 real-world objects with negligible runtime overhead. A user study confirms that developers significantly prefer ProDJ-generated Java code over JSON or XML—particularly for development tasks requiring human involvement, such as test generation.
Software engineers face significant challenges—including difficulty in modeling, lengthy prototyping cycles, and high verification costs—when developing control algorithms for complex dynamic systems such as communication networks. To address these issues, we propose GIPS, the first model-driven engineering framework that tightly integrates graph-structured integer linear programming (ILP) modeling with automated code generation. Using the domain-specific language GIPSL, users declaratively specify constraints and optimization objectives; GIPS then automatically generates functionally complete, executable Java graph-optimization components. This enables end-to-end rapid prototyping—from high-level specifications to runtime deployment. We validate GIPS on a tree-structured peer-to-peer topology control scenario, demonstrating its correctness, efficiency, and scalability. The full implementation—including source code and a ready-to-run virtual machine demonstration environment—is open-sourced, confirming its practical deployability and engineering utility.
Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.
本文提出一种统一的消息模型,用于描述异构串行数据交换协议,并通过工业工具环境实现,支持自动化开发和标准化及弱形式化串行协议的工程应用。
本文通过将Paxos协议的伪代码转化为可执行的DistAlgo语言,解决了分布式系统中复制与共识协议的理解和验证问题。
This work addresses the tension in publish-subscribe systems between standardized message formats, which ensure interoperability but incur redundant bandwidth usage, and custom messages, which improve efficiency at the cost of generality. The paper proposes Selective Field Transmission (SFT), a middleware approach that dynamically trims and transmits only the message fields required by each subscriber—without altering the standard message schema—thereby decoupling interface standardization from transmission efficiency for the first time. SFT achieves this through mechanisms including subscriber demand declaration or automatic inference, on-demand serialization, and differentiated multicast delivery. Experimental results demonstrate that bandwidth savings scale proportionally with the number and size of unused fields, while introducing no measurable per-message latency overhead.
Existing transport-layer hardware struggles to flexibly support the evolution of new protocols due to rigid protocol logic or reliance on protocol-specific assumptions. This work proposes PITA, a novel architecture that reconfigures core components—such as scheduling, packet generation, and data reassembly—through a unified event–state–instruction abstraction model, enabling a protocol-agnostic and line-rate programmable transport-layer datapath. By eliminating protocol-specific assumptions, PITA efficiently supports semantically diverse protocols, including TCP and RoCE, on a single FPGA (Alveo U250) while fully preserving their end-to-end behavioral differences. Experimental results demonstrate that the system meets timing constraints at 250 MHz with low hardware overhead and excellent performance.
This work addresses the complexity, routing ambiguity, and unreliable shutdown commonly introduced by ad hoc glue code in existing modular distributed systems. To overcome these issues, the paper proposes CNS—a lightweight, local-first hybrid event bus that seamlessly bridges local and distributed publish-subscribe contexts through a unified event model and consistent routing semantics. CNS employs an asynchronous fire-and-forget primary path while supporting request-response extensions on the same topic. It integrates typed event keys, family-based serialization and validation, and NATS-backed distributed transport. Prototype evaluation demonstrates low-latency performance: approximately 30 microseconds for local delivery, 1.26–1.37 milliseconds for purely distributed communication, and 1.64–1.89 milliseconds for hybrid bridging, with validation overhead remaining manageable—making CNS suitable for structured inter-process communication and efficient messaging among resource-constrained nodes.