Score
Designs and implements parsers and validators that extract, canonicalize, and sanitize structured fields from model or text outputs, map those parsed outputs to executable representations (for example queries or API payloads), validate them against schemas and constraints, and detect or reject invalid or unsafe outputs.
Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.
Parsing binary data formats (e.g., CBOR, CDDL, COSE) in low-level languages is error-prone and frequently introduces security vulnerabilities. Method: We propose PulseParse, a verifiable parsing and serialization framework based on separation logic. Our approach formally verifies the non-malleability of deterministic CBOR fragments; introduces well-formedness conditions for CDDL and automatically synthesizes verified codecs; delivers the first fully formalized, end-to-end COSE signature protocol; and employs constant-stack-space recursive parsing with robust handling of adversarial inputs. Contributions: We release EverCBOR—a machine-checked, correctness-proven CBOR library—and EverCDDL—a tool for CDDL validation and verified code generation. These constitute the first fully formalized, end-to-end implementations of industrial standards including DICE and COSE. PulseParse supports verified code generation for both C and Rust, enabling high-assurance, cross-language interoperability in safety-critical systems.
The absence of explicit logical schemas in NoSQL applications hinders maintenance and query optimization. Method: This paper proposes a model-driven reverse engineering approach based on static code analysis. It defines an object-oriented language metamodel and a unified schema metamodel (uSchema), and integrates control-flow-driven data access modeling with structural inference to automatically extract logical schemas for NoSQL applications. The approach further supports join-pattern identification and automated refactoring of field redundancy. Contribution/Results: To our knowledge, this is the first domain-specific reverse engineering pipeline for NoSQL, enabling platform-independent model transformations. End-to-end evaluation on MongoDB applications demonstrates high schema extraction accuracy, generation of executable refactored code, elimination of expensive join operations, and significant improvements in both query performance and maintainability.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.
This work addresses the challenge of repairing minimally corrupted structured input files—such as JSON or INI—that frequently fail to parse due to minor syntactic damage, a problem for which existing repair methods often sacrifice content fidelity or introduce semantic errors. To this end, the paper introduces the first Transformer-based approach for this task, proposing a format-aware supervised sequence generation framework. The method integrates oracle-driven validation with a boundary-localized repair mechanism to concentrate generation efforts on faulty regions, thereby producing high-fidelity repairs. Empirical results demonstrate that the approach achieves a 97.57% repair success rate while recovering 94.29% of original content, and it significantly enhances efficiency on long files, operating five times faster than current state-of-the-art techniques.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.
This work addresses the lack of semantic guarantees in existing ladder diagram verification tools, which often leads to false negatives or false positives in safety violation detection due to imprecise translation into model checker inputs. To remedy this, the authors present the first K Framework–based, standards-faithful, and reusable executable formal semantics for IEC 61131-3 ladder diagrams. This semantics uniformly yields both an interpreter and a deductive verifier, serving as an independent audit benchmark for differential testing of translation processes. It accurately models contacts, coils, timers, counters, and retentive scan cycles, with machine-checked correctness verified via kprove. Applying this approach uncovered two real-world flaws in ESBMC: unsound certification of unsafe programs and generation of spurious counterexamples. Furthermore, it formally guarantees input/output behavioral equivalence with the standard for both combinational and latching logic.
This work addresses the limited reliability of structured security artifacts—such as KQL queries and MITRE ATT&CK mappings—generated by large language models, which often fall short of production-grade requirements. To bridge this gap, the authors propose a lightweight verification framework that shifts the focus of quality assurance from generation to validation. The core innovations include a hybrid testing strategy integrating test-driven generation, deterministic program verification, and semantic evaluation by large language models, alongside an interpretable judging mechanism distilled from expert decision distributions. Deployed in Microsoft Sentinel’s production environment, the framework significantly enhances the reliability of three critical types of security artifacts, establishing professional-grade and scalable validation standards.