Score
Designs and implements methods and tools that compose and integrate global constraints across models, programs, or dataflows while applying semantic‑preserving transformations. Builds and analyzes composition operators, transformation pipelines, and integration adapters, verifying semantic equivalence, consistency, and the resulting effects on system behavior and constraints.
Existing conformance checking approaches between process models and reference models suffer from limited semantic expressiveness and insufficient automation, hindering fine-grained compliance verification. This paper proposes a semantic consistency checking method grounded in causal dependency analysis of tasks and events, transcending traditional trajectory-based dependency modeling by formally encoding causal constraints at the semantic level. We establish a unified framework integrating causal dependency modeling, semantic representation, and formal verification, and design an automated conformance checking algorithm implemented in a prototype tool. Empirical evaluation demonstrates that our approach significantly outperforms state-of-the-art techniques in both accuracy and flexibility, achieving— for the first time—the fully automated, high-expressivity semantic conformance verification of process models against reference models.
Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.
This work addresses the challenge that large language models (LLMs) struggle to adhere to formal semantic constraints in real time when generating structured knowledge, often relying on inefficient and error-prone post-hoc validation. To overcome this limitation, the authors propose an ontology-to-tool compilation mechanism that automatically translates domain ontology specifications into executable tool interfaces. By compelling LLM agents to interact with knowledge graphs exclusively through these generated tools, the approach proactively enforces semantic consistency during knowledge generation. Built upon The World Avatar framework, the method integrates the Model Context Protocol, ontology-driven tool synthesis, and agent workflows, substantially reducing the need for manual prompt engineering. Evaluated on the task of processing scientific literature on metal–organic polyhedra synthesis, the system successfully guides LLMs to extract, validate, and repair structured knowledge, demonstrating the feasibility and advantages of this paradigm for scientific text understanding.
Current tool-augmented large language model (LLM) ecosystems suffer from fragmentation—characterized by coexisting heterogeneous protocols (e.g., OpenAI Function Calling, Toolformer), manual schema definition, and complex execution orchestration—leading to low development efficiency and high integration overhead. To address this, we propose a protocol-agnostic unified tool integration framework. Our approach introduces an abstract protocol layer for cross-standard compatibility, an automated schema inference mechanism to eliminate manual specification, and a dual-mode concurrent scheduler enabling seamless synchronous and asynchronous tool execution. Experimental evaluation demonstrates that, compared to baseline approaches, our framework reduces implementation code volume by 60–80%, achieves up to 3.1× improvement in end-to-end execution latency, and maintains full backward compatibility with mainstream LLM tool-calling ecosystems.
AI-augmented Data Processing Systems (DPS) suffer from low trustworthiness in critical applications due to the inherent unreliability of large language model (LLM) outputs, while existing constraint mechanisms are fragmented, imperative, and lack semantic-aware query execution support. Method: We propose Semantic Integrity Constraints (SICs)—a declarative abstraction embedded within relational models—extending classical database integrity constraints to the semantic layer for the first time. We introduce novel constraint classes (e.g., groundedness), unify active constraint decoding with passive verification-and-recovery execution, and extend relational algebra to enable query-aware optimization and execution. Contribution/Results: Our framework significantly improves DPS trustworthiness and performance, supports enterprise-scale deployment, and establishes foundational theory and systems infrastructure for constraint-driven semantic query optimization and adaptive execution.
This work addresses the challenge of irreproducibility in data analysis scripts, which often stems from implicit assumptions—such as specific package versions, expected data formats, or undocumented manual interventions. The paper proposes a static analysis approach tailored to data analysis workflows that, for the first time, unifies diverse implicit assumptions into inferable constraint models. By leveraging customized program analysis and example-driven modeling, the authors develop a prototype system capable of automatically identifying these hidden assumptions, extracting executable preconditions, and generating verifiable constraints. The resulting framework supports runtime validation and automatic documentation generation, substantially enhancing script executability, reproducibility, and interpretability.
This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.
This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.