Score
Designs, specifies, and documents data and file/storage formats and their logical and physical layouts (serialization, metadata/schema, encoding, compression and indexing), and builds or analyzes tools and transformations that optimize layout and representation for performance, space, compatibility and evolution; selects, governs, and manages format choices across systems and workloads to meet interoperability, versioning, and storage/access tradeoffs.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This study addresses the lack of systematic guidance on contextualization strategies for large language model (LLM) agents operating in structured data environments, particularly concerning effectiveness and efficiency across multi-file, large-scale schemas. Using SQL generation as a proxy task, the work presents the first systematic evaluation of eleven models across four context formats—YAML, Markdown, JSON, and TOON—at schema scales ranging from 10 to 10,000 tables. The findings reveal that model capability tiers critically determine optimal context architecture: tailored strategies significantly improve performance, with state-of-the-art models gaining 2.7% accuracy under native file-based contexts, while open-source models average a 7.7% decline. Moreover, native file-based agents scale efficiently to ten-thousand-table schemas while maintaining high navigation accuracy.
To address insufficient semantic description of multidimensional aggregate/summary data, poor adaptability of metadata standards, and cross-source interoperability challenges in big data environments, this paper proposes a multidimensional data source profiling metadata model tailored for data ecosystems. Built upon RDF, the model is the first to support extensible semantic modeling of both aggregate and summary multidimensional data, enabling semantic alignment of dimensions and measures with reference knowledge graphs. It integrates multi-granularity metadata profiles—spanning source-level, attribute-level, and value-distribution characteristics. The model ensures flexible extensibility and cross-source interoperability. Experimental results demonstrate that profile generation time scales linearly with data cardinality, confirming its engineering practicality and predictable performance.
This study addresses the lack of systematic understanding regarding the usability of textual serialization formats such as JSON and XML, particularly concerning the factors that influence cognitive efficiency and user experience. Through a large-scale crowdsourced experiment (N=215) complemented by semi-structured interviews (N=9), the authors conduct a mixed-methods evaluation of multiple formats in realistic editing tasks. While HJSON and YAML demonstrate marginal advantages in specific modification scenarios, these benefits vanish in both simpler and more complex contexts. Crucially, the findings reveal that usability is not primarily determined by syntactic differences but rather by socio-technical ecosystem factors—including tooling support, documentation quality, and community practices. This work provides the first empirical evidence that ecosystem support exerts a more decisive influence on format usability than syntax design alone.
This work addresses the lack of systematic evaluation of behavioral robustness in large language model (LLM) document workflows when confronted with semantically equivalent inputs in varying formats. The authors propose the first format-aware metamorphic testing framework, leveraging three classes of metamorphic relations to conduct a large-scale empirical analysis of end-to-end workflow performance across diverse document formats. They further introduce a lightweight, training-free format adaptation strategy. Across 48,000 experiments, they find that format changes can degrade accuracy by up to 53.63% and induce decision drift in over 41% of instances. Their method recovers up to 44.21% of the lost performance, offering the first systematic evidence of how document formatting impacts LLM reliability and establishing a novel paradigm for testing and mitigating format-induced fragility.
Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.
This study addresses the lack of systematic understanding regarding the creation, application, and organizational impact of data visualization style guides. Through interviews with nine authors from journalism, government, and industry, complemented by a cross-case analysis of 26 published guides, the paper proposes the PRISM socio-technical framework to elucidate their operational logic across four dimensions: Purpose, Rules and mechanisms, Institutional enforcers, and Situational flexibility. The findings reveal an inherent tension between standardization and adaptability, demonstrating that publicly available guides represent only partial manifestations of more comprehensive internal systems. By unpacking how these guides function in practice, the research offers both theoretical grounding and novel practical insights for the future development of visualization design standards.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.
This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.