data format design

Designs, specifies, and documents data and file/storage formats and their logical and physical layouts (serialization, metadata/schema, encoding, compression and indexing), and builds or analyzes tools and transformations that optimize layout and representation for performance, space, compatibility and evolution; selects, governs, and manages format choices across systems and workloads to meet interoperability, versioning, and storage/access tradeoffs.

dataformatdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

This study addresses the lack of systematic guidance on contextualization strategies for large language model (LLM) agents operating in structured data environments, particularly concerning effectiveness and efficiency across multi-file, large-scale schemas. Using SQL generation as a proxy task, the work presents the first systematic evaluation of eleven models across four context formats—YAML, Markdown, JSON, and TOON—at schema scales ranging from 10 to 10,000 tables. The findings reveal that model capability tiers critically determine optimal context architecture: tailored strategies significantly improve performance, with state-of-the-art models gaining 2.7% accuracy under native file-based contexts, while open-source models average a 7.7% decline. Moreover, native file-based agents scale efficiently to ten-thousand-table schemas while maintaining high navigation accuracy.

context engineeringfile-native systemsLLM agents

A metadata model for profiling multidimensional sources in data ecosystems

Mar 20, 2025
CD
C. Diamantini
🏛️ Università Politecnica delle Marche

To address insufficient semantic description of multidimensional aggregate/summary data, poor adaptability of metadata standards, and cross-source interoperability challenges in big data environments, this paper proposes a multidimensional data source profiling metadata model tailored for data ecosystems. Built upon RDF, the model is the first to support extensible semantic modeling of both aggregate and summary multidimensional data, enabling semantic alignment of dimensions and measures with reference knowledge graphs. It integrates multi-granularity metadata profiles—spanning source-level, attribute-level, and value-distribution characteristics. The model ensures flexible extensibility and cross-source interoperability. Experimental results demonstrate that profile generation time scales linearly with data cardinality, confirming its engineering practicality and predictable performance.

Challenges in managing diverse data formats in Big Data ecosystems.Inadequate metadata vocabularies for aggregated or summary data.Need for efficient metadata model for multidimensional data profiling.

This study addresses the lack of systematic understanding regarding the usability of textual serialization formats such as JSON and XML, particularly concerning the factors that influence cognitive efficiency and user experience. Through a large-scale crowdsourced experiment (N=215) complemented by semi-structured interviews (N=9), the authors conduct a mixed-methods evaluation of multiple formats in realistic editing tasks. While HJSON and YAML demonstrate marginal advantages in specific modification scenarios, these benefits vanish in both simpler and more complex contexts. Crucially, the findings reveal that usability is not primarily determined by syntactic differences but rather by socio-technical ecosystem factors—including tooling support, documentation quality, and community practices. This work provides the first empirical evidence that ecosystem support exerts a more decisive influence on format usability than syntax design alone.

cognitive efficiencydata serializationsociotechnical ecosystems

This work addresses the lack of systematic evaluation of behavioral robustness in large language model (LLM) document workflows when confronted with semantically equivalent inputs in varying formats. The authors propose the first format-aware metamorphic testing framework, leveraging three classes of metamorphic relations to conduct a large-scale empirical analysis of end-to-end workflow performance across diverse document formats. They further introduce a lightweight, training-free format adaptation strategy. Across 48,000 experiments, they find that format changes can degrade accuracy by up to 53.63% and induce decision drift in over 41% of instances. Their method recovers up to 44.21% of the lost performance, offering the first systematic evidence of how document formatting impacts LLM reliability and establishing a novel paradigm for testing and mitigating format-induced fragility.

decision driftdocument formatformat robustness

Latest Papers

What's happening recently
View more

Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.

content-awaredocument formattingevaluation benchmark

This study addresses the lack of systematic understanding regarding the creation, application, and organizational impact of data visualization style guides. Through interviews with nine authors from journalism, government, and industry, complemented by a cross-case analysis of 26 published guides, the paper proposes the PRISM socio-technical framework to elucidate their operational logic across four dimensions: Purpose, Rules and mechanisms, Institutional enforcers, and Situational flexibility. The findings reveal an inherent tension between standardization and adaptability, demonstrating that publicly available guides represent only partial manifestations of more comprehensive internal systems. By unpacking how these guides function in practice, the research offers both theoretical grounding and novel practical insights for the future development of visualization design standards.

accountabilityflexibilitygovernance

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

Hot Scholars

TT

Thierry Tambe

Assistant Professor of Electrical Engineering, Stanford University
Computer ArchitectureVLSI
CF

Chao Fang

Shanghai Qi Zhi Institute
efficient MLAI acceleratorhardware-software co-designprecision-scalable computing
CB

Christopher Brix

PhD Candidate in Computer Science, RWTH Aachen University
Verification of Neural NetworksMachine LearningArtificial IntellligenceNeural Networks
JW

Junsong Wang

IBM Reserch, China
Wireless Comminications
SC

Siheng Chen

Shanghai Jiao Tong University
Collective intelligenceLLM agentgraph signal processingcollaborative perception