data serialization

Designs, builds, and evaluates formats, schemas, and pipelines to encode and decode data and machine-learning model artifacts so they can be stored, transmitted, and reconstructed reliably. This includes creating and optimizing model and data serialization formats for token- and bandwidth-efficiency, low-latency and streaming/context serialization, and handling compatibility, versioning, and performance trade-offs.

dataserialization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of empirical guidance in selecting model export formats during AI system development. We systematically evaluate five formats—ONNX, SavedModel, TorchScript, Pickle, and Joblib—across integration efficiency, cross-platform compatibility, and maintenance cost. Employing an embedded multi-case empirical design—including two industrial systems and three distinct technology stacks—we integrate questionnaire surveys (n=17), structured on-site observations, and qualitative thematic analysis. Our findings reveal that ONNX achieves the best overall balance in cross-framework portability and integration efficiency; SavedModel uniquely excels in end-to-end deep learning pipelines, particularly in preprocessing encapsulation; whereas Pickle and Joblib exhibit pervasive security vulnerabilities and environment coupling, incurring the highest integration costs. This work provides the first engineering-oriented, empirically grounded basis for model serialization format selection in production AI deployment.

AI System DevelopmentEfficiency and CompatibilityModel Formats

This study investigates the impact of output format on the performance of large language models in single-turn code generation tasks and its interaction with model identity. Through controlled experiments across four open-source projects, three prominent models—Doubao, DeepSeek, and Qwen—were evaluated using three output formats: JSON Patch, unified diff, and full file, resulting in 4,013 test cases. The findings reveal that output format significantly affects success rates, with no universally optimal format; instead, each model exhibits distinct format preferences—for instance, Doubao achieves a 94% success rate with JSON Patch, while DeepSeek performs best (66%) with unified diff. This work is the first to demonstrate a strong interaction effect between output format and model identity, proposing model-specific output strategies and design principles for tools aimed at preventing format misuse.

coding agentformat misuseinteraction effect

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Aggregating empirical evidence from data strategy studies: a case on model quantization

May 01, 2025
SD
Santiago del Rey
🏛️ Universitat Polit`ecnica de Catalunya | UNIRIO | UFRJ

This study systematically evaluates the impact of model quantization on the correctness and resource efficiency of deep learning systems, while also exploring methodologies for cross-study evidence aggregation in data-driven empirical research. Methodologically, it innovatively applies Structured Synthesis Methods (SSM) for the first time in this domain, integrating findings from six empirical studies covering 19 models through a qualitative-quantitative mixed analysis. Results demonstrate that quantization yields substantial resource gains—average storage compression of ×3.2, inference latency reduction of −41%, and GPU energy consumption decrease of −38%—with only a marginal correctness degradation (−1.7% on average), representing a well-controlled trade-off. The study identifies both consistent patterns and fragmentation bottlenecks in quantization effects, and proposes a refined empirical research framework and methodological guidelines tailored to quantization techniques. These contributions provide foundational methodological support and practical guidance for optimizing trustworthy AI systems.

Assessing model quantization effects on DL correctness and efficiencyEvaluating trade-offs between correctness and resource efficiency in quantizationExploring methodological challenges in aggregating data strategy studies

Latest Papers

What's happening recently
View more

This study addresses a critical yet previously unrecognized issue in knowledge graph construction: the coupling between tabular serialization formats and schema constraints, which significantly degrades both factual coverage and graph fidelity—particularly in country-year statistical tables, where it induces entity inflation or extraction failure. The authors formally identify and name this phenomenon “format-constraint coupling,” introduce a direct graph access evaluation paradigm, and release CSVFidelity-Bench, a benchmark comprising diverse table types and gold-standard facts. Through factorial experiments, bootstrap confidence intervals, token ablation studies, and multi-LLM comparisons, they uncover significant positive coupling effects in four out of six datasets (peak effect size +1.180). Direct graph access reveals a quality gap as large as 47.6 percentage points (p<0.0001), substantially exceeding that of standard retrieval-based approaches.

entity inflationformat-constraint couplingknowledge graph fidelity

"This study addresses the validation issues that may arise when machine learning models are loaded across different software environments, highlighting that relying solely on artifact integrity checks is insufficient. To tackle this, the work proposes Modelstamp, a lightweight Python persistence library designed to verify the consistency of models and their runtime environments before deserialization. By encapsulating SHA-256 digests, runtime metadata, and installed version information into a JSON manifest, and supporting HMAC authentication, Modelstamp ensures secure and reliable cross-environment loading. Experimental results demonstrate that Modelstamp effectively enhances model loading security in 14 environment drift and 8 trust boundary scenarios, with an average verification time ranging from 0.032 seconds to 3.334 seconds for data sizes between 10 MiB and 1 GiB, maintaining a throughput of 307-312 MiB/s."

artifact integritymachine-learning modelsruntime environment

Converting ROS bags into machine learning datasets often relies on ad hoc scripts, resulting in substantial engineering overhead and inefficient iteration. This work introduces, for the first time, the principles of software build systems to robotic dataset construction, proposing a reproducible, incremental generation method grounded in artifact- and dependency-graph semantics. We present Bagzel, an open-source tool built on Bazel, which supports export to the nuScenes format and incorporates Bagzel-xattr for server-side metadata management. Experimental evaluation demonstrates that, on a 20.4 GB dataset, hot builds achieve up to a 386.26× speedup and incremental builds are accelerated by 7.21×, with performance gains further amplified as dataset scale increases.

build processdata pipelinereproducibility

This study addresses the lack of empirical evidence in data quality management for AI systems, where traditional perspectives struggle with model attribution and compliance challenges. Employing reflexive thematic analysis through in-depth interviews with 16 practitioners, this work examines the engineering and organizational dimensions of data quality in AI-driven systems, revealing emergent characteristics including traceability, circularity, and legitimacy. It introduces a novel conceptual framework termed “lifecycle assurance” that integrates fragmented machine learning research agendas and establishes evidence-generation mechanisms supporting specific AI claims. Furthermore, the study identifies six overarching themes and five trust-influencing conditions, offering practice-based, engineering-oriented guidance for managing data quality in AI systems.

AI-driven systemsdata qualityfoundation models

Hot Scholars

LC

Leshem Choshen

MIT, IBM AI research
Model RecyclingEvolving Collaborative PretrainingEvaluationModel Merging
AK

Alois Knoll

Technische Universität München
RoboticsAISensor Data FusionAutonomous Driving
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
ZT

Zhengzhong Tu

Texas A&M University, Google Research, University of Texas at Austin
Agentic AITrustworthy AIEmbodied AI
MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking