file format conversion

Designs and implements tools and pipelines that transform datasets and variables between file and data formats—e.g., general format converters, ETL workflows, structured-to-text and graph serializations, and format conversion pipelines or tools. Builds conversion and packaging processes that ensure data quality and compatibility, preserve analytic signals during serialization, document conversion steps and usage, and integrate serialized variables into downstream workflows.

fileformatconversion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$237K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

PRE-Share Data: Assistance Tool for Resource-aware Designing of Data-sharing Pipelines

Mar 17, 2025
SM
Sepideh Masoudi
🏛️ Technische Universität Berlin

In cross-organizational data sharing, existing multi-pipeline transformation design suffers from low efficiency and severe resource waste under dual constraints of governance compliance and recipient-side adaptability. Method: This paper proposes a reuse-aware pipeline design assistance paradigm that integrates flowchart-based modeling, semantic matching of transformation operations, fine-grained resource consumption modeling, and heuristic configuration optimization. It enables automatic identification of reusable transformation components across pipelines, recommends optimal pipeline structures, and quantifies potential resource savings. Contribution/Results: As the first design assistance framework supporting predictive reporting generation, it achieves, on real-world use cases, an average 37% reduction in computational resource consumption and a 52% reduction in design cycle time, while remaining compatible with self-service data platform deployments.

Designing efficient data-sharing pipelines across organizationsEnsuring compliance with governance policies and recipient requirementsReusing transformation processes to optimize resource consumption

Existing visual analytics workflows are predominantly described in unstructured textual form, hindering systematic comparison, reuse, and practical guidance. This work proposes ATWL, a formal, declarative language for modeling visual analytics workflows through a modular ontology grounded in eight artifact types and standardized intents. For the first time, this approach enables structured, machine-interpretable representations of such workflows. Leveraging large language models, the authors automatically extract workflows from academic papers to construct a reusable repository comprising 17 annotated instances. Empirical evaluation demonstrates that ATWL effectively uncovers cross-workflow structural patterns and yields more compact, structured, and extensible analytical recommendations than original narrative descriptions, thereby facilitating efficient in-context reuse and adaptation.

analytical knowledge representationsystematic comparisonunstructured prose

This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.

black-box convertersconverter orchestrationdata model consistency

Towards Next Generation Data Engineering Pipelines

Jul 18, 2025
KM
Kevin M. Kramer
🏛️ University of Hagen | University of Regensburg

Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.

Achieving self-awareness and self-adaptationEnabling reactivity to data changesImproving data quality in engineering pipelines

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.

auditabilitydata analysisgovernance

This work addresses the challenges of poor readability, reproducibility, and maintainability in R scripts, which are exacerbated by the lack of effective tooling. To tackle this, the authors propose flowR, a plugin integrated into Positron and VS Code that innovatively combines incremental interprocedural data-flow and control-flow analysis to construct a unified data-flow graph accommodating R’s dynamic semantics. The system offers interactive visualization, static backward program slicing, inline value annotations, and linting capabilities, all built upon a modular, extensible architecture. Experimental results demonstrate that flowR constructs complete data-flow graphs in an average of 576 milliseconds, enabling near real-time feedback and substantially enhancing script understandability and maintainability.

comprehensibilitydata analysis scriptsmaintainability

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

This study addresses the lack of systematic understanding regarding the usability of textual serialization formats such as JSON and XML, particularly concerning the factors that influence cognitive efficiency and user experience. Through a large-scale crowdsourced experiment (N=215) complemented by semi-structured interviews (N=9), the authors conduct a mixed-methods evaluation of multiple formats in realistic editing tasks. While HJSON and YAML demonstrate marginal advantages in specific modification scenarios, these benefits vanish in both simpler and more complex contexts. Crucially, the findings reveal that usability is not primarily determined by syntactic differences but rather by socio-technical ecosystem factors—including tooling support, documentation quality, and community practices. This work provides the first empirical evidence that ecosystem support exerts a more decisive influence on format usability than syntax design alone.

cognitive efficiencydata serializationsociotechnical ecosystems

Hot Scholars

JC

Jorge Calvo-Zaragoza

Universidad de Alicante
Optical Music RecognitionHandwritten Text RecognitionMusic Information Retrieval
SP

Shrimai Prabhumoye

Senior Research Scientist @NVIDIA and Adjunct Assistant Professor @Boston University
Natural Language Processing
XZ

Xuanhe Zhou

Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence
DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics