data representation engineering

Designs and implements structured data representations and formatting that convert raw artifacts and metadata into text or vector-ready inputs for models and retrieval systems. This includes selecting and normalizing fields, concatenating and formatting model metadata or document content, and tuning representation choices to optimize retrieval and embedding performance metrics.

datarepresentationengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$189K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work systematically evaluates the efficacy of large language models (LLMs) in automatically converting unstructured textual recipes into the structured Cooklang format. Method: We benchmark GPT-4o, GPT-4o-mini, and Llama3.1 variants under zero- and few-shot settings, and propose the first multidimensional evaluation framework integrating conventional metrics (WER, ROUGE-L, TER) with domain-specific semantic element identification. Contribution/Results: GPT-4o achieves a ROUGE-L score of 0.9722 and WER of 0.0730 in few-shot settings; fine-tuned Llama3.1-8B demonstrates substantial performance gains, confirming the optimization potential of smaller models. This study provides the first empirical validation that LLMs can perform domain-specific structured conversion with high accuracy, establishing a scalable and quantitatively assessable paradigm for standardizing unstructured data across industries.

Assessing performance of LLMs in transforming recipe text to Cooklang format.Evaluating LLMs' ability to convert unstructured text into structured formats.Exploring potential of LLMs for automated structured data generation in various domains.

This study addresses the limitations of semantic similarity–based retrieval in structured, highly repetitive regulatory texts, where linguistic overlap often obscures meaningful content distinctions and undermines the effectiveness of retrieval-augmented generation (RAG). To mitigate this issue, the work systematically investigates metadata-aware retrieval strategies, proposing and evaluating fusion approaches such as unified embedding and prefix concatenation. The findings demonstrate that incorporating metadata enhances intra-document cohesion and reduces inter-document ambiguity, thereby improving retrieval performance. Evaluated on a newly curated benchmark dataset, RAGMATE-10K, both the unified embedding and prefix-based methods significantly outperform pure text baselines across multiple question types and evaluation metrics. Notably, the unified embedding approach achieves superior performance while maintaining greater maintainability.

document retrievalmetadataRetrieval-Augmented Generation

This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.

embedding fine-tuningfield order sensitivitypermutation invariance

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Oct 28, 2024
QZ
Qintong Zhang
🏛️ Shanghai Artificial Intelligence Laboratory | Peking University

This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.

Address challenges in layout detection and multi-modal data integrationConvert unstructured documents into structured machine-readable dataImprove parsing accuracy for complex layouts and high-density text

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Latest Papers

What's happening recently
View more

This study addresses the challenge of natural language–driven simulation model discovery by systematically investigating the impact of data representation, Transformer-based embedding models, and reranking strategies on retrieval performance. By constructing multimodal model metadata and leveraging standard information retrieval metrics, the work presents the first quantitative evaluation of open-source embedding models for this task. Experimental results demonstrate that the proposed approach achieves strong performance in recall@5 and nDCG@5, with reranking substantially enhancing effectiveness on complex queries. These contributions establish the first benchmark framework for AI-enabled model reusability, composability, and interoperability in simulation model retrieval.

AI-driven retrievalmodel discoverymodel reuse

Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.

content-awaredocument formattingevaluation benchmark

This work addresses the challenge that current machine translation systems struggle to preserve alignment between textual content and visual layout—such as tables and mathematical formulas—when processing PDF documents, often resulting in structural distortion. To tackle this issue, the authors construct a multimodal parallel corpus comprising 3,956 PDFs across 15 language pairs, meticulously retaining original layout metadata. They introduce, for the first time, a 45-dimensional geometric feature space to guide K-Medoids sampling, prioritizing visual structural diversity. This dataset enables layout-aware translation that jointly leverages textual and visual context, exposing critical limitations of existing systems in spatial localization and geometric synchronization. It thus establishes a high-fidelity, evaluable benchmark for document translation and reconstruction that faithfully preserves complex layouts.

layout preservationmultimodal machine translationPDF translation

This work addresses the sensitivity of table retrieval to serialization formats such as CSV and HTML, which causes semantically identical tables to yield substantially different embeddings when represented in distinct formats. Treating each format as a noisy view of a shared underlying semantic structure, the authors propose a centroid-based geometric correction method. This approach constructs a canonical representation from the centroid of embeddings across multiple formats and introduces a lightweight residual bottleneck adapter combined with covariance regularization to align single-format embeddings toward this centroid. The adapter is fine-tuned while keeping the base encoder frozen, enabling compatibility with various models including MPNet, BGE-M3, ReasonIR, and SPLADE. Experiments demonstrate that the centroid representation consistently outperforms any single-format embedding, and the adapter significantly enhances robustness in dense retrieval, though it yields limited gains for sparse retrieval.

embedding varianceformat invariancerepresentational stability

This work addresses the limitations of current scientific document retrieval methods, which predominantly rely on document image representations and struggle to effectively leverage critical evidence embedded in structured content such as text, tables, and mathematical formulas. To this end, the authors introduce ArXivDoc, a novel benchmark constructed from LaTeX source code that enables controlled query generation, facilitating a systematic evaluation of textual, visual, and multimodal representations for retrieval. Experimental results demonstrate that textual representations consistently outperform others across diverse query types; multimodal approaches combining text and images achieve substantial gains over image-only methods without requiring specialized training; and image-based representations exhibit significant performance degradation with increasing document length, particularly for structured content. This study thus exposes fundamental shortcomings of image-centric paradigms and establishes a new foundation for advancing scientific document retrieval.

document-as-imageLaTeX sourcesmultimodal documents

Hot Scholars

ZM

Zhipeng Ma

Southwest Jiaotong University
Data-Centric AILarge Language ModelHuman Mobility
BN

Bo Nørregaard Jørgensen

Professor, PhD., Head of Center for Energy Informatics, University of Southern Denmark
Energy InformaticsEnergy-ecosystemsAI AgentsMulti-agent systems
OP

Oiwi Parker Jones

Applied Artificial Intelligence and Clinical Neurosciences, University of Oxford
AINeuroscienceDeep LearningSpeech Recognition
IS

Ioannis Savidis

Associate Professor, Drexel University
VLSI3-D Integrationhardware security and trustlow power circuits
DD

Daniel DeAlcala

PhD Student, Universidad Autónoma de Madrid
Deep LearningSignal Processing