Score
Designs and implements structured data representations and formatting that convert raw artifacts and metadata into text or vector-ready inputs for models and retrieval systems. This includes selecting and normalizing fields, concatenating and formatting model metadata or document content, and tuning representation choices to optimize retrieval and embedding performance metrics.
Text-to-structured generation (e.g., tables, knowledge graphs, charts) for agent-centric AI is a foundational infrastructure enabling context-aware retrieval and autonomous reasoning, yet suffers from fragmented methodologies, scarce standardized datasets, and inconsistent evaluation protocols. Method: We conduct a systematic literature review integrating techniques from NLP, information extraction, knowledge representation, and machine learning to establish the first holistic analytical framework—comprising task taxonomy, benchmark dataset inventory, and unified evaluation metrics. Contribution/Results: We introduce the first general-purpose evaluation framework for structured output generation, explicitly identifying methodological limitations and core challenges (e.g., fidelity, composability, and reasoning-aware assessment). We comprehensively map research gaps and affirm the centrality of this direction in next-generation AI systems, providing both theoretical grounding and practical guidance for future algorithmic development and empirical validation.
This work systematically evaluates the efficacy of large language models (LLMs) in automatically converting unstructured textual recipes into the structured Cooklang format. Method: We benchmark GPT-4o, GPT-4o-mini, and Llama3.1 variants under zero- and few-shot settings, and propose the first multidimensional evaluation framework integrating conventional metrics (WER, ROUGE-L, TER) with domain-specific semantic element identification. Contribution/Results: GPT-4o achieves a ROUGE-L score of 0.9722 and WER of 0.0730 in few-shot settings; fine-tuned Llama3.1-8B demonstrates substantial performance gains, confirming the optimization potential of smaller models. This study provides the first empirical validation that LLMs can perform domain-specific structured conversion with high accuracy, establishing a scalable and quantitatively assessable paradigm for standardizing unstructured data across industries.
This study addresses the limitations of semantic similarity–based retrieval in structured, highly repetitive regulatory texts, where linguistic overlap often obscures meaningful content distinctions and undermines the effectiveness of retrieval-augmented generation (RAG). To mitigate this issue, the work systematically investigates metadata-aware retrieval strategies, proposing and evaluating fusion approaches such as unified embedding and prefix concatenation. The findings demonstrate that incorporating metadata enhances intra-document cohesion and reduces inter-document ambiguity, thereby improving retrieval performance. Evaluated on a newly curated benchmark dataset, RAGMATE-10K, both the unified embedding and prefix-based methods significantly outperform pure text baselines across multiple question types and evaluation metrics. Notably, the unified embedding approach achieves superior performance while maintaining greater maintainability.
This study addresses the sensitivity of structured metadata retrieval models to field ordering, which causes overreliance on positional cues rather than semantic field labels, thereby impairing discoverability in cross-lingual low-resource settings. To mitigate this issue, the authors propose Permutation-Invariant Fine-Tuning (PI-FT), a lightweight approach that randomizes field order and stochastically drops fields during data loading, encouraging the model to attend to semantic labels instead of positional patterns. Implemented with only two lines of code modification in the data loader, PI-FT enables a 118M-parameter CPU-based model to achieve an nDCG@10 of 0.707 on nearly 10,000 development statistics—outperforming all zero-shot baselines, including text-embedding-3-large—and reduces performance degradation under field-order perturbations from 7.4 to just 0.2 points, substantially enhancing robustness and generalization.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This study addresses the challenge of natural language–driven simulation model discovery by systematically investigating the impact of data representation, Transformer-based embedding models, and reranking strategies on retrieval performance. By constructing multimodal model metadata and leveraging standard information retrieval metrics, the work presents the first quantitative evaluation of open-source embedding models for this task. Experimental results demonstrate that the proposed approach achieves strong performance in recall@5 and nDCG@5, with reranking substantially enhancing effectiveness on complex queries. These contributions establish the first benchmark framework for AI-enabled model reusability, composability, and interoperability in simulation model retrieval.
Existing approaches to automatic document formatting suffer from imprecise target localization and redundant content re-reading in content-aware scenarios, compounded by the absence of a dedicated evaluation benchmark. To address these limitations, this work introduces DocFormBench—the first comprehensive evaluation benchmark specifically designed for content-aware document formatting—and proposes DocFormFlow, a decoupled workflow that separates the task into two distinct phases: “what to format” (target localization) and “how to format” (format execution). By integrating large language models with multimodal models, DocFormFlow demonstrates significant improvements in formatting accuracy and substantially reduces token consumption across multiple mainstream models, underscoring precise target localization as a critical factor for high performance.
This work addresses the challenge that current machine translation systems struggle to preserve alignment between textual content and visual layout—such as tables and mathematical formulas—when processing PDF documents, often resulting in structural distortion. To tackle this issue, the authors construct a multimodal parallel corpus comprising 3,956 PDFs across 15 language pairs, meticulously retaining original layout metadata. They introduce, for the first time, a 45-dimensional geometric feature space to guide K-Medoids sampling, prioritizing visual structural diversity. This dataset enables layout-aware translation that jointly leverages textual and visual context, exposing critical limitations of existing systems in spatial localization and geometric synchronization. It thus establishes a high-fidelity, evaluable benchmark for document translation and reconstruction that faithfully preserves complex layouts.
This work addresses the sensitivity of table retrieval to serialization formats such as CSV and HTML, which causes semantically identical tables to yield substantially different embeddings when represented in distinct formats. Treating each format as a noisy view of a shared underlying semantic structure, the authors propose a centroid-based geometric correction method. This approach constructs a canonical representation from the centroid of embeddings across multiple formats and introduces a lightweight residual bottleneck adapter combined with covariance regularization to align single-format embeddings toward this centroid. The adapter is fine-tuned while keeping the base encoder frozen, enabling compatibility with various models including MPNet, BGE-M3, ReasonIR, and SPLADE. Experiments demonstrate that the centroid representation consistently outperforms any single-format embedding, and the adapter significantly enhances robustness in dense retrieval, though it yields limited gains for sparse retrieval.
This work addresses the limitations of current scientific document retrieval methods, which predominantly rely on document image representations and struggle to effectively leverage critical evidence embedded in structured content such as text, tables, and mathematical formulas. To this end, the authors introduce ArXivDoc, a novel benchmark constructed from LaTeX source code that enables controlled query generation, facilitating a systematic evaluation of textual, visual, and multimodal representations for retrieval. Experimental results demonstrate that textual representations consistently outperform others across diverse query types; multimodal approaches combining text and images achieve substantial gains over image-only methods without requiring specialized training; and image-based representations exhibit significant performance degradation with increasing document length, particularly for structured content. This study thus exposes fundamental shortcomings of image-centric paradigms and establishes a new foundation for advancing scientific document retrieval.