Score
Designs and implements systems and pipelines that convert unstructured text into structured numeric indicators by applying LLM-based information extraction and text-to-numeric mapping. This includes prompt engineering and model orchestration to extract entities and aspect scores from noisy or multilingual text and to produce validated structured numeric outputs for downstream analysis.
This work systematically evaluates the efficacy of large language models (LLMs) in automatically converting unstructured textual recipes into the structured Cooklang format. Method: We benchmark GPT-4o, GPT-4o-mini, and Llama3.1 variants under zero- and few-shot settings, and propose the first multidimensional evaluation framework integrating conventional metrics (WER, ROUGE-L, TER) with domain-specific semantic element identification. Contribution/Results: GPT-4o achieves a ROUGE-L score of 0.9722 and WER of 0.0730 in few-shot settings; fine-tuned Llama3.1-8B demonstrates substantial performance gains, confirming the optimization potential of smaller models. This study provides the first empirical validation that LLMs can perform domain-specific structured conversion with high accuracy, establishing a scalable and quantitatively assessable paradigm for standardizing unstructured data across industries.
With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.
Current LLM application development lacks systematic, practice-informed guidelines, leading to a growing gap between academic research and industrial engineering. Method: Drawing on transcribed texts from 189 real-world developer practice videos (2022–2024), we integrate BERTopic-based automated topic modeling with iterative human refinement to construct the first empirically grounded, production-oriented thematic map of LLM application development. Contribution/Results: The map identifies eight core themes—including design & architecture, model enhancement, infrastructure, and ethical risk—spanning 20 key issues. Design & Architecture emerges as the most densely populated theme, with RAG at its architectural center; prompt engineering, fine-tuning, deployment toolchains, and AI ethics are recurrent high-frequency concerns. Critically, the map exposes significant lags in academic research relative to industrial practice and delivers an actionable, empirically validated priority framework—thereby bridging a critical empirical gap in the LLM engineering knowledge base.
This study systematically investigates core challenges impeding large language model (LLM) industrial deployment, identifying 12 representative bottlenecks across four critical dimensions: data scarcity, inefficient inference, complex deployment, and inaccurate evaluation. Method: We employ a mixed-methods approach—structured interviews with frontline practitioners, a research-question-driven review of 68 industrial practice papers, and qualitative content analysis. Contribution/Results: We propose the first “industry-perspective-driven” taxonomy for LLM deployment challenges; establish a dynamically updated GitHub knowledge repository of industrial LLM literature; and deliver an actionable, lifecycle-spanning optimization roadmap. The framework has been adopted by multiple enterprises and serves as a key reference benchmark for industrial LLM adoption.
This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.
This work addresses the lack of formal fidelity verification methods for natural language outputs—such as Gherkin scenarios—generated by large language models (LLMs). To this end, we propose a logic-based consistency verification framework grounded in automated formalization. Methodologically, we introduce automated formalization to LLM output validation for the first time: an LLM-driven formalizer translates both natural language requirements and LLM-generated outputs into first-order logic formulas; formal reasoning is then applied to assess semantic equivalence and detect logical contradictions. Experiments demonstrate that our approach effectively identifies semantic equivalence across paraphrased expressions and uncovers latent logical inconsistencies, thereby significantly enhancing the trustworthiness of generated artifacts. Our primary contribution is establishing the first formal verification paradigm tailored to LLM-generated outputs, providing both theoretical foundations and practical methodology for ensuring the verifiability of automated artifacts in requirements engineering.
This work addresses the challenge of enabling efficient and precise structured querying over unstructured documents, a task hindered by the limitations of existing vector retrieval methods—namely, ambiguous matching and high computational overhead. To overcome these issues, the authors propose an Annotation Index coupled with a SchemaLoop mechanism that automatically constructs hierarchical annotation schemas to transform unstructured text into structured data. They further introduce a SQL-extended query engine that integrates lightweight language models for attribute extraction and large language models for deep semantic reasoning, enhanced by a multi-stage cost-aware execution strategy and incremental index updates. Evaluated on three real-world datasets, the approach achieves an average F1 score of 0.87, substantially outperforming current methods, particularly in complex multi-hop and progressive reasoning queries.
Accurately extracting UK Research Excellence Framework (REF) ratings (1*–4*) from noisy, unstructured text containing missing or invalid values presents a significant challenge, requiring large language models (LLMs) to output only normalized integers (1–4) or a designated missing-value indicator (−1). To address this, this work introduces the first standardized prompt engineering benchmark for complex numerical extraction tasks, accompanied by a publicly available dataset of 1,446 short texts with gold-standard annotations. By integrating semantic understanding with explicit rule-based constraints, an initial prompting strategy achieves 72.6% accuracy. The study clarifies the definition of valid ratings and formalizes a mechanism for handling missing data, thereby advancing research into LLMs’ numerical reasoning and instruction-following capabilities, with the aim of fostering community-driven improvements in structured information extraction from noisy textual sources.
This study addresses the high cost and low efficiency of traditional ontology construction in specialized domains such as casting, which relies heavily on manual annotation and conventional NLP techniques. It presents the first systematic comparison of three few-shot information extraction strategies based on large language models (LLMs)—pretrained model prompting, in-context learning (ICL), and fine-tuning—for automatically extracting domain-specific terms and relations to build ontologies. Through expert validation, the research identifies the most effective LLM-based strategy and successfully constructs a high-quality ontology for the casting domain. The proposed approach significantly improves construction efficiency while maintaining high accuracy, offering a robust and scalable paradigm for automated knowledge modeling in specialized fields.