quantitative indicator extraction

Designs and implements systems and pipelines that convert unstructured text into structured numeric indicators by applying LLM-based information extraction and text-to-numeric mapping. This includes prompt engineering and model orchestration to extract entities and aspect scores from noisy or multilingual text and to produce validated structured numeric outputs for downstream analysis.

quantitativeindicatorextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work systematically evaluates the efficacy of large language models (LLMs) in automatically converting unstructured textual recipes into the structured Cooklang format. Method: We benchmark GPT-4o, GPT-4o-mini, and Llama3.1 variants under zero- and few-shot settings, and propose the first multidimensional evaluation framework integrating conventional metrics (WER, ROUGE-L, TER) with domain-specific semantic element identification. Contribution/Results: GPT-4o achieves a ROUGE-L score of 0.9722 and WER of 0.0730 in few-shot settings; fine-tuned Llama3.1-8B demonstrates substantial performance gains, confirming the optimization potential of smaller models. This study provides the first empirical validation that LLMs can perform domain-specific structured conversion with high accuracy, establishing a scalable and quantitatively assessable paradigm for standardizing unstructured data across industries.

Assessing performance of LLMs in transforming recipe text to Cooklang format.Evaluating LLMs' ability to convert unstructured text into structured formats.Exploring potential of LLMs for automated structured data generation in various domains.

LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows

Oct 06, 2025
SA
Samy Ateia
🏛️ University of Regensburg | University of Bayreuth

With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.

Enabling rapid domain adaptation for scientific information extractionExtracting key concepts from scientific documents using LLMsSupporting FAIR principles in scientific publishing workflows

Themes of Building LLM-based Applications for Production: A Practitioner's View

Nov 13, 2024
AM
Alina Mailach
🏛️ ScaDS.AI Dresden/Leipzig | Leipzig University

Current LLM application development lacks systematic, practice-informed guidelines, leading to a growing gap between academic research and industrial engineering. Method: Drawing on transcribed texts from 189 real-world developer practice videos (2022–2024), we integrate BERTopic-based automated topic modeling with iterative human refinement to construct the first empirically grounded, production-oriented thematic map of LLM application development. Contribution/Results: The map identifies eight core themes—including design & architecture, model enhancement, infrastructure, and ethical risk—spanning 20 key issues. Design & Architecture emerges as the most densely populated theme, with RAG at its architectural center; prompt engineering, fine-tuning, deployment toolchains, and AI ethics are recurrent high-frequency concerns. Critically, the map exposes significant lags in academic research relative to industrial practice and delivers an actionable, empirically validated priority framework—thereby bridging a critical empirical gap in the LLM engineering knowledge base.

Analyze practitioner discussions on LLM deployment challengesIdentify key considerations for LLM-based system developmentProvide systematic overview of LLM application priorities

LLMs with Industrial Lens: Deciphering the Challenges and Prospects - A Survey

Feb 22, 2024
AU
Ashok Urlana
🏛️ TCS Research | IIIT Hyderabad

This study systematically investigates core challenges impeding large language model (LLM) industrial deployment, identifying 12 representative bottlenecks across four critical dimensions: data scarcity, inefficient inference, complex deployment, and inaccurate evaluation. Method: We employ a mixed-methods approach—structured interviews with frontline practitioners, a research-question-driven review of 68 industrial practice papers, and qualitative content analysis. Contribution/Results: We propose the first “industry-perspective-driven” taxonomy for LLM deployment challenges; establish a dynamically updated GitHub knowledge repository of industrial LLM literature; and deliver an actionable, lifecycle-spanning optimization roadmap. The framework has been adopted by multiple enterprises and serves as a key reference benchmark for industrial LLM adoption.

Exploring challenges in industrial LLM applicationsIdentifying opportunities for enhancing LLM utilizationSurveying industry practices and research on LLMs

Latest Papers

What's happening recently
View more

This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.

annotation errorlarge language modelsreproducibility

This work addresses the lack of formal fidelity verification methods for natural language outputs—such as Gherkin scenarios—generated by large language models (LLMs). To this end, we propose a logic-based consistency verification framework grounded in automated formalization. Methodologically, we introduce automated formalization to LLM output validation for the first time: an LLM-driven formalizer translates both natural language requirements and LLM-generated outputs into first-order logic formulas; formal reasoning is then applied to assess semantic equivalence and detect logical contradictions. Experiments demonstrate that our approach effectively identifies semantic equivalence across paraphrased expressions and uncovers latent logical inconsistencies, thereby significantly enhancing the trustworthiness of generated artifacts. Our primary contribution is establishing the first formal verification paradigm tailored to LLM-generated outputs, providing both theoretical foundations and practical methodology for ensuring the verifiability of automated artifacts in requirements engineering.

Developing formal methods to check logical consistency of autoformalized requirementsEnsuring fidelity between informal statements and LLM-generated formal outputsVerifying accuracy of LLM-generated structured outputs from natural language requirements

This work addresses the challenge of enabling efficient and precise structured querying over unstructured documents, a task hindered by the limitations of existing vector retrieval methods—namely, ambiguous matching and high computational overhead. To overcome these issues, the authors propose an Annotation Index coupled with a SchemaLoop mechanism that automatically constructs hierarchical annotation schemas to transform unstructured text into structured data. They further introduce a SQL-extended query engine that integrates lightweight language models for attribute extraction and large language models for deep semantic reasoning, enhanced by a multi-stage cost-aware execution strategy and incremental index updates. Evaluated on three real-world datasets, the approach achieves an average F1 score of 0.87, substantially outperforming current methods, particularly in complex multi-hop and progressive reasoning queries.

analytical queriesinformation extractionprecise querying

Accurately extracting UK Research Excellence Framework (REF) ratings (1*–4*) from noisy, unstructured text containing missing or invalid values presents a significant challenge, requiring large language models (LLMs) to output only normalized integers (1–4) or a designated missing-value indicator (−1). To address this, this work introduces the first standardized prompt engineering benchmark for complex numerical extraction tasks, accompanied by a publicly available dataset of 1,446 short texts with gold-standard annotations. By integrating semantic understanding with explicit rule-based constraints, an initial prompting strategy achieves 72.6% accuracy. The study clarifies the definition of valid ratings and formalizes a mechanism for handling missing data, thereby advancing research into LLMs’ numerical reasoning and instruction-following capabilities, with the aim of fostering community-driven improvements in structured information extraction from noisy textual sources.

information extractionlarge language modelsmessy text

This study addresses the high cost and low efficiency of traditional ontology construction in specialized domains such as casting, which relies heavily on manual annotation and conventional NLP techniques. It presents the first systematic comparison of three few-shot information extraction strategies based on large language models (LLMs)—pretrained model prompting, in-context learning (ICL), and fine-tuning—for automatically extracting domain-specific terms and relations to build ontologies. Through expert validation, the research identifies the most effective LLM-based strategy and successfully constructs a high-quality ontology for the casting domain. The proposed approach significantly improves construction efficiency while maintaining high accuracy, offering a robust and scalable paradigm for automated knowledge modeling in specialized fields.

casting manufacturingdomain-specific knowledgeinformation extraction

Hot Scholars

MG

Meriem Guerar

University of Genova, Italy
PrivacySecurityFederated learningSSI
TF

Tegawendé F. Bissyandé

Chief Scientist II / ERC Fellow / TruX @SnT, University of Luxembourg
Software SecurityProgram RepairCode SearchMachine Learning
JJ

Jin Jin

University of Pennsylvania
Biostatistics
SA

Saja Al-Dabet

United Arab Emirates University
Natural Language ProcessingHealth InformaticsDeep LearningData Mining