Score
Design and build systems that translate rules expressed in natural language into typed intermediate representations (IR), producing formal predicates, slots, and structured nodes from text. These systems identify and normalize entities and actions, resolve temporal and conditional constraints, and output predicate- or IR-based encodings suitable for downstream enforcement, verification, or reasoning.
Text-to-structured generation (e.g., tables, knowledge graphs, charts) for agent-centric AI is a foundational infrastructure enabling context-aware retrieval and autonomous reasoning, yet suffers from fragmented methodologies, scarce standardized datasets, and inconsistent evaluation protocols. Method: We conduct a systematic literature review integrating techniques from NLP, information extraction, knowledge representation, and machine learning to establish the first holistic analytical framework—comprising task taxonomy, benchmark dataset inventory, and unified evaluation metrics. Contribution/Results: We introduce the first general-purpose evaluation framework for structured output generation, explicitly identifying methodological limitations and core challenges (e.g., fidelity, composability, and reasoning-aware assessment). We comprehensively map research gaps and affirm the centrality of this direction in next-generation AI systems, providing both theoretical grounding and practical guidance for future algorithmic development and empirical validation.
This work addresses the challenge of compiling natural language queries into backend query languages in document-centric, hybrid, and heterogeneous data environments, where semantic intent is often ambiguous or incomplete. The authors propose the NLIQ framework, which introduces a “goal sufficiency” criterion to classify queries according to their semantic determinacy. It emphasizes that when intermediate goals must be dynamically constructed, intermediate representations should serve as core semantic objects rather than mere syntactic intermediaries. Through conceptual analysis, case modeling, and formal categorization, the study establishes a unified query paradigm that integrates goal recognition, intermediate representation design, and heterogeneous execution. This framework provides a theoretical foundation for natural language querying in complex data settings and opens new research directions in semantic goal construction, heterogeneous compilation, and answer generation.
This work investigates large language models’ (LLMs) capacity to comprehend compiler intermediate representations (IR), specifically evaluating their performance on four core tasks: control-flow graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Through systematic multi-model benchmarking (GPT-4, LLaMA 3.1, Gemma 2, etc.), a curated structured IR dataset, task-specific prompt engineering, and fine-grained error attribution, the study provides the first empirical evidence of fundamental limitations in LLMs’ IR understanding—particularly in CFG reconstruction (accuracy <42%) and execution reasoning (error rate 68%). Methodologically, it introduces a dual-path enhancement paradigm: (1) IR-domain fine-tuning and (2) explicit control-flow modeling. Experimental results demonstrate that targeted fine-tuning improves task performance by up to 31.5%, establishing a foundational framework for advancing LLM-based IR analysis.
This work addresses the challenge of format and semantic errors produced by small language models when translating natural language into first-order logic (FOL), which undermines the reliability of symbolic reasoning. To mitigate this, the authors propose a staged incremental reasoning framework: first, a large language model synthesizes training data to supervise fine-tuning of the small model; then, the translation process is decoupled into predicate generation and FOL formulation stages. An external verification module is introduced to detect and correct predicate arity errors, thereby enhancing translation accuracy. Evaluated on four logical reasoning benchmarks, the approach significantly reduces error rates, improves predicate coverage, and boosts overall reasoning performance, bringing small models closer to reliable, verifiable symbolic reasoning systems.
Current information retrieval systems struggle with complex queries requiring logical constraints, multi-step reasoning, and evidence synthesis, primarily due to a lack of structured reasoning capabilities. This work proposes the first unified framework for structured reasoning in information retrieval, systematically integrating interdisciplinary approaches—including large language model reasoning strategies, neuro-symbolic systems, probabilistic and Bayesian methods, geometric representations, and energy-based models—to elucidate their inherent trade-offs and complementary mechanisms. By bridging disciplinary boundaries, the framework clarifies the central role of retrieval within general-purpose reasoning systems and provides researchers with conceptual tools and practical guidance to advance the development of verifiable, structured reasoning architectures.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This work addresses the challenge large language models face in accurately translating natural language descriptions into optimization models, particularly when handling composite constraints and intricate business rules. To bridge this gap, the authors propose a Canonical Intermediate Representation (CIR) that decouples business rule logic from its mathematical implementation, serving as a semantic intermediary between natural language and formal optimization models. They further introduce a multi-agent R2C framework that parses input text, retrieves domain knowledge, generates CIR specifications, and subsequently instantiates them into mathematical optimization models. The paper establishes the first systematic benchmark for rule-to-constraint reasoning and demonstrates state-of-the-art performance: on a newly curated complex-rule benchmark, their method achieves 47.2% accuracy, matches or exceeds closed-source models like GPT-5 on existing benchmarks, and sets new best results on select tasks through a self-reflection mechanism.
This study addresses the challenge in regulated public procurement where bid validation must simultaneously ensure factual accuracy and legal auditability—a balance difficult to achieve with conventional approaches that struggle to integrate semantic understanding with rule-based interpretability. To bridge this gap, the authors propose a neuro-symbolic method that synergistically combines large language models (LLMs) with Logic Tensor Networks (LTNs). Specifically, LLMs are employed to extract semantic predicates from textual bid documents, which are then integrated into an LTN framework that encodes domain-specific regulatory rules to perform auditable, logic-driven inference. Experimental evaluation on real-world procurement documents demonstrates that the proposed approach not only maintains competitive performance but also significantly enhances decision interpretability and system modularity, thereby offering robust support for explainable artificial intelligence (XAI) in high-stakes regulatory contexts.
Existing methods for generating detection rules rely on specific input-output pairs and lack a unified framework. This work formalizes the task for the first time as a unified mapping from contextual and target-language inputs to detection rules, introducing UniRule—a novel framework that models semantic distance in a dual semantic projection space encompassing detection intent and detection logic to characterize optimal rules. UniRule integrates an agent-based RAG architecture to achieve generalization across diverse contexts and languages. Experimental results demonstrate that UniRule significantly outperforms pure large language model (LLM) approaches across twelve distinct scenarios, achieving a Bradley-Terry preference coefficient of 0.52, thereby validating its effectiveness and broad applicability.
Existing formal methods incur high costs in specification construction and maintenance and lack scalability, making them ill-suited for verifying modern AI systems. This work proposes a Learning-Integrated Formal Reasoning (LIFR) framework that innovatively combines machine learning with formal verification: it employs natural language processing to automatically generate contracts, leverages graph matching and representation learning to achieve semantic alignment and cross-system reuse of verification artifacts, and establishes a rigorous semantic foundation grounded in Unifying Theories of Programming (UTP) and institution theory. By shifting formal verification from isolated proofs toward a cumulative, knowledge-driven paradigm, the LIFR framework substantially enhances automation and scalability while preserving formal rigor.
This work addresses key limitations of existing Text-to-SQL approaches, which struggle with complex reasoning, integration of domain knowledge, and hypothetical queries, while also incurring high deployment costs in enterprise settings. To overcome these challenges, the authors propose the IESR framework, which leverages a lightweight, non-finetuned large language model to perform information-enhanced structured reasoning. IESR decouples mathematical computation from SQL generation, integrates schema linking with semantic understanding, and introduces a Monte Carlo Tree Search (MCTS)-based multi-path reasoning mechanism coupled with trajectory consistency verification. Without any model fine-tuning, this approach achieves state-of-the-art performance on complex reasoning benchmarks, including LogicQA (24.28 EX) and Archer (37.28 EX), demonstrating significant gains in reasoning accuracy using only lightweight models.