natural-language rule grounding

Design and build systems that translate rules expressed in natural language into typed intermediate representations (IR), producing formal predicates, slots, and structured nodes from text. These systems identify and normalize entities and actions, resolve temporal and conditional constraints, and output predicate- or IR-based encodings suitable for downstream enforcement, verification, or reasoning.

natural-languagerulegrounding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of compiling natural language queries into backend query languages in document-centric, hybrid, and heterogeneous data environments, where semantic intent is often ambiguous or incomplete. The authors propose the NLIQ framework, which introduces a “goal sufficiency” criterion to classify queries according to their semantic determinacy. It emphasizes that when intermediate goals must be dynamically constructed, intermediate representations should serve as core semantic objects rather than mere syntactic intermediaries. Through conceptual analysis, case modeling, and formal categorization, the study establishes a unified query paradigm that integrates goal recognition, intermediate representation design, and heterogeneous execution. This framework provides a theoretical foundation for natural language querying in complex data settings and opens new research directions in semantic goal construction, heterogeneous compilation, and answer generation.

heterogeneous data environmentsintermediate representationnatural language querying

Can Large Language Models Understand Intermediate Representations?

Feb 07, 2025
HJ
Hailong Jiang
🏛️ Kent State University | Huazhong University of Science and Technology | Pacific Northwest National Laboratory | Chongqing University

This work investigates large language models’ (LLMs) capacity to comprehend compiler intermediate representations (IR), specifically evaluating their performance on four core tasks: control-flow graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Through systematic multi-model benchmarking (GPT-4, LLaMA 3.1, Gemma 2, etc.), a curated structured IR dataset, task-specific prompt engineering, and fine-grained error attribution, the study provides the first empirical evidence of fundamental limitations in LLMs’ IR understanding—particularly in CFG reconstruction (accuracy <42%) and execution reasoning (error rate 68%). Methodologically, it introduces a dual-path enhancement paradigm: (1) IR-domain fine-tuning and (2) explicit control-flow modeling. Experimental results demonstrate that targeted fine-tuning improves task performance by up to 31.5%, establishing a foundational framework for advancing LLM-based IR analysis.

Challenges in control flow and execution reasoning.LLMs' understanding of Intermediate Representations.Need for IR-specific enhancements in LLMs.

This work addresses the challenge of format and semantic errors produced by small language models when translating natural language into first-order logic (FOL), which undermines the reliability of symbolic reasoning. To mitigate this, the authors propose a staged incremental reasoning framework: first, a large language model synthesizes training data to supervise fine-tuning of the small model; then, the translation process is decoupled into predicate generation and FOL formulation stages. An external verification module is introduced to detect and correct predicate arity errors, thereby enhancing translation accuracy. Evaluated on four logical reasoning benchmarks, the approach significantly reduces error rates, improves predicate coverage, and boosts overall reasoning performance, bringing small models closer to reliable, verifiable symbolic reasoning systems.

first-order logiclanguage modelslogical reasoning

Current information retrieval systems struggle with complex queries requiring logical constraints, multi-step reasoning, and evidence synthesis, primarily due to a lack of structured reasoning capabilities. This work proposes the first unified framework for structured reasoning in information retrieval, systematically integrating interdisciplinary approaches—including large language model reasoning strategies, neuro-symbolic systems, probabilistic and Bayesian methods, geometric representations, and energy-based models—to elucidate their inherent trade-offs and complementary mechanisms. By bridging disciplinary boundaries, the framework clarifies the central role of retrieval within general-purpose reasoning systems and provides researchers with conceptual tools and practical guidance to advance the development of verifiable, structured reasoning architectures.

evidence synthesisinformation retrievallogical constraints

From Natural Language to SQL: Review of LLM-based Text-to-SQL Systems

Oct 01, 2024
AM
Ali Mohammadjafari
🏛️ University of Louisiana at Lafayette

This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.

Addressing computational and privacy challengesExploring LLM evolution in SQL systemsImproving SQL translation accuracy

Latest Papers

What's happening recently
View more

This work addresses the challenge large language models face in accurately translating natural language descriptions into optimization models, particularly when handling composite constraints and intricate business rules. To bridge this gap, the authors propose a Canonical Intermediate Representation (CIR) that decouples business rule logic from its mathematical implementation, serving as a semantic intermediary between natural language and formal optimization models. They further introduce a multi-agent R2C framework that parses input text, retrieves domain knowledge, generates CIR specifications, and subsequently instantiates them into mathematical optimization models. The paper establishes the first systematic benchmark for rule-to-constraint reasoning and demonstrates state-of-the-art performance: on a newly curated complex-rule benchmark, their method achieves 47.2% accuracy, matches or exceeds closed-source models like GPT-5 on existing benchmarks, and sets new best results on select tasks through a self-reflection mechanism.

code generationcomposite constraintslarge language models

This study addresses the challenge in regulated public procurement where bid validation must simultaneously ensure factual accuracy and legal auditability—a balance difficult to achieve with conventional approaches that struggle to integrate semantic understanding with rule-based interpretability. To bridge this gap, the authors propose a neuro-symbolic method that synergistically combines large language models (LLMs) with Logic Tensor Networks (LTNs). Specifically, LLMs are employed to extract semantic predicates from textual bid documents, which are then integrated into an LTN framework that encodes domain-specific regulatory rules to perform auditable, logic-driven inference. Experimental evaluation on real-world procurement documents demonstrates that the proposed approach not only maintains competitive performance but also significantly enhances decision interpretability and system modularity, thereby offering robust support for explainable artificial intelligence (XAI) in high-stakes regulatory contexts.

explainable AILogic Tensor Networksneurosymbolic AI

Existing methods for generating detection rules rely on specific input-output pairs and lack a unified framework. This work formalizes the task for the first time as a unified mapping from contextual and target-language inputs to detection rules, introducing UniRule—a novel framework that models semantic distance in a dual semantic projection space encompassing detection intent and detection logic to characterize optimal rules. UniRule integrates an agent-based RAG architecture to achieve generalization across diverse contexts and languages. Experimental results demonstrate that UniRule significantly outperforms pure large language model (LLM) approaches across twelve distinct scenarios, achieving a Bradley-Terry preference coefficient of 0.52, thereby validating its effectiveness and broad applicability.

detection rule generationinput-output couplingrule formalization

Existing formal methods incur high costs in specification construction and maintenance and lack scalability, making them ill-suited for verifying modern AI systems. This work proposes a Learning-Integrated Formal Reasoning (LIFR) framework that innovatively combines machine learning with formal verification: it employs natural language processing to automatically generate contracts, leverages graph matching and representation learning to achieve semantic alignment and cross-system reuse of verification artifacts, and establishes a rigorous semantic foundation grounded in Unifying Theories of Programming (UTP) and institution theory. By shifting formal verification from isolated proofs toward a cumulative, knowledge-driven paradigm, the LIFR framework substantially enhances automation and scalability while preserving formal rigor.

AI safetyformal verificationspecification synthesis

This work addresses key limitations of existing Text-to-SQL approaches, which struggle with complex reasoning, integration of domain knowledge, and hypothetical queries, while also incurring high deployment costs in enterprise settings. To overcome these challenges, the authors propose the IESR framework, which leverages a lightweight, non-finetuned large language model to perform information-enhanced structured reasoning. IESR decouples mathematical computation from SQL generation, integrates schema linking with semantic understanding, and introduces a Monte Carlo Tree Search (MCTS)-based multi-path reasoning mechanism coupled with trajectory consistency verification. Without any model fine-tuning, this approach achieves state-of-the-art performance on complex reasoning benchmarks, including LogicQA (24.28 EX) and Archer (37.28 EX), demonstrating significant gains in reasoning accuracy using only lightweight models.

complex reasoningdomain knowledgeenterprise deployment

Hot Scholars

JY

Jiayu Yang

The Australian National University
3D Computer Vision3D AIGC3D ReconstructionMulti-view Stereo
RZ

Rui Zhao

National University of Singapore
Computer VisionMultimodalVision and LanguageVirtual Humans
JZ

Jiwen Zhang

Fudan University
multimodal learningrobotics