Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) can supply sufficient lexical entailment knowledge to support formal proofs in logical natural language inference (NLI) systems. It introduces a novel task of explaining sentence-level entailment through lexical entailment relations and constructs a dedicated dataset to evaluate the capacity of LLMs to generate structured lexical explanations. These explanations are further validated within a neuro-symbolic framework integrated with the LangPro theorem prover to assess their contribution to formal reasoning. Experimental results demonstrate that this task remains highly challenging: the entailment relations generated by LLMs are only partially reliable, frequently restricted to specific instances, and offer limited substantive support for theorem proving. These findings reveal fundamental limitations of current LLMs in performing rigorous logical reasoning.
📝 Abstract
Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.
Problem

Research questions and friction points this paper is trying to address.

Natural Language Inference
Lexical Entailment
Large Language Models
Theorem Proving
Neuro-symbolic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Natural Language Inference
Neuro-symbolic
Lexical Entailment
Theorem Proving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jorryt de Jong
Utrecht University, the Netherlands
S
Stefan Moraca
Utrecht University, the Netherlands
E
Ettore Cesari
Utrecht University, the Netherlands
Lasha Abzianidze
Lasha Abzianidze
Assistant Professor, Utrecht University
Natural Language ProcessingComputational LinguisticsComputational SemanticsNatural Language Inference