Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of large language models (LLMs) in long-context scenarios, where they predominantly perform shallow retrieval and struggle with complex, deep deductive reasoning. To this end, this work proposes ProloNg, the first long-horizon Prolog-based deductive reasoning benchmark. By leveraging Prolog to generate logic puzzles of varying depths combined with large-scale synthetic data techniques, it systematically evaluates the long-range reasoning capabilities of frontier models. The findings reveal a quantified relationship between reasoning depth and performance degradation, demonstrating that most models experience a precipitous decline to random-level accuracy when the reasoning depth exceeds ten steps. These results profoundly expose the fundamental limitations of current LLMs in long-horizon logical reasoning.
📝 Abstract
Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.
Problem

Research questions and friction points this paper is trying to address.

long-horizon reasoning
deductive reasoning
large language models
Prolog
reasoning depth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Horizon Reasoning
Deductive Logic
Prolog
Synthetic Benchmark
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hadeel Al-Negheimish
King Saud University, Massachusetts Institute of Technology
J
Jasna Ilieva
Massachusetts Institute of Technology
Yoon Kim
Yoon Kim
Associate Professor, MIT
Machine LearningNatural Language ProcessingDeep Learning