🤖 AI Summary
This work addresses the challenge that large language models (LLMs) struggle to perform precise and exhaustive mathematical reasoning about program behavior, while existing benchmarks either lack real-world software relevance or semantic rigor. To bridge this gap, we propose CodeLogician—the first neuro-symbolic agent that integrates an LLM with ImandraX, an industrial-grade automated reasoning engine—enabling deep semantic analysis of software logic through LLM-generated explicit formal models. We introduce Code-Logic-Bench, a novel benchmark that systematically evaluates semantic reasoning capabilities across dimensions such as program state space and control flow. Experimental results demonstrate that our approach improves accuracy by 41–47 percentage points over pure LLM baselines, substantially advancing the state of automated, rigorous software understanding.
📝 Abstract
Large Language Models (LLMs) have shown strong performance on code understanding tasks, yet they fundamentally lack the ability to perform precise, exhaustive mathematical reasoning about program behavior. Existing benchmarks either focus on mathematical proof automation, largely disconnected from real-world software, or on engineering tasks that do not require semantic rigor. We present CodeLogician, a neurosymbolic agent for precise analysis of software logic, integrated with ImandraX, an industrial automated reasoning engine deployed in financial markets and safety-critical systems. Unlike prior approaches that use formal methods primarily to validate LLM outputs, CodeLogician uses LLMs to construct explicit formal models of software systems, enabling automated reasoning to answer rich semantic questions beyond binary verification outcomes. To rigorously evaluate mathematical reasoning about software logic, we introduce code-logic-bench, a benchmark targeting the middle ground between theorem proving and software engineering benchmarks. It measures reasoning correctness about program state spaces, control flow, coverage constraints, and edge cases, with ground truth defined via formal modeling and region decomposition. Comparing LLM-only reasoning against LLMs augmented with CodeLogician, formal augmentation yields substantial improvements, closing a 41-47 percentage point gap in reasoning accuracy. These results demonstrate that neurosymbolic integration is essential for scaling program analysis toward rigorous, autonomous software understanding.