🤖 AI Summary
The absence of systematic evaluation frameworks for large language models’ (LLMs) abstract causal reasoning capabilities hinders progress in understanding their philosophical and cognitive foundations.
Method: We propose the first standardized evaluation framework grounded in philosophical causal theory—particularly Lewis’s neuron diagrams—introducing, for the first time, a formally rigorous definition of generalized validity for neuron-diagram-based causation, thereby refuting the long-standing consensus that such definitions are inherently non-formalizable. Integrating formal philosophical modeling with zero-shot causal discrimination tasks, we empirically assess leading LLMs—including ChatGPT, DeepSeek, and Gemini.
Results: Our evaluation reveals that LLMs accurately resolve complex, long-debated philosophical causal cases, demonstrating nascent yet robust abstract causal reasoning. Beyond establishing a scalable, theory-informed causal competence benchmark, this work unveils a novel human–AI collaborative paradigm for advancing causal philosophy.
📝 Abstract
We propose a test for abstract causal reasoning in AI, based on scholarship in the philosophy of causation, in particular on the neuron diagrams popularized by D. Lewis. We illustrate the test on advanced Large Language Models (ChatGPT, DeepSeek and Gemini). Remarkably, these chatbots are already capable of correctly identifying causes in cases that are hotly debated in the literature. In order to assess the results of these LLMs and future dedicated AI, we propose a definition of cause in neuron diagrams with a wider validity than published hitherto, which challenges the widespread view that such a definition is elusive. We submit that these results are an illustration of how future philosophical research might evolve: as an interplay between human and artificial expertise.