🤖 AI Summary
Current AI systems exhibit weak abstract reasoning capabilities under data scarcity and distributional shift. Method: We propose the first causal-augmented reasoning evaluation framework integrating observational, interventional, and counterfactual feedback, grounded in structural causal models (SCMs) to automatically generate diverse reasoning tasks. It supports multidimensional assessment—including few-shot prompting, in-context learning, program synthesis, and logical reasoning. Contribution/Results: By embedding a causal world model into abstract reasoning evaluation, our framework enables the first systematic measurement of high-level capabilities—causal discovery, program generation, and counterfactual reasoning—in language models. Evaluated across four distinct LLM scenarios, it significantly improves out-of-distribution generalization and robustness in abstract reasoning. This work establishes a novel paradigm for trustworthy AI reasoning evaluation.
📝 Abstract
Reasoning requires adaptation to novel problem settings under limited data and distribution shift. This work introduces CausalARC: an experimental testbed for AI reasoning in low-data and out-of-distribution regimes, modeled after the Abstraction and Reasoning Corpus (ARC). Each CausalARC reasoning task is sampled from a fully specified causal world model, formally expressed as a structural causal model. Principled data augmentations provide observational, interventional, and counterfactual feedback about the world model in the form of few-shot, in-context learning demonstrations. As a proof-of-concept, we illustrate the use of CausalARC for four language model evaluation settings: (1) abstract reasoning with test-time training, (2) counterfactual reasoning with in-context learning, (3) program synthesis, and (4) causal discovery with logical reasoning.