counterfactual benchmark design

Designs and builds benchmark suites and evaluation protocols that test models under counterfactual or parallel-world interventions by creating tasks or simulated "worlds" with altered causal laws, relations, priors, or off-path scenarios. Develops and applies quantitative and qualitative evaluation methods and metrics—such as prior resistance rate and reasoning retention rate—to measure models' ability to generalize to novel relations, resist default priors, retain correct reasoning, and reveal slips back to standard relations.

counterfactualbenchmarkdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

CounterBench: A Benchmark for Counterfactuals Reasoning in Large Language Models

Feb 16, 2025
YC
Yuefei Chen
🏛️ Rutgers University | Case Western Reserve University

Large language models (LLMs) exhibit significant deficiencies in formal counterfactual reasoning, yet no rigorous benchmark exists to assess this capability. Method: We introduce CounterBench—the first causal-rule-based counterfactual reasoning benchmark (1K instances)—covering diverse causal graph structures and semantic confounding variants. We propose CoIn, a novel inference paradigm integrating iterative reasoning, backtracking search, and prompt-guided exploration of the counterfactual solution space. Our evaluation framework is grounded in formal causal rules, and we publicly release both the dataset and evaluation code. Contribution/Results: Experiments reveal that state-of-the-art LLMs perform near-chance on CounterBench (~20% accuracy), while CoIn boosts their performance by an average of +35.2%, with consistent cross-model generalization. This work establishes a principled benchmark and methodology for quantitatively evaluating and controllably enhancing causal reasoning capabilities in LLMs.

Evaluate LLMs in counterfactual reasoningIntroduce CounterBench benchmark datasetPropose CoIn to improve LLM reasoning

This work addresses the absence of evaluation benchmarks for assessing large language models’ capacity to perform intervention reasoning and causal research design in real-world social systems. We introduce InterveneBench, the first end-to-end benchmark constructed from 744 empirical social science papers, which challenges models to infer policy intervention effects and articulate identification assumptions without access to predefined causal graphs. To enhance model performance on this task, we propose STRIDES, a multi-agent collaborative framework that substantially improves causal study design capabilities. Experimental results demonstrate that state-of-the-art large language models exhibit limited proficiency on InterveneBench, whereas STRIDES significantly outperforms existing approaches.

benchmarkingcausal inferenceintervention reasoning

Benchmarking Counterfactual Image Generation

Mar 29, 2024
TM
Thomas Melistas
🏛️ National and Kapodistrian University of Athens | Archimedes/Athena RC | The University of Edinburgh | Imperial College London | Spotify | The University of Essex

This work addresses the problem of counterfactual image generation methods producing edits that violate intrinsic causal logic within images. To this end, we propose the first unified benchmark framework for systematically evaluating causal consistency and visual fidelity—without requiring ground-truth labels. Methodologically, we integrate structural causal models (SCMs) with hierarchical variational autoencoders (Hierarchical VAEs), establishing a multi-model, multi-dataset, and multi-causal-graph evaluation paradigm. We further introduce customized metrics, including causal consistency, to quantify alignment with underlying causal mechanisms. Experimental results demonstrate that Hierarchical VAEs significantly outperform GAN- and flow-based baselines on both natural and medical imaging domains, highlighting their generalizability across modalities. The framework is released as an open-source, extensible Python benchmark package, enabling community-wide reproducibility, validation, and extension.

Causal ConsistencyCounterfactual Image GenerationPerformance Evaluation

Latest Papers

What's happening recently
View more

This study addresses the disconnect between academic research and industrial practice in treatment effect estimation, where prevailing evaluation paradigms hinder real-world applicability. Through a large-scale empirical analysis, we systematically compare diverse meta-learners, base learners, and specialized causal models across semi-synthetic benchmarks and real-world datasets. Our findings reveal a pronounced inconsistency between counterfactual and observable performance metrics, and demonstrate that model rankings derived from semi-synthetic data fail to generalize to real settings. Notably, simple meta-learners paired with strong base models consistently outperform purpose-built causal models on real data, underscoring the critical importance of validation on real-world outcomes and observable metrics. These results challenge the dominant reliance on semi-synthetic evaluations and call for a paradigm shift toward more empirically grounded assessment protocols.

counterfactual metricsevaluation gapreal-world datasets

This work addresses the lack of effective evaluation of causal intervention responsiveness in existing video generation models, as conventional benchmarks focus solely on the plausibility of individual videos and fail to assess physical consistency. The study proposes the first causal world model evaluation framework tailored for embodied scenarios, constructing 319 prompt pairs from real-world nuScenes and DROID datasets that differ only in a single physical variable. A four-dimensional scoring mechanism—Adherence, Physics, Environment, and Outcome (APEO)—is introduced to systematically evaluate output consistency under interventions. Experiments reveal that even state-of-the-art models achieve only 52% pairwise accuracy, while open-source models average around 28%. Performance strongly correlates with the visual salience of interventions, with subtle changes yielding success rates as low as 14.2%, underscoring significant limitations in current models’ causal reasoning capabilities.

causal reasoningembodied scenariosphysical consistency

This study investigates whether text-to-image (T2I) generation models possess genuine causal reasoning capabilities or merely rely on statistical visual-linguistic associations. To this end, the authors introduce CF-World, the first counterfactual benchmark for T2I models, featuring a three-tiered task hierarchy designed to evaluate generation performance under conditions that violate real-world priors. They also propose CF-Eval, an automated evaluation framework based on vision-language models (VLMs), and define two novel metrics—Prior Resistance Rate and Reasoning Retention Rate—to systematically quantify a model’s ability to resist ingrained commonsense priors and retain causal reasoning. Experimental results demonstrate that state-of-the-art T2I models exhibit significant performance degradation in counterfactual scenarios, revealing their strong dependence on co-occurrence patterns in training data and a lack of true causal understanding.

causal reasoningcommonsense priorscounterfactual reasoning

Existing benchmarks struggle to comprehensively evaluate the causal reasoning capabilities of data science agents, often lacking either realistic analytical scenarios or rigorous causal generative mechanisms. To address this gap, this work proposes CausalDS, a novel benchmark that systematically integrates structural causal models (SCMs), realistic synthetic data, and natural language narratives. CausalDS spans Pearl’s three levels of causal reasoning and is embedded within an authentic data science workflow, supporting tool invocation, programming interfaces, and uncertainty quantification. By grounding scenario generation in real-world data distributions, the benchmark mitigates the “causal parroting” problem and explicitly incorporates active abstention as a core evaluation dimension, thereby enabling a holistic assessment of agents’ abilities in causal inference, practical execution, and principled refusal to answer.

benchmarkingcausal reasoningdata-science agents

Existing evaluation methods for world models often misinterpret state copying as accurate prediction in environments with sparse changes, failing to assess a model’s understanding of executable causal relationships. To address this, this work proposes ScratchWorld—a novel offline diagnostic benchmark built upon the Scratch virtual machine—that enables structured replay of state transitions, latent variables, causal trajectories, and counterfactual outcomes through multimodal inputs and diverse diagnostic tasks. The study introduces a value-aware Changed-Field F1 metric that effectively distinguishes mere state replication from genuine causal reasoning. Evaluations on 659 samples reveal that even the best-performing among seven state-of-the-art models achieves only 13.8% Changed-Field F1, whereas a trivial copy strategy attains 98.0% full-state accuracy yet scores 0.0% on the Changed-Field F1, highlighting a fundamental deficiency in current models’ ability to adhere to executable rules and demonstrating the efficacy of the proposed evaluation framework.

causal reasoningevaluation benchmarkexecutable consequences

Hot Scholars

BG

Ben Glocker

Imperial College London
Medical Image AnalysisComputer VisionMachine Learning
DL

Dongha Lee

Yonsei University
Data miningInformation retrievalNatural language processing
KR

Keyun Ruan

Alphabet Inc; www.ruankeyun.com
Risk EconomicsDisruptive TechnologySocietal RiskDigital Risk Measurement