Artificial intelligences and human scientists exhibit complementary strengths in theory building
This study investigates the comparative efficacy of large language models (LLMs) and human scientists in social science theory construction, prediction, and revision. By evaluating 25 LLMs against 73 scholars on gender and racial inequality topics, the research employs large-scale blind reviews, empirical predictive accuracy metrics, and Bayesian belief updating models for quantitative assessment. Results indicate that individual LLMs outperform most humans in theory generation, excelling at handling theoretical complexity yet exhibiting ornamental redundancy. Conversely, aggregated human judgments demonstrate greater creativity, superior predictive efficiency, and more precise error-driven belief updating. These findings reveal complementary mechanisms between artificial intelligence and human cognition in scientific discovery, providing empirical foundations for AI-assisted social science research.