CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入CogGym框架,旨在解决大规模比较人类与机器认知的问题,采用统一的任务无关实验标记语言促进可重复性比较。
📝 Abstract
Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.
Problem

Research questions and friction points this paper is trying to address.

human and machine cognition
scalable comparison
cognitive science
Innovation

Methods, ideas, or system contributions that make the work stand out.

cognitive science
scalable framework
human-in-the-loop pipeline
Experiment Markup Language (EML)
commonsense reasoning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Lance Ying
Lance Ying
Harvard University / MIT
Computational CognitionSocial IntelligenceHuman-like IntelligenceMultimodal Reasoning
J
Jinzhou Wu
Cornell University
Y
Yingshan Susan Wang
Massachusetts Institute of Technology
S
Shivam Aarya
Johns Hopkins University
Luca M. Schulze Buschoff
Luca M. Schulze Buschoff
Helmholtz Munich
Machine LearningCognitive Science
H
Harry Chen
Massachusetts General Hospital
Katherine M. Collins
Katherine M. Collins
Machine Learning PhD Student at the University of Cambridge
Cognitive ScienceMachine LearningBayesian StatisticsHuman-AI Interaction
A
Andrea de Varda
Massachusetts Institute of Technology
S
Shuhao Fu
Santa Fe Institute
Sean Dae Houlihan
Sean Dae Houlihan
Dartmouth College
Cognitive ScienceSocial CognitionEmotionArtificial Intelligence
Akshay K. Jagadish
Akshay K. Jagadish
Princeton University
Cognitive ScienceLarge Language ModelsMeta-learningIn-context Learning
Guangyuan Jiang
Guangyuan Jiang
MIT
cognitive scienceartificial intelligencecomputational linguistics
Samuel Kiegeland
Samuel Kiegeland
Graduate Student, ETH Zurich
T
Tetsu Kurumisawa
Massachusetts Institute of Technology
R
Rongzhi Liu
University of Chicago
Ryan Liu
Ryan Liu
PhD Student in Computer Science, Princeton University
Large Language ModelsComputational Cognitive ScienceNLP Applications
N
Ningshan Ma
Massachusetts Institute of Technology
K
Kathryn McGregor
Princeton University
Younes Strittmatter
Younes Strittmatter
Princeton University
Polina Tsvilodub
Polina Tsvilodub
University of Tübingen
cognitive sciencepragmaticsNLP
J
Jacob Hoover Vigly
CHI-FRO
S
Sarah Wu
Stanford University
E
Enjie Xu
University of California, Los Angeles
Y
Yiling Yun
University of California, Los Angeles
Kelsey Allen
Kelsey Allen
Research Scientist, DeepMind
Artificial IntelligenceCognitive ScienceComputational NeuroscienceCollective BehaviorPhysics