CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
The paper frames understanding and modeling human intelligence as parallel goals shared by AI and cognitive science. As AI systems become more capable, it asks where model responses resemble human responses and where they systematically diverge, noting that the breadth and diversity of human tasks make scalable, rigorous human–model comparison difficult.
CogGym addresses this with a scalable, unified framework grounded in cognitive science for comparing model and human behavior on matched experimental trials. It uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For its initial release, the authors curate and standardize 258 cognitive experiments from 100 papers focused on human commonsense reasoning, then evaluate 50 large language models against human responses.
The results show a clear scaling trend: larger and more recent AI models better reproduce human judgments. However, improvement on these commonsense reasoning tasks is considerably slower than gains seen on formal-reasoning benchmarks such as math and coding. Model–human fit also remains well below human split-half reliability—$R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video—with the best models reaching only $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments.
The authors intend CogGym to be a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.