🤖 AI Summary
This study challenges the long-standing view that embodied social experience is necessary for acquiring cultural cognition, asking whether large language models (LLMs) can acquire human everyday social norms solely through statistical learning over linguistic data. Method: Using 555 real-world social scenarios, we evaluated GPT-4.5, GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4 on their ability to predict population-level appropriateness ratings of social behaviors along a continuous scale, benchmarking model outputs against large-scale human judgments. Contribution/Results: GPT-4.5 outperformed all individual human raters (100th percentile) in predicting population means; other models also significantly exceeded ≥96% of human participants. This constitutes the first empirical demonstration that purely text-based statistical learning suffices for high-fidelity modeling of social norms—highlighting language’s robust capacity as a primary medium for cultural transmission and norm representation.
📝 Abstract
A fundamental question in cognitive science concerns how social norms are acquired and represented. While humans typically learn norms through embodied social experience, we investigated whether large language models can achieve sophisticated norm understanding through statistical learning alone. Across two studies, we systematically evaluated multiple AI systems' ability to predict human social appropriateness judgments for 555 everyday scenarios by examining how closely they predicted the average judgment compared to each human participant. In Study 1, GPT-4.5's accuracy in predicting the collective judgment on a continuous scale exceeded that of every human participant (100th percentile). Study 2 replicated this, with Gemini 2.5 Pro outperforming 98.7% of humans, GPT-5 97.8%, and Claude Sonnet 4 96.0%. Despite this predictive power, all models showed systematic, correlated errors. These findings demonstrate that sophisticated models of social cognition can emerge from statistical learning over linguistic data alone, challenging strong versions of theories emphasizing the exclusive necessity of embodied experience for cultural competence. The systematic nature of AI limitations across different architectures indicates potential boundaries of pattern-based social understanding, while the models' ability to outperform nearly all individual humans in this predictive task suggests that language serves as a remarkably rich repository for cultural knowledge transmission.