π€ AI Summary
This study addresses the paradox wherein artificial intelligence enhances task performance yet fails to improve usersβ self-assessments, resulting in overconfidence and impaired metacognitive calibration. By distinguishing between performance augmentation and metacognitive augmentation, this work employs large-scale empirical studies, computational modeling, and statistical controls to compare the Dunning-Kruger effect across human-AI collaboration and independent task execution. The findings elucidate the mechanisms underlying calibration failure in AI-assisted contexts, demonstrating that AI-supported groups exhibit significantly greater confidence bias than those working independently. Based on these insights, the authors propose verification-supportive interface design recommendations, offering both a theoretical foundation and practical guidance for mitigating AI-induced metacognitive distortions.
π Abstract
AI assistance can improve performance without improving self-assessment. We report a study (N=366) comparing Human alone and Human+AI performance on reasoning tasks, for which the AI model is benchmarked on the same items. Participants estimated global and block performance and rated confidence in their answers. Human+AI achieved higher scores, but self-estimates tracked performance weakly. Average overestimation was similar across groups, covering individual errors. Across tasks, confidence distinguished correct from incorrect answers less accurately in the Human+AI group, while within-task differences remained uncertain. The Dunning-Kruger pattern was found in both groups, with a larger observed contrast in Human+AI. Controls for score noise reduced but did not eliminate the pattern, with the controlled group difference remaining inconclusive. An extended computational account describes global and block estimates. Our findings distinguish performance augmentation from metacognitive augmentation and motivate interfaces that support verification, communicate task-specific AI model performance, and help users evaluate the quality of their joint work rather than produce answers.