Score
The capability (often modeled in AI systems) to infer human intentions, social norms, and contextual cues from multimodal inputs and to generate diverse plausible hypotheses from sparse descriptions. It also includes conceptually grounding and defining higher-order notions like second-order bias to support robust, human-centered inference and decision-making.
As AI capabilities surpass human performance, human feedback becomes increasingly unreliable, leading to the failure of scalable oversight. Method: We formally model the role of human evaluators’ belief structures in value inference for the first time, introduce the “belief coverage” relaxation framework, and establish a novel paradigm for constructing coverage-aware belief models using large language models. Contribution: Theoretically, we derive sufficient conditions under which belief uncertainty vanishes. Practically, we provide formal guarantees and a viable pathway for robust supervision without requiring precise prior beliefs—thereby significantly enhancing the scalability and trustworthiness of alignment for high-capability AI systems.
This study addresses the gap in understanding how multimodal large language models (MLLMs) reason about social norms in complex, real-world scenarios that integrate both textual and visual information. While existing normative reasoning approaches predominantly rely on symbolic logic and struggle with multimodal social contexts, the norm comprehension capabilities of MLLMs remain underexplored. The authors present the first systematic evaluation of five leading MLLMs—including GPT-4o and Qwen-2.5VL—on a dataset of 30 textual and 30 visual social stories, benchmarking their performance against human judgments. Results reveal that all models perform significantly better on textual than visual norm inference, with GPT-4o achieving the strongest overall results and Qwen-2.5VL emerging as the top-performing open-source model. Nevertheless, substantial limitations persist across all models when handling nuanced or complex social norms, highlighting both the promise and the challenges of current MLLMs in multimodal normative reasoning.
This paper addresses the challenge of modeling capability disparities among heterogeneous decision-makers—human experts and AI models—in collaborative decision-making. We propose the first unified capability modeling framework: (1) representing human and AI decision capabilities via learnable capability vectors, (2) dynamically assigning context-aware weighted fusion coefficients for decision integration, and (3) introducing a learning-free global collaboration baseline as a zero-shot reference. Our method comprises capability embedding, context-sensitive weighted fusion, and multi-source decision ensemble. Evaluated on image classification and hate speech detection, our approach significantly outperforms state-of-the-art methods—particularly when non-expert participants exhibit relatively strong capabilities—yielding substantial gains in both accuracy and robustness. Empirical results validate that explicit capability modeling markedly enhances human-AI collaborative decision-making performance.
Current AI benchmarking relies heavily on static scores, which inadequately reflect true model capabilities and suffer from questionable reliability. To address this, we propose an inference-based evaluation paradigm grounded in capability theory—formulating capability assessment as a theory-driven statistical inference problem rather than a mere measurement task. Our approach innovatively integrates psychometrics, Bayesian inference, and uncertainty modeling to construct a rigorous framework that quantifies both sensitivity and sample-level uncertainty. We further design an adaptive sampling algorithm to reduce sample complexity. Experiments demonstrate that our method significantly improves the reliability and interpretability of evaluations while reducing required sample size by over 50%. This work establishes a novel, principled paradigm for trustworthy AI capability assessment.
This study challenges the long-standing view that embodied social experience is necessary for acquiring cultural cognition, asking whether large language models (LLMs) can acquire human everyday social norms solely through statistical learning over linguistic data. Method: Using 555 real-world social scenarios, we evaluated GPT-4.5, GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4 on their ability to predict population-level appropriateness ratings of social behaviors along a continuous scale, benchmarking model outputs against large-scale human judgments. Contribution/Results: GPT-4.5 outperformed all individual human raters (100th percentile) in predicting population means; other models also significantly exceeded ≥96% of human participants. This constitutes the first empirical demonstration that purely text-based statistical learning suffices for high-fidelity modeling of social norms—highlighting language’s robust capacity as a primary medium for cultural transmission and norm representation.
Current large language models lack a perceptible sense of “mindedness” in extended dialogue, often appearing flat and devoid of inner life. This work proposes a “dimensional integrity” framework centered on four first-person behavioral stances—time, truth, entropy, and love—grounded in empirical evidence of human cognition, complemented by observable behavioral layers such as proactivity and conversational rhythm. Rather than prioritizing task performance, the framework enhances the perceived mindedness of artificial interlocutors. It shifts the focus of credible AGI development from capability to dimensional integrity, clearly distinguishing perception engineering from theories of machine consciousness. A prototype implementing temporal dimensionality through behavioral modeling and rhythm control is presented, alongside six falsifiable predictions supporting preregistered experiments. Key behavioral features have already been deployed in a production-level companion application.
This work addresses the challenge in human-agent collaboration where human cognitive biases often lead to inaccurate beliefs about the agent’s knowledge state, thereby impeding effective coordination. The paper presents the first approach that integrates second-order theory of mind (ToM-2) with explicit modeling of cognitive biases within an interactive partially observable Markov decision process (I-POMDP) framework. This enables the agent to infer not only the human’s mistaken beliefs about its own mental state but also the underlying heuristic mechanisms driving those biases, allowing it to generate targeted, adaptive feedback. Experimental results demonstrate that the proposed method significantly increases the informational value of teaching actions provided by the agent, and user evaluations confirm that the generated feedback is perceived as more useful, offering a novel pathway toward cognitively aligned human-agent collaboration.
This study demonstrates that large language models do not merely reflect human cognitive biases implicitly during alignment but systematically amplify them and propagate these amplified biases back to humans, thereby threatening fairness and safety in decision-making. Challenging the prevailing assumption that AI bias is a passive mirror of human prejudice, this work establishes for the first time that such models act as active amplifiers of bias—a phenomenon that intensifies across successive model generations. Drawing on social cognitive theory and behavioral analyses of cross-generational models, the research develops a diagnostic framework that elucidates the mechanisms underlying the隐蔽ness, amplification, and transmissibility of AI-induced bias. Building on these insights, the study proposes a multi-layered intervention strategy spanning diagnostic, regulatory, and operational dimensions to mitigate the adverse societal impacts of AI bias on human judgment and policy decisions.
This study critically examines the assumption that early AI intervention enhances collaborative sensemaking, arguing that premature provision of AI-generated insights may lead users to uncritically adopt suggestions, thereby undermining their capacity for independent interpretation and validation. Integrating perspectives from cognitive science and human-computer interaction, the research employs qualitative analysis and human-subject experiments to systematically uncover the risks of cognitive overreliance on AI outputs during the initial stages of meaning construction. It further identifies underlying psychological and contextual factors that predispose users to defer to AI interpretations. Building on these findings, the work proposes three key reflective questions to guide the design of more responsible AI-assisted systems, offering a theoretical foundation to mitigate cognitive biases and foster more autonomous human reasoning.
Current vision foundation models lack reliable and unified evaluation standards for human interpretability, particularly in high-stakes scenarios. This work proposes the first quantifiable and comparable assessment framework that integrates psychophysical experiments—measuring localization and naming capabilities—with features extracted via sparse autoencoders and a chance-calibrated scoring mechanism based on random baselines, enabling interpretability measurement on a unified scale. Empirical analysis across six vision Transformers and over 15,000 human responses reveals that contemporary foundation models generally exhibit lower interpretability than supervised counterparts. Crucially, interpretability is not determined by overall model capability but hinges on the locality of feature activations and their alignment with coarse-grained semantic concepts.