🤖 AI Summary
This work addresses the limitation of existing evaluation benchmarks that oversimplify societal risk as a single scalar mean, thereby neglecting its multidimensionality, distributional structure, and tail extremities. To overcome this, the authors propose the SHARP framework, which models social harm as a multidimensional random variable encompassing bias, fairness, ethical considerations, and epistemic reliability. Risk aggregation is performed via an additive cumulative log-risk formulation, and tail-sensitive statistics—particularly Conditional Value-at-Risk at the 95th percentile (CVaR95)—are introduced for robust assessment. Experiments across eleven state-of-the-art large language models reveal that models with similar average risk exhibit more than twofold differences in tail exposure and volatility, uncovering heterogeneous failure modes invisible to conventional scalar metrics and demonstrating SHARP’s unique capacity to identify high-risk behaviors.
📝 Abstract
Large language models (LLMs) are increasingly deployed in high-stakes domains, where rare but severe failures can result in irreversible harm. However, prevailing evaluation benchmarks often reduce complex social risk to mean-centered scalar scores, thereby obscuring distributional structure, cross-dimensional interactions, and worst-case behavior. This paper introduces Social Harm Analysis via Risk Profiles (SHARP), a framework for multidimensional, distribution-aware evaluation of social harm. SHARP models harm as a multivariate random variable and integrates explicit decomposition into bias, fairness, ethics, and epistemic reliability with a union-of-failures aggregation reparameterized as additive cumulative log-risk. The framework further employs risk-sensitive distributional statistics, with Conditional Value at Risk (CVaR95) as a primary metric, to characterize worst-case model behavior. Application of SHARP to eleven frontier LLMs, evaluated on a fixed corpus of n=901 socially sensitive prompts, reveals that models with similar average risk can exhibit more than twofold differences in tail exposure and volatility. Across models, dimension-wise marginal tail behavior varies systematically across harm dimensions, with bias exhibiting the strongest tail severities, epistemic and fairness risks occupying intermediate regimes, and ethical misalignment consistently lower; together, these patterns reveal heterogeneous, model-dependent failure structures that scalar benchmarks conflate. These findings indicate that responsible evaluation and governance of LLMs require moving beyond scalar averages toward multidimensional, tail-sensitive risk profiling.