🤖 AI Summary
This study addresses the limitations of existing evaluations of moral reasoning in large language models (LLMs), which are typically confined to utilitarian-deontological conflicts and lack sufficient scale, revealing a pervasive omission bias wherein models reverse decisions under altered framing. To address this, we construct a multi-ethical benchmark integrating five philosophical perspectives, leveraging LLMs to simulate diverse ethical personas for generating paired conflicting scenarios. We further propose a method to quantify framing-invariant bias based on philosophical disagreement signals and systematically evaluate various inference-time intervention strategies. Our findings indicate that while omission bias is prevalent, it diminishes with increasing model scale. Moreover, principle-prioritized interventions significantly enhance response consistency but may introduce novel action biases.
📝 Abstract
As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.