🤖 AI Summary
This study addresses the challenge of evaluating the contextual adaptability of large language models (LLMs) in applying values within specialized scenarios by proposing a cross-domain value alignment framework. Methodologically, it introduces a novel "value × domain" matrix derived from regulatory documents, integrating multi-domain scenario simulations, factorial experiments, and action-rationale dual evaluation to systematically assess LLMs' context-sensitive execution of honesty, autonomy, and confidentiality principles in medical and legal settings. Results indicate that while the models achieve an average behavioral appropriateness of 95.6%, their accuracy in citing domain-specific justifications fluctuates significantly. This reveals a localized failure mechanism driven by risk sensitivity, thereby underscoring the necessity for fine-grained evaluations of contextual alignment in LLMs.
📝 Abstract
Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision is made. In this paper, we ask whether large language models (LLMs) appropriately adapt the application of a value across professional settings, while remaining consistent when contextual changes do not alter the relevant professional norm. To study this, we introduce ContextAdapt, an evaluation framework covering honesty, autonomy, and confidentiality across medicine, law, finance, and national security. Drawing on primary-source professional and regulatory documents, we construct a value x domain framework and use this to develop scenarios testing both default professional rules and recognised exceptions. We evaluate 12 LLMs on both the actions they recommend and the justifications they provide. In our main experiment, models achieve 95.6% mean appropriateness, although the use of the correct domain-specific justification varies substantially across models, from 25.6% to 76.9%. In a separate factorial experiment, explicitly naming the professional domain and changing the role of the model have limited effect on behaviour. Varying stakes, however, reveals severe but localised failures: in some cases, models alter their responses even though the underlying professional obligation remains unchanged. In particular, perceived severity appears to act as a cue for disclosure across both honesty and confidentiality scenarios. These results show that evaluating value alignment requires us to consider not only whether models follow abstract principles, but whether they apply them appropriately across different contexts.