🤖 AI Summary
This study addresses the oversight in existing LLM evaluations regarding models’ impact on users’ real-world interpersonal relationships and the lack of mechanisms to measure their support orientation. To this end, this work introduces the concept of “relational orientation” and proposes the RELATE auditing framework, which operationalizes, for the first time, inward (AI-dependent) and outward (encouraging human connection) scaffolding dimensions. Grounded in psychological theory, it establishes sentence-level auditing criteria and enables quantitative evaluation through personality-conditioned simulated dialogues coupled with a dual-LLM auto-judging technique. Experimental results demonstrate that models’ inward tendencies intensify as conversations progress, while user interaction styles significantly moderate the proportion of outward support. These findings provide effective signals for auditing the social influence of large language models.
📝 Abstract
Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.