MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过建立MIS-Bench评估多模态大语言模型在心理治疗人际技能评价中的可靠性,并提出MIS-RAFT方法以提高模型与专家评分的一致性。
📝 Abstract
Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of assessing psychotherapeutic interpersonal skills and introduce MIS-Bench, a Multimodal Interpersonal Skills (MIS) benchmark comprising 996 psychotherapy response videos annotated across 8 dimensions of Facilitative Interpersonal Skills. Across 9 MLLMs with multiple modality and prompting settings, we find that current models show only modest agreement with human experts, inconsistent gains from multimodal input, and limited benefits from reasoning-based prompting. To mitigate this gap, we propose MIS-RAFT, a regression-aware fine-tuning method inspired by RAFT and tailored to fine-grained interpersonal skill scoring at one-decimal precision. MIS-RAFT addresses the mismatch between autoregressive token prediction and scalar-valued expert assessment, significantly improving agreement with human ratings. Overall, MIS-Bench reveals a clear gap between general multimodal capability and expert-level interpersonal judgment, while MIS-RAFT offers a promising path toward more reliable model-based assessment.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Psychotherapeutic Interpersonal Skills
Expert Judgment
Benchmarking
Assessment Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

MIS-RAFT
regression-aware fine-tuning
MIS-Bench
interpersonal skills assessment
🔎 Similar Papers