Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the scarcity and high cost of expert annotations in counseling dialogue quality assessment by proposing a scale-anchored LLM-as-a-Judge framework. The approach leverages an ensemble of small-scale open-source models to extract session-level construct scores, integrating nonverbal bimodal features to enable cross-domain transfer prediction. Our findings demonstrate that cross-domain training outperforms within-domain training, revealing how linguistic features encode expert impressions while quantifying the impact of recording settings. Experimentally, the cross-domain nested Spearman correlation reaches 0.54, surpassing the within-domain baseline of 0.48, with single-judge performance improved to 0.41. These results establish a low-cost, highly reliable paradigm for automated evaluation.
πŸ“ Abstract
Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $ρ= 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.
Problem

Research questions and friction points this paper is trying to address.

counseling-quality assessment
cross-domain transfer
data scarcity
LLM judges
expert impression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Instrument-Anchored LLM Judges
Cross-Domain Transfer
Counseling-Quality Assessment
Ensemble Learning
Expert Rating Constructs
πŸ’Ό Related Jobs
No related jobs found.