π€ AI Summary
This study addresses the scarcity and high cost of expert annotations in counseling dialogue quality assessment by proposing a scale-anchored LLM-as-a-Judge framework. The approach leverages an ensemble of small-scale open-source models to extract session-level construct scores, integrating nonverbal bimodal features to enable cross-domain transfer prediction. Our findings demonstrate that cross-domain training outperforms within-domain training, revealing how linguistic features encode expert impressions while quantifying the impact of recording settings. Experimentally, the cross-domain nested Spearman correlation reaches 0.54, surpassing the within-domain baseline of 0.48, with single-judge performance improved to 0.41. These results establish a low-cost, highly reliable paradigm for automated evaluation.
π Abstract
Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $Ο= 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.