E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决自动视频面试评估中可解释性不足的问题,E-AVI框架通过提取多模态证据并结合维度条件注意机制来提高评分准确性。
📝 Abstract
Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.
Problem

Research questions and friction points this paper is trying to address.

Automated Video Interview
Multimodal Assessment
Evidence-Grounded
Inspectable Support
Numerical Predictions
Innovation

Methods, ideas, or system contributions that make the work stand out.

timestamped multimodal evidence
dimension-conditioned evidence attention
source-level embeddings
natural-language feedback
predictive performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoshen Wang
The Hong Kong Polytechnic University
D
Dongbo Che
The Hong Kong Polytechnic University
Z
Zeyi Xie
Independent Researcher
Y
Yuanjie Du
The Hong Kong Polytechnic University
S
Shicheng Hua
The Hong Kong Polytechnic University
Xingyu Wang
Xingyu Wang
Nanjing University of Posts and Telecommunications
NLP