🤖 AI Summary
This study addresses the persistent stagnation in automatic speech deception detection, where increasing model complexity has failed to yield performance breakthroughs. Through a systematic review, meta-analysis, and nested multilevel statistical modeling, this work provides the first quantitative validation of performance ceilings in this field. The analysis reveals a pooled accuracy of 74.4%, exposing an "accuracy ceiling" between 70% and 75%, with only a minority of studies demonstrating verifiable ground truth and independent evaluation. Challenging the prevailing assumption that large-scale models can substantially enhance detection capabilities, these findings indicate that data quality and methodological rigor are more critical than model complexity. By confirming that existing paradigms struggle to overcome current bottlenecks, this research offers evidence-based guidance for future investigations in automated deception detection.
📝 Abstract
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines. We systematically reviewed 25 years of research (289 reports, 6,136 classification models) and meta-analyzed 3,653 models nested within 97 datasets. Pooled accuracy was 74.4% (95% CI: 71.2%-77.4%) with substantial heterogeneity. Accuracy was driven by methodological quality (ground truth, data source, class balance, evaluation procedure) more than by model complexity: the adoption of embeddings and large language models has not translated into improved predictive performance. Only 12.46% of reports used data with verifiable ground-truth, and only 23.96% of models were evaluated on independent data. The pooled accuracy aligns with meta-analyses of manual approaches, suggesting a ceiling of 70-75%, unlikely to be lifted by current research conventions.