A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video
This study addresses the decision instability of vision-language models (VLMs) in autism screening and the scarcity of early-stage expert resources by proposing a grounding-aware, certainty-scoring architecture. By freezing VLM parameters, the method constructs a deterministic evidence layer through timestamped event extraction and textual confidence calibration, effectively decoupling perception from decision-making. Furthermore, a weighted evidence scoring algorithm is designed to achieve stable and interpretable risk stratification, enabling decisions to be decomposed into feature-level contributions with support for rescoring. Evaluated on home videos, the model attains an AUC of 85.1% and an accuracy of 86.0%, with a fully correct annotation rate of 74.4%. Notably, it yields zero false positives among typically developing children while significantly reducing prediction variance.