🤖 AI Summary
This study addresses the decision instability of vision-language models (VLMs) in autism screening and the scarcity of early-stage expert resources by proposing a grounding-aware, certainty-scoring architecture. By freezing VLM parameters, the method constructs a deterministic evidence layer through timestamped event extraction and textual confidence calibration, effectively decoupling perception from decision-making. Furthermore, a weighted evidence scoring algorithm is designed to achieve stable and interpretable risk stratification, enabling decisions to be decomposed into feature-level contributions with support for rescoring. Evaluated on home videos, the model attains an AUC of 85.1% and an accuracy of 86.0%, with a fully correct annotation rate of 74.4%. Notably, it yields zero false positives among typically developing children while significantly reducing prediction variance.
📝 Abstract
Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0\% accuracy, and F$_1$ 71.8 over three runs. It labels 74.4\% of clips correctly in every run (60.5\% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.