A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decision instability of vision-language models (VLMs) in autism screening and the scarcity of early-stage expert resources by proposing a grounding-aware, certainty-scoring architecture. By freezing VLM parameters, the method constructs a deterministic evidence layer through timestamped event extraction and textual confidence calibration, effectively decoupling perception from decision-making. Furthermore, a weighted evidence scoring algorithm is designed to achieve stable and interpretable risk stratification, enabling decisions to be decomposed into feature-level contributions with support for rescoring. Evaluated on home videos, the model attains an AUC of 85.1% and an accuracy of 86.0%, with a fully correct annotation rate of 74.4%. Notably, it yields zero false positives among typically developing children while significantly reducing prediction variance.
📝 Abstract
Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0\% accuracy, and F$_1$ 71.8 over three runs. It labels 74.4\% of clips correctly in every run (60.5\% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.
Problem

Research questions and friction points this paper is trying to address.

Autism spectrum disorder screening
Vision-language models
Prediction instability
Home video analysis
Deterministic decision-making
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deterministic Evidence Layer
Vision-Language Models
Autism Spectrum Disorder Screening
Grounded Perception Constraints
Weight-of-Evidence Scorer
🔎 Similar Papers
No similar papers found.