ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of multimodal infant autism screening data characterized by longitudinal coverage, expert annotations, and privacy protection. To this end, it introduces the ASD-FEAT dataset, which integrates facial, eye-tracking, and audio features, and develops an end-to-end computer vision pipeline for the automated identification of social interaction behavioral markers. The core innovation lies in pioneering the combination of longitudinal developmental data with privacy-aware multimodal representations, alongside the introduction of intra-visit peer contrastive signals. Experimental results demonstrate that the proposed automated pipeline achieves an accuracy of 76.2% and an AUROC of 0.82. Furthermore, incorporating the contrastive signals yields a Matthews Correlation Coefficient of 0.49, significantly surpassing manual coding baselines.
📝 Abstract
Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from video recordings of infant-adult interaction sessions. The key contribution of ASD-FEAT is the combination of longitudinal coverage from infancy through 36 months, repeated interaction sessions, clinically validated developmental outcomes, expert frame-level behavioral annotations, and privacy-conscious multimodal feature representations. To the best of our knowledge, existing ASD behavioral datasets do not jointly provide these characteristics at comparable scale. To demonstrate the utility of ASD-FEAT, we use it to evaluate a computer-vision-based end-to-end pipeline relying on machine learning techniques to automatically identify ASD risk. ASD-FEAT integrates both expert-defined and deep-learned features, including face and eye landmarks, facial action units, gaze direction, head position, mel-spectrogram audio representations, and optical flow, to identify behavioral markers of social interaction. Our automated pipeline achieves an ASD classification accuracy of 76.2% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.82, compared to classifiers trained on manually labeled behaviors, which yielded 81.3% accuracy and an AUROC of 0.88. We further introduce a within-visit partner contrast: a per-visit signal contrasting examiner-directed and parent-directed social behavior which, when added to the classifier, lifts the fully automated Look Face + Smile and Look Face + Vocal configurations to Matthews correlations of 0.49 and 0.45 respectively, exceeding the human-coded single-partner baseline of 0.42.
Problem

Research questions and friction points this paper is trying to address.

Autism Spectrum Disorder
Early Screening
Multi-Modal Dataset
Risk Prediction
Infant Behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Dataset
Longitudinal Tracking
End-to-End Pipeline
Feature Fusion
Within-Visit Partner Contrast
💼 Related Jobs
No related jobs found.
S
Sidrah Liaqat
Department of Electrical and Computer Engineering, University of Kentucky, Lexington, KY 40506, USA
H
Halil Helvaci
Department of Electrical and Computer Engineering, University of Kentucky, Lexington, KY 40506, USA
S
Sen-Ching Cheung
Department of Electrical and Computer Engineering, University of Kentucky, Lexington, KY 40506, USA
Chongruo Wu
Chongruo Wu
UC Davis
Computer Vision
D
Dongjie Chen
Department of Electrical and Computer Engineering, University of California, Davis, CA 95616, USA
C
Chen Nee Chuah
Department of Electrical and Computer Engineering, University of California, Davis, CA 95616, USA
S
Sally Ozonoff
Department of Psychiatry and Behavioral Sciences, MIND Institute, University of California, Davis, CA 95616, USA