Institution profile

Seattle Children's Hospital

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

TimeNet: An Extensible Unified Data Infrastructure for Next-Generation Temporal Foundation Models

Oct 03, 2026

This study addresses the challenges of data format fragmentation and task-specific pipelines that hinder cross-domain generalization in time-series foundation model development. To this end, it proposes a scalable, unified, open-source data standard that decouples multimodal signals from task definitions. By leveraging a shared data model, heterogeneous task families are expressed as reusable views of a single record, enabling the construction of a configuration-driven training pipeline for joint training. The project efficiently transcodes 1.5 million instances, and cross-dataset joint training yields a 14% improvement in F1 score. Ultimately, this work provides a universal data infrastructure to facilitate the large-scale research and development of time-series foundation models.

0 citationsRead paper

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

Oct 01, 2026

This study addresses the vulnerability of open-source large language models to counterfactual biases driven by non-clinical factors in pediatric emergency triage, which threatens decision fairness. To investigate this, we construct paired counterfactual cases altering only demographic or social variables to audit the sensitivity of Emergency Severity Index predictions across ten open-source models. We propose a lightweight, interpretable framework incorporating hierarchical correlation analysis, evaluated alongside QLoRA fine-tuning and medical-specific models such as MedGemma. Our findings reveal that scaling model size or employing medical pretraining does not necessarily mitigate bias, uncovering latent directional failure modes. Notably, the fine-tuned Qwen2.5-7B achieves the lowest bias rate at 5.27%. This work provides empirical evidence and methodological support for conducting fairness audits prior to clinical deployment.

0 citationsRead paper

Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark

Aug 03, 2026

This study addresses the challenge of cross-temporal-window distribution shift in functional near-infrared spectroscopy (fNIRS)-based autism classification, caused by inter-individual variability in hemodynamic response delays. The work formalizes this temporal transfer problem for the first time and establishes a benchmark encompassing varying window lengths and shifts. Leveraging fNIRS topographic maps processed through vision-based models, the authors systematically evaluate eight strategies—including zero-shot inference, adversarial domain adaptation, and self-supervised learning—under settings with no target data or minimal fine-tuning. Results reveal that zero-shot performance remains limited (54–69%), underscoring individual variability as a critical bottleneck. However, fine-tuning on merely ~5% of target subjects restores accuracy to 90–96%, while unsupervised domain adaptation achieves 78–90%, demonstrating that even short 2.5-second windows retain sufficient discriminative information.

0 citationsRead paper
Recent publications

Latest Papers

TimeNet: An Extensible Unified Data Infrastructure for Next-Generation Temporal Foundation Models

Oct 03, 2026

This study addresses the challenges of data format fragmentation and task-specific pipelines that hinder cross-domain generalization in time-series foundation model development. To this end, it proposes a scalable, unified, open-source data standard that decouples multimodal signals from task definitions. By leveraging a shared data model, heterogeneous task families are expressed as reusable views of a single record, enabling the construction of a configuration-driven training pipeline for joint training. The project efficiently transcodes 1.5 million instances, and cross-dataset joint training yields a 14% improvement in F1 score. Ultimately, this work provides a universal data infrastructure to facilitate the large-scale research and development of time-series foundation models.

0 citationsRead paper

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

Oct 01, 2026

This study addresses the vulnerability of open-source large language models to counterfactual biases driven by non-clinical factors in pediatric emergency triage, which threatens decision fairness. To investigate this, we construct paired counterfactual cases altering only demographic or social variables to audit the sensitivity of Emergency Severity Index predictions across ten open-source models. We propose a lightweight, interpretable framework incorporating hierarchical correlation analysis, evaluated alongside QLoRA fine-tuning and medical-specific models such as MedGemma. Our findings reveal that scaling model size or employing medical pretraining does not necessarily mitigate bias, uncovering latent directional failure modes. Notably, the fine-tuned Qwen2.5-7B achieves the lowest bias rate at 5.27%. This work provides empirical evidence and methodological support for conducting fairness audits prior to clinical deployment.

0 citationsRead paper

Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark

Aug 03, 2026

This study addresses the challenge of cross-temporal-window distribution shift in functional near-infrared spectroscopy (fNIRS)-based autism classification, caused by inter-individual variability in hemodynamic response delays. The work formalizes this temporal transfer problem for the first time and establishes a benchmark encompassing varying window lengths and shifts. Leveraging fNIRS topographic maps processed through vision-based models, the authors systematically evaluate eight strategies—including zero-shot inference, adversarial domain adaptation, and self-supervised learning—under settings with no target data or minimal fine-tuning. Results reveal that zero-shot performance remains limited (54–69%), underscoring individual variability as a critical bottleneck. However, fine-tuning on merely ~5% of target subjects restores accuracy to 90–96%, while unsupervised domain adaptation achieves 78–90%, demonstrating that even short 2.5-second windows retain sufficient discriminative information.

0 citationsRead paper