Training and Scaling Compute-Optimal Physiological Waveform Foundation Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of compute-optimal scaling laws for physiological waveform foundation models by training over one hundred Aether models. Through linear probing evaluations on MIMIC-III tasks and multivariate scaling law fitting, it systematically investigates the quantitative relationships among model size, pretraining duration, and clinical predictive performance. The findings reveal that, under compute-optimal conditions, data scaling outperforms parameter scaling, and they confirm the complementary synergistic effects of pretraining and labeled supervision. Furthermore, an optimized FLOPs allocation strategy is proposed. Ultimately, a 720M-parameter model surpasses existing baselines with scaling law prediction errors below 1%, establishing a reliable training paradigm for efficient clinical prediction.
📝 Abstract
We investigate the scaling laws and compute-optimal training of physiological waveform foundation models (FMs). We train Aether, a family of over one hundred FMs ranging from 20M to 2.1B parameters, on up to 36.3M hours of physiological waveforms. We construct eight clinical prediction tasks from MIMIC-III and evaluate the FMs through linear probing. The 720M FM outperforms all existing baseline FMs across all eight tasks. A scaling law of model size, pretraining hours, and labeled patients predicts downstream ranking error, i.e. $1-\mathrm{AUROC}$, effectively with $0.5\%$ prediction MAE at held-out resource scales and $0.9\%$ MAE when extrapolating to 2.1B parameters. We present three findings: (1) Compute-optimal training scales both FM size and pretraining hours. Under the fitted law, a $10.0\times$ increase in compute FLOPs scales model size by $1.2\times$ and pretraining hours by $8.2\times$. (2) Larger FMs use waveform data more efficiently, and greater pretraining exposure increases the benefit of model scaling. Starting from 25M parameters and 4.8M pretraining hours, doubling FM size reduces the predicted hours needed for the same performance by $51.8\%$. (3) Pretraining and clinical supervision reinforce each other: more labeled patients increase the return to pretraining, while larger FMs and longer pretraining reduce labeling requirements. For the example of the 720M FM, extending pretraining from 120K to 36.3M hours reduces the predicted patient requirement by $61\%$ at a target ranking error. These findings provide a quantitative training recipe and a promising and durable scaling path for physiological waveform modeling and downstream clinical prediction.
Problem

Research questions and friction points this paper is trying to address.

physiological waveform foundation models
scaling laws
compute-optimal training
clinical prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physiological Waveform Foundation Models
Scaling Laws
Compute-Optimal Training
Clinical Prediction
Linear Probing