🤖 AI Summary
This study addresses the challenge of temporal context fragmentation and long-horizon prediction in federated world models, where trajectory ownership boundaries inherently limit information continuity. Using the MIMIC-IV dataset, we construct clients partitioned by illness severity and conduct hourly clinical forecasting experiments employing ten federated algorithms, including FedAvg and FedProx, integrated with local history rules and paired rolling evaluation techniques. Our contributions reveal how participation rates and data ownership constrain long-window coverage, demonstrating that 32-step window coverage reaches merely 7.55%–21.36%. Furthermore, we quantify the significant error escalation in FedAvg induced by fine-grained client partitioning. Ultimately, this work establishes a multidimensional evaluation framework for federated learning in clinical time-series forecasting, providing critical insights into the interplay between federation dynamics and predictive horizon limitations.
📝 Abstract
World models learn state evolution from trajectories, making access to temporal context a central training requirement. Federated learning can use distributed records, while ownership boundaries within a trajectory restrict the examples each client can construct. Our study benchmarks this cross-time setting through hourly action-conditioned clinical prediction on eight MIMIC-IV disease cohorts, comprising 40.87 million transition memberships. We specify severity-based client ownership, patient-separated construction, local history and future-window rules, and paired rollout evaluation from one to 32 hours. A matrix of ten federated algorithms covers 32 disease--partition configurations under five rounds of ten-percent participation. Three findings emerge from existing results and training logs. First, client ownership and participation jointly restrict long-window coverage: only 7.55\%--21.36\% of pooled-available 32-step windows have a locally complete anchor visited during training, averaged across diseases. Second, finer severity partitions accompany higher FedAvg error in 15 of 16 paired comparisons, while algorithm gains are small and horizon-dependent: FedProx reduces mean error by 0.56\%, with no consistent improvement at 32 steps. Third, algorithm labels conceal distinct update behavior, including inactive extrapolation and orders-of-magnitude differences in update scale. Cached-update performance also varies strongly across trajectory partitions under the same benchmark protocol. These results establish temporal access, participation coverage, optimization behavior, and horizon-resolved prediction as complementary dimensions for evaluating federated clinical world models.