Score
Design and implement pretraining methods and foundation models for network telemetry: build data ingestion and self-supervised training pipelines that operate on multivariate, bursty and zero‑inflated telemetry signals, learn cross‑layer metric coupling, and produce representations that improve long‑horizon forecasting and downstream network analytics.
This work addresses the limitations of traditional operations that rely on multiple isolated predictors for heterogeneous telemetry data, which hinder unified modeling and degrade decision-making performance. The paper proposes the first generative telemetry foundation model tailored for HPC job scheduling and network resource allocation. Treating telemetry as an event-driven stream of irregular temporal entities, the model employs a single-pass inference architecture that requires no future timestamps and incorporates calibrated conditional quantile forecasting. This design enables flexible time-horizon predictions and zero-shot transfer across time intervals, achieving effective cross-domain adaptation with only a few hours of target-domain data. Replay experiments on real-world HPC job logs and network traffic demonstrate that the approach reduces average bounded slowdown by 77% and cuts policy violation rates by approximately 50%.
Supervised learning in cybersecurity suffers from heavy reliance on large-scale labeled data and poor generalization. Method: We propose the first general-purpose foundation model tailored for cybersecurity deployment, introducing a novel protocol-aware tokenization mechanism, integrating multimodal network flow representations, and designing a hierarchical Transformer architecture to enable self-supervised pretraining on massive unlabeled network traffic data. Contributions/Results: Our model outperforms four state-of-the-art methods across five representative security tasks; demonstrates significantly improved robustness against noisy labels and spurious correlations (i.e., learning shortcuts); and effectively captures real-world network contextual semantics. This work establishes a new paradigm for building AI infrastructure in cybersecurity that minimizes labeling dependency while maximizing generalization capability.
This work addresses the poor transferability of general-purpose temporal foundation models to wireless network telemetry data, which exhibit burstiness, zero-inflation, and cross-layer coupling. To overcome these challenges, the authors propose APEX—a decoder-only Transformer architecture pretrained on large-scale real-world enterprise access point (AP) telemetry data—specifically designed for edge-side forecasting and anomaly detection. The paper introduces a network-native pretraining paradigm and presents two variants: APEX-Edge, a lightweight model optimized for edge deployment, and APEX-Large, a high-performance version for cloud inference. Experimental results demonstrate that APEX-Large reduces mean absolute error (MAE) by 18% over the strongest baseline Toto (and by 38% over SARIMA) on DHCP degradation prediction, achieving an F1 score of 0.93 in anomaly detection. Meanwhile, APEX-Edge enables sub-second, privacy-preserving inference directly on AP devices.
Traditional time-series forecasting typically minimizes global prediction error, overlooking downstream tasks’ heterogeneous requirements across different forecast horizons. This paper introduces a task-oriented forecasting paradigm that employs a task-driven, dynamically segmented weighting training mechanism, enabling models to adaptively prioritize learning for time segments according to their importance to downstream applications (e.g., resource scheduling). The method comprises three key components: (i) segmented forecast decomposition, (ii) importance-aware dynamic weight fusion, and (iii) an end-to-end loss function explicitly aligned with decision-making objectives. Evaluated on multiple standard benchmarks and a custom wireless communication dataset, the approach achieves significant improvements in both forecasting accuracy and downstream task performance. Notably, it is the first to realize joint optimization of prediction fidelity and decision-level objectives.
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
This work addresses the substantial computational waste in hyperparameter optimization caused by training runs that are destined to fail early on. The authors propose a method to accurately predict final model performance using only telemetry data—such as loss, accuracy, gradient signal-to-noise ratio, weight norm dynamics, and activation saturation—from the first few epochs of a single training run, along with its hyperparameters, without requiring information from other runs. For the first time, they systematically validate the predictive power of early-training signals and demonstrate the incremental value of gradient- and weight-level metrics. Leveraging a gradient-boosted tree model across 23,788 experiments, they achieve R² values of 0.92–0.99 in predicting final accuracy and ROC-AUC scores of 0.983–0.998 for relative performance ranking using just the first five epochs, with useful predictive signals emerging as early as after a single epoch.
This study addresses the limited training potential and architectural inflexibility of general-purpose time series foundation models by constructing a billion-parameter foundation model. Methodologically, it proposes a pattern-guided Mixture-of-Experts mechanism and an implicit quantile network head to enable sparsely activated routing and arbitrary quantile probabilistic forecasting. By integrating channel-independent pretraining, progressive curriculum learning, and variable-resolution post-training, the model supports flexible inference across variates, covariates, and multiple resolutions, while incorporating test-time scaling and parallel decoding to enhance efficiency. The proposed model achieves state-of-the-art performance on benchmarks such as GIFT-Eval, comprehensively outperforming existing pretrained and task-specific supervised models.
This study evaluates whether foundation models can replace traditional supervised methods for real-world time series forecasting without task-specific training. It introduces a novel operational perspective by categorizing forecasting tasks into four representative scenarios and proposes a sequence-feature-based complexity-aware routing mechanism to automatically select the optimal model. Through extensive cross-domain benchmarking and zero-shot inference comparisons against supervised baselines, the work demonstrates that foundation models excel in settings with transferable periodic structures or cold-start conditions. The proposed routing strategy not only maintains competitive prediction accuracy but also significantly reduces inference overhead, outperforming uniform deployment of foundation models across all tasks.
This study addresses the lack of systematic evaluation of pretraining methodologies and data-scaling effects for electrocardiogram (ECG) foundation models. Within a unified framework, it presents the first comprehensive comparison of five self-supervised learning objectives—including contrastive predictive coding and Joint Embedding Predictive Architecture (JEPA)—combined with three dominant architectures: structured state space models (SSMs), Transformers, and CNNs, evaluated on ECG datasets up to 11 million samples. The results demonstrate consistent performance gains with increasing data scale, with contrastive predictive coding slightly outperforming JEPA. Notably, structured state space models significantly surpass both Transformers and CNNs across multiple clinical downstream tasks, highlighting the critical role of their inductive bias in enabling superior transferability.
This work addresses the limitations of task-specific, data-hungry, and poorly generalizable temporal-aware modules in digital twin and Prognostics and Health Management (PHM) systems, which also suffer from integration challenges. To overcome these issues, the authors propose a modular foundation model based on an ensemble of pretrained encoders. The model leverages self-supervised learning to acquire transferable temporal representations, employs a gating mechanism for dynamic encoder selection, and utilizes Transformer-based self-attention to enable cross-encoder interaction and representation fusion. Innovatively, it adopts a shared latent space alignment combined with an adaptive aggregation strategy, allowing lightweight multi-task adaptation and conditional computation while keeping the pretrained encoders frozen. The approach demonstrates superior performance on the ETT benchmark and validates its practical utility in an industrial virtual sensing application for hydro-generator rotor temperature monitoring.