🤖 AI Summary
To enable high-fidelity digital twins of radio access networks (RAN), lightweight packet-level traffic generators are needed that accurately reproduce the heavy-tailed payload sizes and inter-packet arrival time distributions of real traffic, while ensuring model compactness and re-calibratability.
Method: We propose a novel hybrid architecture combining hidden Markov models (HMMs) with Student’s t mixture density networks: the HMM explicitly models idle states to anchor heavy-tailed behavior, while each latent state adaptively modulates the tail thickness of its associated Student’s t distribution; flow-level cumulative distribution functions are modeled, and fidelity is evaluated via Wasserstein distance.
Results: Evaluated on Web, smart-home, and encrypted media datasets, our method significantly outperforms neural network and Transformer baselines—reducing parameter count by 1–2 orders of magnitude, achieving a model size of only 0.2 MB—and yields the closest match to ground-truth traffic in terms of CDF fidelity, autocorrelation structure, and flow descriptor distributions.
📝 Abstract
Digital twins of radio access networks require packet-level traffic generators that reproduce the size and timing of packets while remaining compact and easy to recalibrate as traffic changes. We address this need with a hybrid generator that combines a small hidden Markov model, which captures buffering, streaming, and idle states, with a mixture density network that models the joint distribution of payload length and inter-arrival time (IAT) in each state using Student-t mixtures. The state space and emission family are designed to handle heavy-tailed IAT by anchoring an explicit idle state in the tail and allowing each component to adapt its tail thickness. We evaluate the model on public traces of web, smart home, and encrypted media traffic and compare it with recent neural network and transformer based generators as well as hidden Markov baselines. Across most datasets and metrics, including average per-flow cumulative distribution functions, autocorrelation based measures of temporal structure, and Wasserstein distances between flow descriptors, the proposed generator matches the real traffic most closely in the majority of cases while using orders of magnitude fewer parameters. The full model occupies around 0.2 MB in our experiments, which makes it suitable for deployment inside digital twins where memory footprint and low-overhead adaptation are critical.