🤖 AI Summary
This work addresses the challenge of preserving high-order statistical structures in highly compressed high-dimensional flow fields, where conventional autoencoders such as VAEs exhibit limited performance. The authors propose DiffCoder, a novel framework that, for the first time, integrates a diffusion model as a generative prior into the flow field compression and reconstruction task. By jointly training a convolutional ResNet encoder with a conditional diffusion decoder in an end-to-end manner, DiffCoder effectively maintains spectral characteristics and statistical consistency of the flow fields even under severe information bottlenecks. Experiments on the Kolmogorov flow dataset demonstrate that DiffCoder significantly outperforms VAEs under strong compression, achieving both higher pointwise accuracy and superior statistical fidelity.
📝 Abstract
Data-driven flow-field reconstruction typically relies on autoencoder architectures that compress high-dimensional states into low-dimensional latent representations. However, classical approaches such as variational autoencoders (VAEs) often struggle to preserve the higher-order statistical structure of fluid flows when subjected to strong compression. We propose DiffCoder, a coupled framework that integrates a probabilistic diffusion model with a conventional convolutional ResNet encoder and trains both components end-to-end. The encoder compresses the flow field into a latent representation, while the diffusion model learns a generative prior over reconstructions conditioned on the compressed state. This design allows DiffCoder to recover distributional and spectral properties that are not strictly required for minimizing pointwise reconstruction loss but are critical for faithfully representing statistical properties of the flow field. We evaluate DiffCoder and VAE baselines across multiple model sizes and compression ratios on a challenging dataset of Kolmogorov flow fields. Under aggressive compression, DiffCoder significantly improves the spectral accuracy while VAEs exhibit substantial degradation. Although both methods show comparable relative L2 reconstruction error, DiffCoder better preserves the underlying distributional structure of the flow. At moderate compression levels, sufficiently large VAEs remain competitive, suggesting that diffusion-based priors provide the greatest benefit when information bottlenecks are severe. These results demonstrate that the generative decoding by diffusion offers a promising path toward compact, statistically consistent representations of complex flow fields.