Score
Designs and implements encoders and compression pipelines that map inputs into latent representations expressed in a spectral or frequency domain (e.g., Fourier or log‑spectral coefficient matrices) and produces position‑ or component‑wise latent vectors. Builds and analyzes training procedures and regularizers for probabilistic latent models (such as VAEs) that enforce spectral constraints or zone‑weighted spectral compression (spectrally regularized VAE training, spectrally regularized latent compression, Fourier latent encoding).
研究通过分析自编码器参数矩阵的谱特性,将其视为数据样本的向量表示,从而区分不同训练子集的模型,无需使用原始样本。
This work addresses the limited diffusibility of latent spaces in latent diffusion models, which stems from a mismatch between the spectral distribution of latent variables and that of natural images, leading to degraded generation quality. The authors propose the Spectral Matching Hypothesis, introducing Encoder Spectral Matching (ESM) to enforce a flat power-law power spectral density in latent representations and Decoder Spectral Matching (DSM) to preserve frequency-wise semantic consistency. They establish, for the first time, a unified theoretical framework for latent diffusibility from a spectral perspective, explaining phenomena such as over-noising or over-smoothing in existing methods and unifying several approaches under this framework. Furthermore, they extend the spectral principle to representation alignment tasks (REPA). Experiments on CelebA and ImageNet demonstrate that the proposed method significantly outperforms current state-of-the-art techniques, validating the efficacy of spectral matching in enhancing diffusion model performance.
This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.
This study investigates geometric differences among latent manifolds learned by convolutional (CAE), denoising (DAE), and variational autoencoders (VAE) on complex data. We propose a theoretical framework grounded in matrix manifold modeling and Hilbert-space isometric embeddings, integrating distance-preserving embeddings with the geometry of symmetric positive definite (SPD) and symmetric positive semidefinite (SPSD) matrices to systematically analyze manifold continuity and hierarchical structure under input perturbations. We establish, for the first time, that the VAE latent manifold is a smooth product manifold composed of two SPD and one SPSD matrix manifolds; in contrast, CAE and DAE latent manifolds exhibit hierarchical, non-smooth structures—each layer being itself a smooth product manifold. This work provides a unified differential-geometric explanation for fundamental smoothness disparities across autoencoder architectures, offering a novel theoretical foundation for interpretable latent spaces and controllable generative modeling.
To address the high computational overhead of video VAE encoding and latent-space discontinuities induced by block-wise inference in high-resolution, long-duration videos, this paper proposes WaveFlow-VAE—a wavelet-driven energy-flow video VAE. The method integrates multilevel discrete wavelet transform (DWT), variational autoencoding, and latent-space energy-flow modeling. Its core contributions are: (1) a novel multilevel wavelet decomposition scheme that guides low-frequency energy toward compact latent representations, enabling energy-efficient modeling; and (2) a causal caching mechanism ensuring temporal consistency and completeness of latent features during block-wise inference. Experiments demonstrate that WaveFlow-VAE outperforms state-of-the-art video VAEs in reconstruction quality (higher PSNR and lower LPIPS), achieves 2× higher throughput, and reduces GPU memory consumption by 4×, while maintaining superior visual fidelity.
To address the challenges of massive data volume and storage/transmission bottlenecks associated with geostationary hyperspectral satellites (e.g., NASA’s TEMPO), this paper proposes a variational autoencoder (VAE)-based joint compression and atmospheric retrieval framework. The method unifies hyperspectral data compression with latent-space extraction of Level-2 atmospheric products—including NO₂, O₃, HCHO, and cloud fraction—for the first time. We discover that atmospheric information exhibits semi-linear latent encoding, where nonlinear latent probes significantly outperform linear models, while explicit supervision yields marginal gains in latent-space quality—revealing fundamental information-preservation constraints in neural compression. Our approach achieves an ultra-high compression ratio of 514×, with reconstruction errors reduced by one to two orders of magnitude. Retrieval accuracy reaches R² = 0.93 for cloud fraction and R² = 0.81 for total ozone, demonstrating efficient preservation of critical atmospheric signals.
This work addresses the challenge of integrating variational autoencoders (VAEs) as trainable layers within neural networks. It proposes a general framework for flexibly embedding VAEs into arbitrary network architectures, accompanied by an end-to-end training strategy that leverages the reparameterization trick and probabilistic modeling to ensure full differentiability throughout the pipeline. For the first time, this approach enables VAEs to function as plug-and-play modules akin to standard neural network layers, substantially enhancing their compatibility and representational capacity within complex models. Experimental results demonstrate that the proposed VAE layer consistently achieves stable performance across diverse tasks and outperforms conventional standalone VAE models, thereby significantly expanding the applicability of VAEs in deep learning systems.
This work addresses the trade-off in audio variational autoencoders (VAEs) between over-regularization, which degrades generation quality, and under-regularization, which harms downstream task performance. The authors propose a target KL regularization mechanism that, for the first time, enables controllable training of audio VAEs at specified bitrates. Leveraging rate–distortion theory, they construct rate–distortion curves over the continuous latent space, establishing a fair basis for comparison with discrete neural audio codecs. The approach is evaluated in text-to-sound generation, where the impact of varying compression rates on synthesis quality is systematically analyzed, revealing an optimal configuration that achieves a significant balance between audio fidelity and the predictability of latent representations.
This work addresses the challenge of learning disentangled and interpretable latent representations from complex, non-stationary, high-dimensional time-varying signals, which exhibit rich time-frequency structures that conventional variational autoencoders (VAEs) struggle to model effectively. To this end, we propose the Decompositional Variational Autoencoder (DecVAE), which uniquely integrates signal decomposition priors directly into the VAE framework. DecVAE employs an encoder-only architecture and jointly leverages signal decomposition models, contrastive self-supervised tasks, and variational inference to learn multi-subspace latent representations aligned with the intrinsic time-frequency characteristics of the data. Extensive experiments on synthetic data and three scientific datasets demonstrate that DecVAE substantially outperforms existing VAE approaches, achieving significant improvements in disentanglement quality, cross-task generalization, and interpretability of the learned latent representations.
Existing video VAEs overly prioritize reconstruction fidelity while neglecting the impact of latent-space spectral structure on diffusion training, resulting in slow convergence and suboptimal generation quality. This work identifies two intrinsic properties of video latent representations: pronounced low-frequency spatiotemporal distribution and channel-wise modality concentration. Leveraging these insights, we propose two lightweight, general-purpose structured regularization techniques—local correlation regularization and latent masking reconstruction. These methods jointly guide the VAE to learn a spectrally controllable and structurally coherent latent space, significantly improving both training efficiency and generative performance of subsequent diffusion models. Experiments demonstrate a 3× speedup in diffusion training, a 10% improvement in video reward scores, and consistent superiority over leading open-source video VAEs across multiple quantitative metrics.