Score
Design and train encoder–decoder systems (e.g., VAEs, quantized/continuous encoders, and related architectures) that compress high‑dimensional inputs into compact continuous or discrete latent vectors and associated decoders. Guide the latent-space learning with supervised, metric, or architectural constraints so the resulting geometry and codes reflect target attributes (for example separability of outcomes or embedding of performance), and implement the loss functions, quantizers, and evaluation analyses needed to enforce and measure those properties.
Diffusion models on high-dimensional data suffer from poor generation quality, low training efficiency, and—critically—failure to preserve the intrinsic geometric structure of the data distribution. Method: This paper proposes a geometry-preserving encoder-decoder framework to replace conventional VAEs, enabling efficient and stable diffusion modeling in the latent space. Contribution/Results: We introduce the first theoretical encoder-decoder framework with rigorous differential-geometric constraints (e.g., isometry or conformality); provide a convergence proof for the encoder and demonstrate its acceleration effect on decoder convergence; and design a theory-guided encoder optimization strategy. Experiments show that our method significantly improves joint training stability and convergence speed, achieves superior generation quality across multiple benchmarks, and reduces training time.
This paper addresses the fundamental mismatch between the continuous latent space of standard Variational Autoencoders (VAEs) and the inherently discrete nature of data such as text. To resolve this, we propose the Discrete VAE—a VAE explicitly designed for categorical latent variables. Methodologically, we derive the evidence lower bound (ELBO) rigorously from first principles of variational inference under categorical latents and employ the Gumbel-Softmax reparameterization to enable differentiable gradient estimation in discrete latent spaces. Our key contributions are threefold: (1) a tutorial-style, unified theoretical framework for discrete VAEs; (2) a robust and reproducible training paradigm; and (3) publicly released, fully functional code. Experiments demonstrate that the Discrete VAE significantly improves interpretability and structural coherence in discrete data generation, outperforming continuous-latent baselines while preserving principled probabilistic modeling.
This paper addresses the conceptual gap between variational autoencoders (VAEs) and vector-quantized VAEs (VQ-VAEs) in modeling continuous versus discrete latent representations. Methodologically, it proposes a novel autoencoder framework that eliminates both the KL divergence term and the reparameterization trick; instead, it explicitly enforces latent space compactness via learnable clustering centers and employs multi-vector outputs to jointly support continuous interpolation and discrete reconstruction. Key contributions include: (1) uncovering an intrinsic relationship between autoencoder generative fidelity and latent space compactness; (2) establishing a deterministic transition from VAEs to VQ-VAEs without relying on probabilistic assumptions; and (3) empirically validating smooth interpolation and stable reconstruction on MNIST, CelebA, and FashionMNIST. Experiments further reveal that naively increasing the number of output vectors leads to model degradation—manifesting as localized, patchwise discrete encoding—highlighting the critical role of architectural design.
To address the uncontrollable topology of latent spaces (LS) in autoencoders (AEs), this paper proposes a geometric-loss-guided supervised AE co-optimization framework—the first to explicitly configure LS topology in supervised AEs. The method jointly optimizes encoder architecture and geometric constraint losses, enabling user-defined cluster positions and shapes, decoder-free label prediction, and cross-sample similarity assessment. Key innovations include zero-shot cross-dataset generalization, similarity estimation for unseen classes, and text-driven image retrieval without classifiers or language models. Experiments demonstrate 12–19% improvements in zero-shot accuracy on LIP, Market-1501, and WildTrack, and achieve 78.3% mAP in cross-modal retrieval.
This study systematically investigates the impact of latent-space dimensionality on IoT botnet detection performance, comparing Vision Transformers (ViTs) and Variational Autoencoder (VAE) encoders for dimensionality reduction of structured network traffic data (CSV format). Addressing the absence of spatial locality and hierarchical structure in IoT traffic—key inductive biases assumed by ViTs—we reveal, for the first time, that ViTs underperform significantly relative to VAEs due to model-structure mismatch. Under a unified framework, we couple ViT and VAE encoders with downstream classifiers (MLP, LSTM) and evaluate them on the N-BaIoT and CICIoT2022 datasets. Results demonstrate that VAE consistently outperforms ViT across all tested latent dimensions, achieving average improvements of 5.2%–13.8% in four core metrics: accuracy, precision, recall, and F1-score. This confirms VAE’s superior suitability for unsupervised representation learning on non-image sequential traffic data.
This study investigates the joint impact of latent dimensionality and frame rate in continuous audio encoders on downstream task performance. By employing matched training protocols, frozen-model PCA interventions, and automatic speech recognition (ASR) probing techniques, it systematically evaluates representational disparities across varying width and frame rate configurations. The research reveals the interaction mechanisms between these factors regarding downstream utility, demonstrating that representational organization is more critical than mere reconstruction fidelity. Experimental findings indicate that a moderate latent width paired with a high frame rate optimally benefits ASR performance, while high-dimensional models exhibit performance bottlenecks under specific compression ratios. These results challenge the conventional assumption that higher dimensionality is inherently superior, thereby establishing a new paradigm for encoder design.
This study investigates whether information compression—characterized by low mutual information—and geometric compression, manifested as intra-class clustering, are inherently linked in deep learning. Employing the Conditional Entropy Bottleneck (CEB) framework, continuous dropout, theoretically sound mutual information estimators, controlled noise injection, and detailed intra-class clustering analysis, the work systematically examines the interplay between these two forms of compression and generalization. The findings reveal no stable correspondence between information and geometric compression; instead, they exhibit a nonlinear negative relationship modulated by training configurations. Moreover, generalization appears to act more as a confounding factor than a direct consequence of either compression mechanism. These results challenge prevailing assumptions and offer a refined perspective on the mechanisms underlying generalization in representation learning.
Standard variational autoencoders employ Gaussian priors, which struggle to align with data manifolds exhibiting non-Euclidean topologies—such as periodicity or boundedness—leading to distorted representations. This work proposes a topology-aware latent space modeling framework that constructs factorized prior distributions tailored to manifolds decomposable into products of circles, intervals, and lines, along with their finite group quotients. This design enables disentangled latent representations and analytically tractable KL divergences. By integrating differentiable coordinate transformations, group-invariant decoding, and anchor-point constraints, the approach ensures smooth gradients and topological consistency. To our knowledge, this is the first method to systematically align latent variable distributions with the intrinsic topology of data manifolds, supporting reparameterizable encoder–prior pairs and significantly outperforming Gaussian-prior baselines on synthetic manifolds as well as rotation- and cyclic-translation variants of MNIST.
This work addresses the challenge in variational autoencoders (VAEs) of simultaneously achieving high representational capacity and disentangled, low-dimensional latent representations. The authors formulate VAE training as a soft-constrained optimization problem, introducing an entropy-based soft constraint mechanism to regulate the information content of individual latent variables. Coupled with weight filtering, this approach enables automatic pruning of low-entropy dimensions. The proposed method enhances representation efficiency while preserving disentanglement. Experiments demonstrate significant improvements: on dSprites, activation scores increase by 43–62%, FactorVAE score reaches 0.891, and reconstruction error decreases by 38%; on MNIST, over 90% classification accuracy is achieved using only two latent dimensions—reducing input dimensionality by 80% compared to baselines—and training convergence accelerates by 37%.
研究通过对比不同编码器和解码器线性程度的自编码器架构,发现使用线性编码器与非线性解码器在保持降维精度的同时提高了模型简洁性和解释性。