Score
Designs and implements encoding schemes and latent-variable models that map inputs into binary-valued latent codes drawn from a learned stochastic distribution (stochastic binary latent coding, binary latent coding, stochastic binary representations) or their continuous relaxations (softbinary coding). Builds training algorithms and network architectures to learn these discrete/relaxed representations end-to-end, and analyzes their quantization train–test mismatch, rate–distortion behavior, and effects on representation quality and vector-quantization performance.
This work addresses key challenges in nonlinear transform coding (NTC)—namely, the training–test mismatch, smoothness bias in continuous transforms, and the difficulty of achieving shaping gain with high-dimensional vector quantization—by proposing SoftBinary Coding (SBC), an end-to-end neural compression framework based on a stochastic binary latent space. SBC introduces, for the first time, an information-theoretically optimal binary discrete structure into neural compression, combining a differentiable training framework with efficient binary channel simulation to theoretically guarantee rate optimality. Experimental results demonstrate that SBC outperforms classical trellis-coded quantization (TCQ) on i.i.d. source vector quantization tasks, achieving state-of-the-art performance.
This work addresses the lack of compositional structure in conditional representations of diffusion models, which hinders generalization to out-of-distribution (OOD) samples. We propose Discrete Latent Codes (DLC), a self-supervised discrete image representation learned via simplex embedding. DLC encodes images into semantically compositional discrete token sequences, achieving high-fidelity reconstruction while significantly improving generation efficiency. Its core innovation is the first explicitly compositional discrete image representation, enabling unconditional high-quality synthesis, OOD novel image generation, and efficient text-to-image generation when integrated with large language models. On ImageNet, DLC achieves state-of-the-art performance in unconditional image generation. Empirical results validate its rationality for cross-distribution generation and effectiveness for text-guided synthesis.
This work addresses the challenge of balancing accuracy and generalization in neural network quantization by proposing a soft quantization method during training. The approach introduces short-range attractive coupling among weights to encourage automatic discretization of the weight distribution, enabling mixed-precision compression without complex scheduling schemes. Relying on only two hyperparameters, the method offers both simplicity and flexibility, providing a novel tool for studying the trade-off between model compression and generalization. Experimental results on ResNet-20 with CIFAR-10 demonstrate that the proposed technique outperforms post-training quantization methods based on histogram equalization, achieving higher accuracy in compressed models.
This work addresses the limitations of existing image generation models in generation quality, visual representation capability, and multi-task generalization. We propose BiGR—the first conditional image generation model that unifies generative and discriminative capabilities within a single architecture. Its key contributions are: (1) a novel compact binary latent code modeling scheme, incorporating a binary tokenizer and masked autoregressive prediction; (2) joint optimization of diverse generative tasks—including inpainting, editing, outpainting, and interpolation—alongside discriminative evaluation via linear probing; and (3) entropy-ordered sampling for efficient, high-fidelity generation. Experiments demonstrate significant improvements: FID-50k substantially decreases, and linear probe accuracy increases markedly. BiGR enables zero-shot cross-task generalization and successfully extends to text-to-image generation, achieving state-of-the-art generation performance while preserving strong representation learning capacity.
Deep generative models (DGMs) suffer from statistical non-identifiability and semantic entanglement of latent variables, severely limiting interpretability—especially when modeling heterogeneous data (e.g., text, images, response times). Method: We propose a class of identifiable and interpretable multilayer binary latent variable models, unifying heterogeneous data representation. We establish the first rigorous identifiability theory for discrete deep generative models, proving that latent layer sizes must decrease strictly across layers. We introduce inter-layer nonlinear spectral initialization and a sparse-penalized stochastic approximation EM algorithm, integrating directed graphical modeling, nonlinear spectral analysis, and hierarchical binary encoding. Contribution/Results: Our approach yields semantically transparent, interpretable outcomes in topic modeling, image representation learning, and educational response-time analysis. Simulation studies confirm parameter estimation consistency and demonstrate efficient estimation of exponentially many latent components.
This work addresses the inconsistent performance of binary quantization in embedding spaces, which excels in contrastive learning embeddings but degrades sharply in others, and resolves the lack of a unified theoretical foundation between the “random rotation” and “axis-aligned” quantization strategies. The study identifies the heterogeneity of coordinate-wise variances as the key factor governing quantization efficacy and establishes, for the first time, an analytical framework under a Gaussian structural assumption. This framework yields a closed-form solution for rank fidelity, quantitatively linking the information content of magnitude bits to variance heterogeneity, and unifies the conditions under which the two seemingly opposing strategies are optimal. Theoretical predictions are validated across 13 datasets and 6 embedding types, providing the first principled design guidelines for binary quantization systems.
This work addresses the high storage and transmission costs of neural network parameters by proposing an extreme compression method based on random seeds and quantized latent variables. The model weights are represented as mappings generated jointly by a fixed random basis, trainable quantized latent vectors, and seed-based initialization. This approach eliminates the need to store full weight matrices or projection matrices, requiring only a compact set of latent vectors for model reconstruction, while preserving accuracy through quantization-aware fine-tuning. Leveraging a block-wise scalable basis design, the method efficiently supports deployment of large-scale models. Experiments demonstrate that, under comparable accuracy to low-bitwidth quantized models, the compressed model size is drastically reduced—determined solely by the dimensionality and bitwidth of the latent variables rather than the original parameter count.
This work proposes Permutation-Invariant Vector Quantization (PI-VQ), a novel discrete representation framework that overcomes the entanglement between codebook entries and spatial positions inherent in conventional approaches like VQ-VAE. By enforcing permutation invariance in the latent codes, PI-VQ learns global semantic features independent of location, enabling direct interpolation-based image generation without requiring a trained prior model. The method introduces a matching quantization algorithm based on optimal bipartite graph matching, which increases bottleneck capacity by 3.5× and allows synthesis of novel images in a single forward pass. Experiments on CelebA, CelebA-HQ, and FFHQ demonstrate that PI-VQ achieves superior performance in terms of precision, density, and coverage, validating the efficacy of position-agnostic discrete representations for generative modeling.
This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.
This work addresses the suboptimal reconstruction in lossy compression caused by the mismatch between the encoder’s assumed source distribution and the true data distribution. To mitigate this issue without modifying the encoder, the authors propose a generative decompression framework that leverages prior knowledge of the true source distribution at the decoder. By employing Bayesian estimation and conditional expectation, the method achieves optimal reconstruction under the fixed encoder constraint. This study is the first to introduce Bayes-optimal decoding into mismatched compression scenarios and extends the approach to noisy channels and task-oriented compression. The framework integrates Gaussian source modeling, maximum a posteriori detection, and deep semantic classification, significantly narrowing the performance gap with jointly optimized benchmarks while enabling high-fidelity, adaptive reconstruction.