Score
Design and implement variational inference and optimization frameworks that combine contrastive objectives with a variational formulation to learn, discriminate, and manipulate latent representations and their temporal evolution; build the associated loss functions, inference procedures, and optimization algorithms that enforce variational evolution constraints and use latent contrastive signals to train unsupervised intervention- or policy-like mappings.
Generalized variational inference suffers from tight coupling between modeling and optimization, hindering modularity and composability. Method: We propose the first composable, unified framework for variational inference. Leveraging a newly discovered chain rule—akin to reverse-mode automatic differentiation—that governs the interplay between Bayesian inference and variational objectives, we design composable operators that decouple model structure, inverse modeling, local loss binding, and parameter exposure. Further, we introduce a statistical game-theoretic perspective to enable localized optimization. Contribution/Results: Experiments across canonical Bayesian models demonstrate that our framework significantly improves modularity, interpretability, and construction efficiency of variational inference. It supports flexible, plug-and-play assembly of arbitrary subcomponents and end-to-end optimization, establishing a novel paradigm for complex probabilistic modeling.
For constrained optimization problems (COPs) with unknown cost functions, this paper proposes an inverse optimization latent variable model that jointly learns the latent distribution of cost functions and structured, constraint-satisfying decodings. Methodologically, it is the first to integrate latent variable modeling with inverse optimization, employing the Fenchel–Young loss to enable gradient backpropagation through non-differentiable, deterministic solvers—thus supporting end-to-end optimization in latent space. Key contributions include: (1) modeling the *distribution*—rather than a point estimate—of cost functions, enabling behavior diversity modeling and generalizable inference across multi-agent or multi-scenario settings; and (2) guaranteeing strictly feasible and interpretable decoding paths. Evaluated on real-world vessel and taxi trajectory datasets, as well as synthetic graph-path data, the method achieves significant improvements in path reconstruction accuracy, predictive distribution quality, and interpretability of latent representations.
High-dimensional black-box constrained optimization faces challenges including intractable feasible region identification, the curse of dimensionality in Bayesian optimization (BO), and poor scalability and mode collapse in generative modeling. Method: We propose a novel framework integrating generative modeling with BO, where candidate solution sampling is formulated as posterior inference in a latent space. Leveraging normalizing flow-based generative models, our approach enables constraint-aware distribution learning and uncertainty quantification, while amortized inference ensures efficient sampling. Contribution/Results: The method effectively mitigates optimization difficulties arising from multimodality and hard constraints, avoids mode collapse, and significantly improves scalability to high dimensions. Extensive experiments on synthetic benchmarks and real-world tasks demonstrate superior convergence speed and robustness compared to state-of-the-art methods.
This paper addresses the challenges of variational parameter optimization, poor posterior approximation quality, and instability in variational inference (VI). We formulate VI as a Wasserstein gradient flow optimization problem over the space of variational parameters—marking the first such characterization. By endowing this parameter space with a probability measure and evolving it along the associated Wasserstein gradient flow, we provide a unifying interpretation of black-box VI and natural-gradient VI. We propose an efficient numerical solver based on particle discretization, preserving theoretical rigor while enhancing computational feasibility. We prove that classical VI methods arise as special cases within our framework. Experiments on synthetic data demonstrate substantial improvements in convergence speed, optimization stability, and posterior approximation accuracy compared to standard approaches.
This work addresses likelihood-free simulation-based inference (SBI), where the likelihood is intractable. To model complex, simulator-induced posterior distributions—often highly structured and multi-modal—we propose an efficient variational autoencoder (VAE)-based posterior estimation framework. Our method incorporates two complementary prior mechanisms: (1) a data-adaptive multivariate prior network to improve generalization across queries, and (2) a standard Gaussian prior to preserve model simplicity and expressiveness. End-to-end variational inference enables scalable, generative posterior approximation. On standard SBI benchmarks, our approach matches the accuracy of state-of-the-art normalizing flow–based methods while substantially reducing training and inference costs—yielding superior computational efficiency and scalability. The core contribution is the systematic integration of the VAE paradigm into SBI, establishing a lightweight, robust, and easily deployable solution for large-scale, high-dimensional simulation-based inference.
This work addresses three key challenges in constrained optimization with variational autoencoders: inefficient latent space sampling, difficulty in identifying critical variables, and training instability. To overcome these issues, the authors propose a Multi-stage Constrained Optimization Framework (MCOF) that integrates an entropy-constrained VAE to identify salient latent variables, applies a probability integral transform to uniformize the posterior distribution, and employs a constraint-priority filtering strategy to alternately optimize objectives and constraints. Diversity of solutions is preserved through resampling of non-critical variables. MCOF innovatively combines feature selection, posterior regularization, and multiplier-free constraint handling, enabling efficient optimization within a low-dimensional subspace while avoiding posterior collapse and Gaussian mixture bias. Experiments demonstrate that the method exactly recovers analytical optima on synthetic problems and generates fully novel, constraint-compliant molecular structures in the ZINC250k drug design task.
This work addresses policy mode collapse, fragile exploration, and distributional shift in reinforcement learning from human feedback by proposing a geometry-driven proximal policy optimization method. The approach models the policy as a particle-based variational inference process within a mixture-of-experts architecture, updated via Stein variational gradient descent. It introduces a geometric proximal control mechanism grounded in functional kernels and an expert orthogonality loss, thereby eliminating reliance on fixed clipping or KL divergence scheduling. Evaluated on 33B/4B sparse mixture-of-experts models, the method achieves substantial performance gains: a +179 ELO improvement on Codeforces programming tasks and a 32% reduction in token consumption on AIME mathematical reasoning benchmarks.
This work addresses the limitations of traditional variational inference, which struggles to calibrate posterior distributions due to the representational constraints of the evidence lower bound (ELBO). The authors propose a novel single-parameter variational objective that introduces, for the first time, an adjustable score-based posterior, thereby establishing a flexible variational framework capable of jointly learning hierarchical structures and Bayesian posteriors. By incorporating analytically tractable gradient computation, the method significantly improves posterior calibration in mixture models and achieves higher ELBO values in variational autoencoders (VAEs). This enhancement promotes better alignment between the decoder and the prior distribution, ultimately strengthening the model's probabilistic representation capabilities.
This work proposes an end-to-end gradient-driven Bayesian optimization framework to address the high computational cost associated with posterior sampling and acquisition function optimization in traditional Bayesian neural network–based approaches. By introducing variational mutual information estimation into Bayesian optimization for the first time and integrating it within an actor-critic architecture, the method jointly optimizes input exploration and information gain assessment. This design eliminates the inner-loop acquisition function optimization, yielding a fully differentiable and computationally efficient optimization pipeline. Empirical evaluations demonstrate that the proposed approach achieves performance comparable to or better than existing baselines across multiple high-dimensional synthetic and real-world tasks, while reducing computational overhead by up to two orders of magnitude.
Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.