Score
Analyzing and estimating convergence rates of Markov chains, proving iteration-complexity bounds in metrics like total variation, and designing reversible chains or samplers that efficiently mix in high-dimensional settings for sampling and model selection.
This paper addresses the slow convergence of finite ergodic Markov chains by proposing an acceleration framework based on permutations and projections. Methodologically, it establishes—for the first time—the equivalence between the mixing time of permuted Markov chains and information-theoretic projections (measured via KL divergence or squared Frobenius norm); introduces an alternating projection algorithm that geometrically unifies the mixing process and reveals intrinsic connections to the Sylvester equation and assignment problem, yielding a trace-based criterion for stationarity. Theoretical contributions include: reducing the relaxation time from exponential to polynomial order for bimodal distributions; achieving logarithmic mixing time under random permutations—surpassing the linear bound of Diaconis–Holmes–Neal; and demonstrating significantly improved mixing efficiency over standard Metropolis–Hastings in physical model experiments.
This work investigates the decomposability and geometric structure of transition matrices for multivariate Markov chains to accelerate mixing. We model factor chains as information projections onto submanifolds under the KL divergence, establishing—for the first time—a rigorous correspondence between factorization structure and information-geometric projection. This framework unifies Han–Shearer-type inequalities and submodularity of entropy rate, while revealing intrinsic connections among entropy-rate submodularity, large-deviation rate functions, and mixing times. Building on this insight, we design a projection sampler parameterized by temperature coordinates—a geometric refinement of the exchange algorithm. Theoretical analysis shows that its mixing time improves over the classical exchange algorithm by a factor proportional to the product of the number of temperatures and the dimension of the state space, substantially enhancing sampling efficiency for high-dimensional multivariate chains.
This work studies the adaptive complexity of parallel sampling from log-concave distributions—specifically, the minimal number of sequential rounds required to achieve a prescribed accuracy, assuming polynomially many queries can be executed in parallel per round. We establish, for the first time, tight lower bounds on the number of iterations under total variation distance for both unconstrained and box-constrained settings: in the unconstrained case, exponentially small error necessitates superlinear rounds; under box constraints, even superpolynomially small error cannot be achieved by any nearly linear-round algorithm. Methodologically, we introduce a novel hardness potential framework based on chain-structured random partitions and classical smoothing techniques, which uniformly captures distributional properties including log-smoothness/Lipschitzness and strong/non-strong concavity. This is the first systematic characterization of fundamental computational bottlenecks for log-concave sampling in the parallel model.
Metropolis–Hastings samplers—especially informed variants—exhibit slow convergence and dimension-dependent relaxation times in high-dimensional discrete parameter spaces under multimodal posteriors. Method: We develop the first dimension-free mixing time theory for information-driven MCMC on discrete domains, integrating multicommodity flow analysis with single-site drift conditions and leveraging high-dimensional statistical structure to derive verifiable sufficient conditions. Contribution/Results: Our analysis yields a tight, dimension-independent upper bound on the relaxation time, breaking the conventional dimensional dependence bottleneck in convergence analysis. This provides the first theoretically grounded, computationally efficient sampling framework for discrete-parameter inference tasks—such as Bayesian model selection—where scalability with dimension is critical. The bound is constructive and applicable to a broad class of informed discrete MCMC algorithms, enabling rigorous performance guarantees without restrictive assumptions on posterior geometry or sparsity.
This work investigates the non-asymptotic convergence and statistical concentration of constant-step-size stochastic gradient descent (SGD) for smooth strongly convex optimization. Methodologically, it models SGD iterates as a Markov chain and establishes, for the first time, non-asymptotic convergence rates to its unique invariant distribution under both total variation and Wasserstein-2 distances. Theoretically, it shows that the invariant distribution inherits the tail properties (sub-Gaussian or sub-exponential) of the gradient noise, enabling high-confidence error bounds on the final iterate. In linear regression, it further derives a dimension-free upper bound on the bias of Polyak–Ruppert averaging. Compared to classical asymptotic analyses, this framework delivers sharper, more practical finite-step statistical guarantees—enhancing theoretical understanding of SGD’s robustness and generalization behavior.
This study addresses the problem of efficiently identifying the initial state in partially observable Markov chains, particularly under passive observation and limited-efficiency constraints. To this end, the authors introduce and systematically analyze, for the first time, a novel model termed “Markov chains with support for backtracking,” which permits algorithms to strategically revert the process to prior states to accelerate learning or decision-making. The key contributions include establishing the equivalence between non-adaptive and adaptive backtracking strategies in terms of state distinguishability, and constructing a non-adaptive strategy whose query complexity exceeds that of the optimal adaptive strategy by only a polynomial factor—a gap proven to be unavoidable. The theoretical analysis integrates probability theory, information theory, and computational complexity, thereby establishing a new analytical framework for reasoning about backtracking mechanisms.
This work addresses the slow convergence and strong dimension dependence of traditional MCMC methods in high- and infinite-dimensional Bayesian posterior sampling by proposing and analyzing two novel multi-proposal preconditioned Crank–Nicolson algorithms, termed mpCN and MTpCN. Leveraging parallelized proposal mechanisms, these algorithms achieve enhanced sampling efficiency and are shown to converge under non-convex, high-dimensional settings. The study establishes, for the first time, rigorous dimension-independent and proposal-number-uniform exponential convergence rates for both methods. By innovatively constructing two coupling schemes, the authors derive Wasserstein contraction, an $L^2$ spectral gap, and non-asymptotic statistical guarantees. The theory demonstrates that dimension-independent mixing is attainable without convexity assumptions, provided the log-likelihood is bounded and Lipschitz. Numerical experiments confirm faster warm-up and more robust parameter tuning, significantly outperforming standard pCN and independent parallel-chain approaches.
This work addresses the challenge of diagnosing convergence in Markov chain Monte Carlo (MCMC) methods by proposing an efficient diagnostic framework based on multi-marginal coupling. By introducing shared randomness across multiple Metropolis–Hastings chains, the authors construct a Poisson Monte Carlo estimator and develop an adaptive point-process update rule alongside a distributed matching algorithm, substantially alleviating computational bottlenecks in high-dimensional settings. The approach establishes theoretical connections to list-level distribution coupling and distributed matching problems, yielding a natural optimization objective tailored for multi-chain coupling. Experimental results demonstrate that, across various dimensional configurations, the proposed method reduces coupling time by up to 50% compared to existing baselines, achieving significantly improved diagnostic efficiency.
This work addresses the challenge of analyzing mixing times for Markov chain Monte Carlo (MCMC) algorithms under non-convex potentials and heavy-tailed target distributions by introducing a unified contraction framework based on $\mathsf{E}_\gamma$-divergence. The framework incorporates Gaussian smoothing to derive an explicit global contraction coefficient and introduces a local contraction coefficient to handle unbounded importance weights, thereby effectively accommodating heavy-tailed settings. It applies broadly to algorithms such as projected Langevin Monte Carlo and independent Metropolis–Hastings. For the former, it establishes, for the first time, explicit exponential convergence rates under KL, $\chi^2$, and Rényi divergences; for the latter, it recovers or improves existing convergence bounds under either finite-moment or heavy-tailed conditions.
This work investigates the efficient learning and application of preconditioners to enhance the sampling efficiency of Markov Chain Monte Carlo (MCMC) methods within a finite time horizon. Focusing on the Unadjusted Langevin Algorithm (ULA), the study provides non-asymptotic theoretical guarantees for preconditioners constructed using either the covariance or the expected Hessian of the target distribution, under Wasserstein-2 contraction conditions. It presents the first non-asymptotic time complexity analysis for learned preconditioned MCMC, rigorously quantifying how the upfront cost of learning a preconditioner can be amortized by subsequent gains in sampling efficiency. Furthermore, the analysis integrates classical heuristics—such as effective sample size and mixing time—into the modern theoretical framework of MCMC, thereby ensuring the generation of N approximately independent samples within a provably finite number of steps.