mixing time analysis

Analyzing and estimating convergence rates of Markov chains, proving iteration-complexity bounds in metrics like total variation, and designing reversible chains or samplers that efficiently mix in high-dimensional settings for sampling and model selection.

mixingtimeanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Improving the convergence of Markov chains via permutations and projections

Nov 13, 2024
MC
Michael C. H. Choi
🏛️ National University of Singapore | University College London

This paper addresses the slow convergence of finite ergodic Markov chains by proposing an acceleration framework based on permutations and projections. Methodologically, it establishes—for the first time—the equivalence between the mixing time of permuted Markov chains and information-theoretic projections (measured via KL divergence or squared Frobenius norm); introduces an alternating projection algorithm that geometrically unifies the mixing process and reveals intrinsic connections to the Sylvester equation and assignment problem, yielding a trace-based criterion for stationarity. Theoretical contributions include: reducing the relaxation time from exponential to polynomial order for bimodal distributions; achieving logarithmic mixing time under random permutations—surpassing the linear bound of Diaconis–Holmes–Neal; and demonstrating significantly improved mixing efficiency over standard Metropolis–Hastings in physical model experiments.

Analyzing mixing properties and convergence rates of projection samplersDemonstrating improved mixing performance over Metropolis-Hastings in examplesImproving convergence of ergodic Markov chains via permutations and projections

Geometry and factorization of multivariate Markov chains with applications to the swapping algorithm

Apr 19, 2024
MC
Michael C. H. Choi
🏛️ National University of Singapore | Waseda University

This work investigates the decomposability and geometric structure of transition matrices for multivariate Markov chains to accelerate mixing. We model factor chains as information projections onto submanifolds under the KL divergence, establishing—for the first time—a rigorous correspondence between factorization structure and information-geometric projection. This framework unifies Han–Shearer-type inequalities and submodularity of entropy rate, while revealing intrinsic connections among entropy-rate submodularity, large-deviation rate functions, and mixing times. Building on this insight, we design a projection sampler parameterized by temperature coordinates—a geometric refinement of the exchange algorithm. Theoretical analysis shows that its mixing time improves over the classical exchange algorithm by a factor proportional to the product of the number of temperatures and the dimension of the state space, substantially enhancing sampling efficiency for high-dimensional multivariate chains.

Analyzes factorizability of Markov chainsDemonstrates mixing time improvement with swappingIntroduces projection sampler for acceleration

Adaptive complexity of log-concave sampling

Aug 23, 2024
HZ
Huanjian Zhou
🏛️ The University of Tokyo | RIKEN AIP | The Chinese University of Hong Kong | Vector Institute

This work studies the adaptive complexity of parallel sampling from log-concave distributions—specifically, the minimal number of sequential rounds required to achieve a prescribed accuracy, assuming polynomially many queries can be executed in parallel per round. We establish, for the first time, tight lower bounds on the number of iterations under total variation distance for both unconstrained and box-constrained settings: in the unconstrained case, exponentially small error necessitates superlinear rounds; under box constraints, even superpolynomially small error cannot be achieved by any nearly linear-round algorithm. Methodologically, we introduce a novel hardness potential framework based on chain-structured random partitions and classical smoothing techniques, which uniformly captures distributional properties including log-smoothness/Lipschitzness and strong/non-strong concavity. This is the first systematic characterization of fundamental computational bottlenecks for log-concave sampling in the parallel model.

Analyze limitations of linear iteration algorithms for samplingInvestigate error bounds for constrained and unconstrained distributionsStudy adaptive complexity of parallelized log-concave sampling

Dimension-free Relaxation Times of Informed MCMC Samplers on Discrete Spaces

Apr 05, 2024
HC
Hyunwoong Chang
🏛️ The University of Texas at Dallas | Texas A&M University

Metropolis–Hastings samplers—especially informed variants—exhibit slow convergence and dimension-dependent relaxation times in high-dimensional discrete parameter spaces under multimodal posteriors. Method: We develop the first dimension-free mixing time theory for information-driven MCMC on discrete domains, integrating multicommodity flow analysis with single-site drift conditions and leveraging high-dimensional statistical structure to derive verifiable sufficient conditions. Contribution/Results: Our analysis yields a tight, dimension-independent upper bound on the relaxation time, breaking the conventional dimensional dependence bottleneck in convergence analysis. This provides the first theoretically grounded, computationally efficient sampling framework for discrete-parameter inference tasks—such as Bayesian model selection—where scalability with dimension is critical. The bound is constructive and applicable to a broad class of informed discrete MCMC algorithms, enabling rigorous performance guarantees without restrictive assumptions on posterior geometry or sparsity.

Addressing multimodal posteriors in Bayesian model selection problemsAnalyzing convergence of MCMC in high-dimensional discrete spacesEstablishing dimension-free relaxation times for Metropolis-Hastings algorithms

Convergence and concentration properties of constant step-size SGD through Markov chains

Jun 20, 2023
IM
Ibrahim Merad
🏛️ LPSM | Université Paris Cité | DMA | École normale supérieure

This work investigates the non-asymptotic convergence and statistical concentration of constant-step-size stochastic gradient descent (SGD) for smooth strongly convex optimization. Methodologically, it models SGD iterates as a Markov chain and establishes, for the first time, non-asymptotic convergence rates to its unique invariant distribution under both total variation and Wasserstein-2 distances. Theoretically, it shows that the invariant distribution inherits the tail properties (sub-Gaussian or sub-exponential) of the gradient noise, enabling high-confidence error bounds on the final iterate. In linear regression, it further derives a dimension-free upper bound on the bias of Polyak–Ruppert averaging. Compared to classical asymptotic analyses, this framework delivers sharper, more practical finite-step statistical guarantees—enhancing theoretical understanding of SGD’s robustness and generalization behavior.

Analyzes constant step-size SGD convergence via Markov chain theoryDerives dimension-free concentration bounds for SGD iterates and averagesEstablishes non-asymptotic convergence in total variation and Wasserstein distances

Latest Papers

What's happening recently
View more

This study addresses the problem of efficiently identifying the initial state in partially observable Markov chains, particularly under passive observation and limited-efficiency constraints. To this end, the authors introduce and systematically analyze, for the first time, a novel model termed “Markov chains with support for backtracking,” which permits algorithms to strategically revert the process to prior states to accelerate learning or decision-making. The key contributions include establishing the equivalence between non-adaptive and adaptive backtracking strategies in terms of state distinguishability, and constructing a non-adaptive strategy whose query complexity exceeds that of the optimal adaptive strategy by only a polynomial factor—a gap proven to be unavoidable. The theoretical analysis integrates probability theory, information theory, and computational complexity, thereby establishing a new analytical framework for reasoning about backtracking mechanisms.

Markov ChainsPartial ObservabilityQuery Complexity

This work addresses the slow convergence and strong dimension dependence of traditional MCMC methods in high- and infinite-dimensional Bayesian posterior sampling by proposing and analyzing two novel multi-proposal preconditioned Crank–Nicolson algorithms, termed mpCN and MTpCN. Leveraging parallelized proposal mechanisms, these algorithms achieve enhanced sampling efficiency and are shown to converge under non-convex, high-dimensional settings. The study establishes, for the first time, rigorous dimension-independent and proposal-number-uniform exponential convergence rates for both methods. By innovatively constructing two coupling schemes, the authors derive Wasserstein contraction, an $L^2$ spectral gap, and non-asymptotic statistical guarantees. The theory demonstrates that dimension-independent mixing is attainable without convexity assumptions, provided the log-likelihood is bounded and Lipschitz. Numerical experiments confirm faster warm-up and more robust parameter tuning, significantly outperforming standard pCN and independent parallel-chain approaches.

dimension-free samplinginfinite-dimensional samplingMarkov Chain Monte Carlo

This work addresses the challenge of diagnosing convergence in Markov chain Monte Carlo (MCMC) methods by proposing an efficient diagnostic framework based on multi-marginal coupling. By introducing shared randomness across multiple Metropolis–Hastings chains, the authors construct a Poisson Monte Carlo estimator and develop an adaptive point-process update rule alongside a distributed matching algorithm, substantially alleviating computational bottlenecks in high-dimensional settings. The approach establishes theoretical connections to list-level distribution coupling and distributed matching problems, yielding a natural optimization objective tailored for multi-chain coupling. Experimental results demonstrate that, across various dimensional configurations, the proposed method reduces coupling time by up to 50% compared to existing baselines, achieving significantly improved diagnostic efficiency.

coalescenceconvergence diagnosisMarkov chain Monte Carlo

This work addresses the challenge of analyzing mixing times for Markov chain Monte Carlo (MCMC) algorithms under non-convex potentials and heavy-tailed target distributions by introducing a unified contraction framework based on $\mathsf{E}_\gamma$-divergence. The framework incorporates Gaussian smoothing to derive an explicit global contraction coefficient and introduces a local contraction coefficient to handle unbounded importance weights, thereby effectively accommodating heavy-tailed settings. It applies broadly to algorithms such as projected Langevin Monte Carlo and independent Metropolis–Hastings. For the former, it establishes, for the first time, explicit exponential convergence rates under KL, $\chi^2$, and Rényi divergences; for the latter, it recovers or improves existing convergence bounds under either finite-moment or heavy-tailed conditions.

contraction coefficientsconvergence guaranteesheavy-tailed distributions

This work investigates the efficient learning and application of preconditioners to enhance the sampling efficiency of Markov Chain Monte Carlo (MCMC) methods within a finite time horizon. Focusing on the Unadjusted Langevin Algorithm (ULA), the study provides non-asymptotic theoretical guarantees for preconditioners constructed using either the covariance or the expected Hessian of the target distribution, under Wasserstein-2 contraction conditions. It presents the first non-asymptotic time complexity analysis for learned preconditioned MCMC, rigorously quantifying how the upfront cost of learning a preconditioner can be amortized by subsequent gains in sampling efficiency. Furthermore, the analysis integrates classical heuristics—such as effective sample size and mixing time—into the modern theoretical framework of MCMC, thereby ensuring the generation of N approximately independent samples within a provably finite number of steps.

computational costMCMCnon-asymptotic analysis

Hot Scholars

YY

Yitong Yin

Nanjing University
theoretical computer science
WF

Weiming Feng

The University of Hong Kong
randomized algorithms
RS

Robert Schober

Friedrich-Alexander-University Erlangen-Nuremberg
MB

Mario Beraha

Department of Economics, Management and Statistics, University of Milano-Bicocca
Bayesian statisticsBayesian nonparametricsWasserstein metric
VJ

Vahid Jamali

Technical University of Darmstadt
Wireless CommunicationsMolecular CommunicationResilience