mutual information estimation

Designs, implements, and evaluates algorithms and estimators that compute and analyze mutual information and pointwise mutual information between variables, including numeric, nonparametric, and neural estimators (e.g., MINE), and methods that handle continuous–discrete mixtures, quantized or repeated values, and cross-modal signals. Also develops time-aware estimators and analyses to measure mutual information across time slices or events, and techniques for mutual-information minimization used for representation debiasing and for quantifying uncertainty reduction and dependence without discretization.

mutualinformationestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural Mutual Information Estimation with Vector Copulas

Oct 23, 2025
YC
Yanzhi Chen
🏛️ University of Cambridge | Imperial College London | Alan Turing Institute | University of Edinburgh

Mutual information (MI) estimation is fundamental in data science, yet existing approaches struggle to reconcile model flexibility with statistical interpretability: neural estimators require substantial data, while classical models (e.g., Gaussian copulas) fail to capture complex, high-order dependencies. This paper introduces the first neural MI estimation framework grounded in vector vine structures—marking the first integration of vector vine theory into deep learning–based estimation. Methodologically, we employ neural networks to parameterize marginal transformations and leverage vector vines to explicitly model intricate multivariate dependence structures; training is performed end-to-end via a variational lower bound and density-ratio estimation. Evaluated on synthetic benchmarks and multimodal real-world datasets, our approach achieves significant gains in estimation accuracy, robustness, and generalization—effectively balancing expressive power and statistical interpretability.

Achieving better trade-off between model complexity and capacityAddressing limitations of existing overly complex or simplified estimatorsEstimating mutual information in data science and machine learning

Accurate Estimation of Mutual Information in High Dimensional Data

May 31, 2025
EA
Eslam Abdelaleem
🏛️ Emory University

Accurate mutual information (MI) estimation in high-dimensional data has long been hindered by severe undersampling and the absence of principled reliability assessment criteria. This paper introduces the first MI estimation protocol with explicit reliability testing and statistical consistency guarantees. Methodologically, it integrates self-supervised representation learning, embedding dimensionality analysis, training dynamics monitoring, and statistical consistency verification. Crucially, we establish the first theoretical result showing that reliable MI estimation remains feasible—even under extreme undersampling—whenever the high-dimensional data admits an exact low-dimensional representation. Building upon this insight, we develop the first verifiable framework for quantifying MI estimator quality. Extensive evaluation on multiple high-dimensional benchmark datasets demonstrates controlled estimation error, strong robustness, and significantly expanded practical applicability of MI in real-world high-dimensional data analysis.

Accurate MI estimation in high-dimensional data is challengingExisting methods lack reliability checks for estimator accuracyLow-dimensional representations enable reliable MI estimation in undersampled data

Existing methods struggle to robustly quantify dependencies between continuous time series and discrete event sequences, often suffering from quantization errors, repeated values, and event redundancy. This work proposes the first non-parametric mutual information estimation framework that operates without data transformation or learning, directly modeling the continuous–discrete duality. It further introduces latent event clustering to mitigate co-occurrence bias. Evaluated on both synthetic and real-world datasets, the method consistently outperforms state-of-the-art approaches across tasks including causal analysis, recurring pattern discovery, and covariate selection, achieving notable improvements in accuracy, robustness, and interpretability.

dependence quantificationheterogeneous datamutual information

FMMI: Flow Matching Mutual Information Estimation

Nov 11, 2025
IB
I. Butakov
🏛️ Applied AI Institute | Institute of Numerical Mathematics, RAS

Estimating mutual information (MI) in high-dimensional settings suffers from low efficiency, poor accuracy, and limited stability—especially when MI values span several orders of magnitude. To address this, we propose the first normalized flow-based MI estimator grounded in flow matching (FM). Unlike prevailing discriminative approaches, our method directly models the invertible density transformation between the joint and marginal distributions via normalizing flows, enabling end-to-end, differentiable MI estimation. By integrating flow matching into the MI estimation framework, we achieve both theoretical rigor—guaranteeing unbiased gradient estimation—and computational scalability. Experiments on multivariate benchmark tasks demonstrate that our estimator significantly outperforms baselines including InfoNCE and MINE in estimation accuracy, converges faster, incurs lower computational overhead, and exhibits strong robustness to both extremely small and large MI values.

Creating scalable mutual information estimation across various ground-truth valuesEstimating mutual information efficiently in high dimensionsTransforming distributions using normalizing flows instead of classifiers

Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions

Jul 03, 2024
CK
Christian Keup
🏛️ École Polytechnique Fédérale de Lausanne (EPFL)

This paper investigates the fundamental limits and algorithmic design of joint inference in high-dimensional multimodal learning: given two noisy, correlated spiked data matrices, how can shared latent variables be optimally recovered? First, it rigorously characterizes the Bayesian optimal recovery threshold under generalized priors and heterogeneous noise. Second, it proves that classical methods—Partial Least Squares (PLS) and Canonical Correlation Analysis (CCA)—exhibit suboptimal phase transitions and fail to achieve this theoretical limit. Third, it proposes a joint estimation algorithm based on Approximate Message Passing (AMP), with rigorous performance characterization via state evolution analysis and numerical validation. The AMP algorithm achieves Bayesian-optimal recovery with linear time complexity, significantly outperforming PLS and CCA. Collectively, this work establishes precise statistical boundaries for multimodal fusion and provides a principled algorithmic pathway to attain them.

Analyzing multi-modal inference performance gain over isolated modalitiesComparing AMP algorithm performance against sub-optimal PLS and CCA methodsDeriving optimal thresholds for recovering latent structures from noisy data

Latest Papers

What's happening recently
View more

Existing benchmarks for mutual information estimation are largely confined to low-dimensional, simplified distributions, limiting their ability to evaluate estimator performance on complex real-world data. This work proposes a unified benchmark framework grounded in copula theory, comprising two complementary test suites: the first systematically controls mutual information, dimensionality, and marginal complexity using synthetic data and flow-based models; the second integrates real images with controllable dependency structures, extending the classic same-pair paradigm. For the first time, this framework jointly encompasses both synthetic and real data while accounting for both dependency structure and marginal complexity. Systematic evaluation of diverse discriminative and generative estimators reveals that no single method dominates across all settings, and that fundamental limitations exist across different estimator families—limitations more effectively exposed by the newly designed tests.

BenchmarkingEstimator EvaluationHigh-dimensional Data

This work addresses the challenge of accurately estimating mutual information (MI) between discrete random variables by proposing a novel approach based on Discrete Bridge Matching. For the first time, discrete bridge models are introduced into MI estimation, reframing the problem as a domain translation task. The resulting DBMI estimator integrates generative modeling with information-theoretic techniques, specifically tailored for discrete data. Experimental results demonstrate that the proposed method significantly outperforms existing approaches in both low-dimensional and image-based MI estimation tasks, confirming its effectiveness and strong generalization capability.

discrete datadiscrete random variablesinformation theory

Estimating mutual information between continuous high-dimensional variables remains challenging due to its unbounded nature and sensitivity to dimensionality, which compromises both accuracy and comparability; existing normalization methods further suffer from numerical instability in high dimensions. This work proposes the first fully neural normalized mutual information estimator, leveraging the Donsker–Varadhan representation within a MINE-like architecture to estimate mutual information. It introduces the MI-NEE framework, wherein neural networks learn the divergence between marginal distributions and a uniform reference distribution to stably recover entropy terms, thereby yielding a bounded and comparable normalized metric. Experiments on Gaussian data across dimensions 1 to 8 demonstrate that the proposed method significantly outperforms the KSG baseline, confirming its accuracy and robustness in estimating normalized dependence for high-dimensional continuous variables.

continuous multidimensional variablesdimensionality sensitivitymutual information estimation

This study addresses the fundamental challenge of achieving optimal function approximation under limited evaluation data—a central problem in numerical analysis and machine learning. From the perspective of information-based complexity, the work systematically investigates function recovery under generalized sampling by integrating information-theoretic analysis, optimal recovery theory, and nonlinear, adaptive, and randomized sampling mechanisms. It uncovers intrinsic connections among diverse sampling strategies, characterizes the information-theoretic limits of function approximation given finite data, and proposes efficient algorithms and sampling schemes that approach these limits. The results provide foundational insights and a unified framework for optimal sampling theory.

function approximationfunction evaluationinformation-based complexity

Hot Scholars

WZ

Weizhan Zhang

Professor,Department of Computer Science and Technology, Xi'an Jiaotong University
Multimedia networking
JG

Jiafeng Guo

Professor, Institute of Computing Techonology, CAS
Information RetrievalMachine LearningText AnalysisNeuIR
XC

Xueqi Cheng

Ph.D. student, Florida State University
Data miningLLMGNNComputational social science
TP

Tiago Pimentel

ETH Zurich
Natural Language ProcessingComputational LinguisticsInformation TheoryMachine Learning
AO

Alfonso Ortega

Universidad de Zaragoza, Instituto de Investigación en Ingeniería de Aragón (I3A)
Speech ProcessingSpeech TechnologiesSpeaker RecognitionSpeaker Diarization