🤖 AI Summary
This study addresses the challenge of predicting conditional marginal distributions for missing variables under incomplete information, a task that conventionally relies on modeling the full joint distribution or incorporating domain-specific prior knowledge. To this end, this work proposes Marformer, which leverages a BERT-like attention mechanism to construct hidden representations and iteratively refine distributional features. By directly predicting the conditional marginals of missing variables via a Transformer architecture, inference is accomplished in a single forward pass. This approach transcends the limitations of traditional generative models by eliminating the need for domain expertise. Extensive experiments on both synthetic and real-world datasets demonstrate that Marformer matches or surpasses classical methods and generative baselines in predictive performance while achieving substantially improved inference efficiency.
📝 Abstract
Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one more observation before deciding. We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values. Like BERT, which is trained to predict missing words from context, the Marformer constructs a hidden-vector representation for each distribution $p(X_i)$ and iteratively refines it through attention to other distributions $p(X_j)$. Unlike generative approaches, the Marformer does not model the full joint distribution, requires no domain knowledge of the data-generating process, and makes all predictions in a single forward pass. We evaluate across three synthetic domains with missing data---Bayesian networks, discretized multivariate Gaussians, and structured annotation data. The Marformer can match or outperform classical missing-data methods, even when those methods are given the true model family and prior that generated the synthetic data. We also evaluate on a real annotation dataset, where the Marformer outperforms the evaluated baselines at the largest training size. In both cases, the Marformer is substantially faster than the evaluated generative baselines.