🤖 AI Summary
This study addresses the trade-off between data acquisition cost and information preservation by proposing an information-theoretic framework grounded in discrete diffusion priors. The core innovation lies in pioneering the use of a frozen D3PM as an entropy surrogate, combined with side information to train a one-shot mask generator. A sequential greedy algorithm is then employed to maximize mutual information under budget constraints, enabling active data selection. Experimental results demonstrate that the proposed method reduces reconstruction error eightfold on MNIST and improves PSNR by 3.4 dB on CIFAR-10. Furthermore, for MRI reconstruction, its static masks significantly outperform mainstream approaches such as variable-density sampling and LOUPE. Ultimately, this work achieves efficient, task-agnostic observation across diverse applications.
📝 Abstract
Acquiring data is costly: higher measurement fidelity costs power and storage and risks collecting irrelevant content, while aggressive cost reduction can discard information that later analysis needs. We address this trade-off with an information-theoretic framework that acquires data relevant to a broad set of tasks rather than to one model. A mask policy, conditioned on side information, chooses which pixels to measure so as to maximize the mutual information between a discrete image and its partial observation under a budget; since the image entropy does not depend on the mask, this is equivalent to minimizing the conditional entropy. A frozen discrete denoising diffusion model (D3PM) supplies the posterior, and we use it in two ways: as an entropy surrogate for training a one-shot mask generator, and as the criterion for sequential greedy acquisition. The one-shot generator outperforms random masks only with care, including an unbiased gradient estimator for binary masks. With sequential acquisition, on MNIST the prior makes $8\times$ fewer errors than random at a $10\%$ budget, and on CIFAR-10 it gains $0.9$--$3.4$~dB. On fastMRI, our proposed technique using a static mask outperforms the well-known methods such as variable density and LOUPE.