🤖 AI Summary
This study addresses the inability of conventional metrics to distinguish genuine recovery from hallucinated details in ultra-low-field MRI super-resolution. For the first time, this work quantifies the upper bound of subject-specific information at 64 mT and proposes an auditing protocol based on paired empirical data to systematically evaluate GAN, diffusion, and Transformer architectures. The research establishes an in-plane information limit of 3–4 mm and reveals that synthetic data substantially overestimate recoverability. Furthermore, it demonstrates that fine details generated by existing models are largely fictitious and exhibit no significant correlation with individual subjects. By open-sourcing the codebase, this project provides a rigorous evaluation paradigm for low-field MRI reconstruction.
📝 Abstract
Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject's own 3T scan than with a stranger's, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.