🤖 AI Summary
This study addresses the challenge of factor entanglement in unsupervised subtype discovery and controllable generation for medical imaging. We propose directly disentangling salient and common factors within the high-dimensional latent space of a frozen representation autoencoder, eliminating the need for subtype labels. By integrating visual foundation model representations, a conditional diffusion Transformer is employed to guide the synthesis of novel samples. The proposed approach achieves precise separation of OCT disease categories and enables unsupervised subtype discovery. Experimental results demonstrate a reconstruction rFID below 2, a probe accuracy of 0.950, and an improved subtype accuracy of 90.5%. These findings validate the effectiveness of latent space disentanglement for high-quality, controllable image generation in medical applications.
📝 Abstract
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.