SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of factor entanglement in unsupervised subtype discovery and controllable generation for medical imaging. We propose directly disentangling salient and common factors within the high-dimensional latent space of a frozen representation autoencoder, eliminating the need for subtype labels. By integrating visual foundation model representations, a conditional diffusion Transformer is employed to guide the synthesis of novel samples. The proposed approach achieves precise separation of OCT disease categories and enables unsupervised subtype discovery. Experimental results demonstrate a reconstruction rFID below 2, a probe accuracy of 0.950, and an improved subtype accuracy of 90.5%. These findings validate the effectiveness of latent space disentanglement for high-quality, controllable image generation in medical applications.
📝 Abstract
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.
Problem

Research questions and friction points this paper is trying to address.

contrastive analysis
salient factor discovery
unsupervised subtype discovery
visual representation
conditional generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Analysis
Salient Factor Discovery
Diffusion Transformer
Visual Foundation Representations
Unsupervised Subtype Discovery
💼 Related Jobs
No related jobs found.
S
Shuang Liang
HKU
L
Lejun Liao
Boston College
S
Shiyuan Zhang
University of Virginia
M
Max C. Zhang
Boston College
X
Xiaolong Luo
Harvard University
H
Han Wang
HKU
Stefano Anzellotti
Stefano Anzellotti
Boston College
Yuan Yuan
Yuan Yuan
Assistant Professor in Computer Science at Boston College, previously at MIT.
Machine LearningComputer VisionMedical AIArtificial Intelligence