🤖 AI Summary
This study addresses the asymmetry and coupling of sequential and morphological conditions in fluorescence microscopy image generation by proposing FACET, a framework for high-fidelity localization synthesis of unimaged proteins. FACET introduces explicit inductive biases to disentangle conditional explanatory power through a factorized asymmetric encoding architecture. It leverages cross-protein semantic memory to share coarse-grained patterns while modeling fine-grained bounded residuals, combined with variance-preserving state projection for efficient diffusion transport. Experimental results demonstrate that, with minimal parameter overhead, FACET improves spatial overlap by 34.3%, reduces FID by 27.2%, and decreases network evaluation calls by 75%, significantly enhancing the recovery of biologically relevant structures and predictive calibration.
📝 Abstract
Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged proteins. The two conditions, however, play asymmetric roles: morphological context is spatially aligned with the target, whereas sequence is non-spatial and must specify protein-dependent localization within it, with recurring coarse patterns shared across proteins and finer protein-specific variation. Existing generators condition on both jointly, without separating what each explains. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic generative framework that encodes this structure as an explicit inductive bias: sequence semantics are learned from what context leaves unexplained, coarse localization regularities are shared across proteins through a semantic memory, and protein-specific variation is a bounded residual around them. A variance-preserving state projection further lets FACET perform continuous stochastic transport through a pretrained diffusion predictor with minimal parameter overhead. On held-out proteins, FACET improves spatial overlap by 34.3% on the Human Protein Atlas and 14.0% on OpenCell over a backbone-matched baseline, and reduces FID by 27.2% and 46.5%, respectively, with 75% fewer network evaluations. It also substantially improves protein-association structure recovery and yields better-calibrated predictions, while detailed ablations show complementary contributions from its design choices. These results identify factorized asymmetric conditioning, rather than generator capacity alone, as a key lever for high-fidelity, efficient, and biologically meaningful cellular image synthesis.