🤖 AI Summary
This study addresses the difficulty of pretrained text-to-image models in reliably acquiring specific visual identities from few-shot references while preserving compositional control. Building upon Stable Diffusion 3.5, this work proposes an external memory mechanism indexed by trigger words, which injects concept memories as relative residuals into the frozen text encoder and MMDiT contextual states via gated directions, thereby decoupling backbone adaptation from external memory. The proposed approach supports prompt-selective and merge-free multi-concept access. Experimental results demonstrate that subject fidelity is comparable to DreamBooth-LoRA with superior context preservation, significantly reduced additional loading overhead, and effective paired-trigger composition alongside intra-category disentanglement.
📝 Abstract
Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.