Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback
This study addresses the insufficient inter-modal information interaction in multimodal representation learning by proposing the MEQ architecture. This method introduces a continuous bidirectional reciprocal feedback mechanism to replace conventional feature concatenation, coupling inputs from different modalities into mutually representative embeddings and yielding their interactive fixed point as the final output. Furthermore, theoretical analysis is provided to guide network design, effectively preventing failure modes. Experimental results demonstrate that MEQ surpasses or matches existing baseline methods on classification and visual grounding tasks, achieving particularly significant improvements in localization accuracy under cross-modal complementary scenarios.