🤖 AI Summary
This paper addresses the critical challenge of enabling autonomous capability enhancement in multimodal large language models (MLLMs) under low-human-effort constraints. We present the first systematic survey of self-improvement mechanisms in MLLMs, proposing a three-dimensional framework encompassing data augmentation (generation, feedback integration, and filtering), data organization (curriculum learning, memory mechanisms, and reinforcement learning), and model optimization. Innovatively, we establish a hierarchical “Data–Organization–Optimization” taxonomy to unify mainstream methodologies, evaluation paradigms, and application scenarios. We explicitly identify key bottlenecks—including cross-modal alignment bias, heavy dependence on feedback quality, and limited generalizability—as open challenges. Our work delivers the first structured research blueprint for MLLM autonomous evolution, advancing the development of cost-efficient and sustainable multimodal agents.
📝 Abstract
Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potential for leveraging diverse data sources and developing more general self-improving models. This survey is the first to provide a comprehensive overview of self-improvement in Multimodal LLMs (MLLMs). We provide a structured overview of the current literature and discuss methods from three perspectives: 1) data collection, 2) data organization, and 3) model optimization, to facilitate the further development of self-improvement in MLLMs. We also include commonly used evaluations and downstream applications. Finally, we conclude by outlining open challenges and future research directions.