Self-Improvement in Multimodal Large Language Models: A Survey

📅 2025-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the critical challenge of enabling autonomous capability enhancement in multimodal large language models (MLLMs) under low-human-effort constraints. We present the first systematic survey of self-improvement mechanisms in MLLMs, proposing a three-dimensional framework encompassing data augmentation (generation, feedback integration, and filtering), data organization (curriculum learning, memory mechanisms, and reinforcement learning), and model optimization. Innovatively, we establish a hierarchical “Data–Organization–Optimization” taxonomy to unify mainstream methodologies, evaluation paradigms, and application scenarios. We explicitly identify key bottlenecks—including cross-modal alignment bias, heavy dependence on feedback quality, and limited generalizability—as open challenges. Our work delivers the first structured research blueprint for MLLM autonomous evolution, advancing the development of cost-efficient and sustainable multimodal agents.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPMultiagent Systems: Multiagent Learning

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potential for leveraging diverse data sources and developing more general self-improving models. This survey is the first to provide a comprehensive overview of self-improvement in Multimodal LLMs (MLLMs). We provide a structured overview of the current literature and discuss methods from three perspectives: 1) data collection, 2) data organization, and 3) model optimization, to facilitate the further development of self-improvement in MLLMs. We also include commonly used evaluations and downstream applications. Finally, we conclude by outlining open challenges and future research directions.
Problem

Research questions and friction points this paper is trying to address.

Surveying self-improvement methods for Multimodal Large Language Models
Structuring literature on data collection and model optimization
Identifying challenges and future directions for MLLM development
Innovation

Methods, ideas, or system contributions that make the work stand out.

Surveying multimodal self-improvement methods systematically
Organizing methods by data collection and model optimization
Outlining evaluations for multimodal self-improving models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.