Score
Designs and builds training pipelines, datasets, and model variants that align vision and vision–language models to follow natural‑language instructions by converting annotations into instruction–response pairs, generating or mixing synthetic multimodal corpora, and training compact VLMs or multimodal LLMs on mixed text–image targets. Analyzes and implements tuning strategies—including decoupled or task‑specific instruction schemas, auxiliary‑data mixing, and separate optimization schedules—to preserve task‑specific trajectories, reduce interference, and improve joint training stability for instruction‑following across visual and multimodal tasks.
This work addresses the parameter efficiency challenge in visual–language large models (VLLMs) for effective cross-modal fusion. We systematically analyze 34 state-of-the-art VLLMs and, for the first time, unify their training paradigms into three categories—single-stage fine-tuning, two-stage fine-tuning, and direct adaptation—establishing the first taxonomy of VLLM efficiency grounded in training methodology. Our study fills a critical gap by providing the first systematic analysis of direct adaptation, empirically demonstrating that it achieves over 90% of two-stage fine-tuning performance with less than 1% parameter overhead. We comprehensively examine core components—including LLM backbones, vision encoders, multimodal fusion architectures, parameter-efficient adaptation techniques (e.g., LoRA, Adapters), and evaluation protocols—and synthesize key benchmarks and metrics. The work delivers both a theoretical framework and empirical evidence to advance efficient multimodal modeling.
To address slow convergence and suboptimal performance in multimodal large language model (MLLM) instruction tuning—caused by imbalanced learning between the language backbone and feature encoders—this paper proposes a collaborative optimization framework. Methodologically, it introduces: (1) a novel learning balance quantification metric that explicitly models gradient dynamics of both components; (2) a state-aware generative distribution regularization loss to enhance cross-modal semantic alignment; and (3) a dynamic learning rate scheduler with a model-agnostic architecture, supporting instruction tuning across visual and audio modalities. Evaluated across diverse LLM backbones and multimodal encoders, the framework significantly accelerates training convergence and achieves state-of-the-art performance on both vision and audio understanding benchmarks.
Existing evaluation benchmarks inadequately assess multimodal large language models’ (MLLMs) capability to follow complex, hierarchical instructions. Method: We introduce MIA-Bench—a rigorously constructed benchmark comprising 400 image-instruction pairs—and propose “instruction fidelity” as a novel, core evaluation dimension. We design structured instruction challenges and a fine-grained compliance assessment protocol. High-quality test samples are generated via human-crafted templates and pattern constraints; model adherence is improved through supervised fine-tuning and instruction-augmented data. Contribution/Results: Experiments reveal substantial performance gaps among state-of-the-art MLLMs on MIA-Bench. Targeted fine-tuning boosts average instruction compliance by 23.6% without degrading general vision-language capabilities. This work establishes a new paradigm for systematic evaluation and optimization of MLLM instruction-following behavior.
This work addresses the challenge of enhancing zero-shot generalization in multimodal large language models (MLLMs) through principled visual instruction design. We propose a novel three-stage automated instruction construction paradigm—“Synthesize-Complexify-Reconstruct”—and empirically establish, for the first time, a positive correlation between instruction complexity and model performance. We introduce ComVint, the first high-quality instruction-tuning dataset tailored for complex visual reasoning, comprising 32K samples. Our method integrates vision-language joint modeling with LLM-based instruction reconstruction and quality assurance. On the MME-Perception and MME-Cognition benchmarks, our approach improves LLaVA’s performance by 27.86% and 27.60%, respectively, and delivers consistent, significant gains across four mainstream MLLMs. The code and ComVint dataset are publicly released.
To address the scarcity of high-quality, human-annotated video data for training large video-language models, this paper introduces the first large-scale synthetic data generation paradigm tailored for video instruction following. We construct LLaVA-Video-178K—a high-fidelity synthetic video instruction dataset comprising fine-grained captioning, open-ended question answering, and multiple-choice QA—and propose a unified cross-task format with mixed-data training. Our approach leverages multimodal large models for automated data synthesis and employs end-to-end video-language joint instruction tuning to align visual understanding with linguistic instructions. The resulting model, LLaVA-Video, achieves state-of-the-art performance across multiple video understanding benchmarks, demonstrating that synthetically generated data can match the efficacy of real human annotations. To foster reproducibility and community advancement, we fully open-source the dataset, data generation pipeline, and model checkpoints.
To address pervasive hallucination in large vision-language models (LVLMs) during cross-modal generation, this paper proposes a post-training alignment framework leveraging synthetically generated human preference data. The method introduces two key innovations: (1) the first controllable synthetic preference data generation mechanism specifically designed for LVLMs; and (2) the first use of a learnable reward model—replacing manual annotations or fixed metrics (e.g., CLIP)—as a human preference proxy, enabling teacher-free multimodal Direct Preference Optimization (DPO). Evaluated on LLaVA-1.5-7B, the approach achieves 87.6% accuracy and 97.8% precision on POPE, improves MMHal-Bench score from 2.36 to 3.49, and reduces hallucination rate by 50.98% (from 51.0% to 25.0%).
This work addresses the underperformance of multimodal large language models on fine-grained visual reasoning tasks, which stems primarily from their overreliance on linguistic priors during instruction tuning at the expense of visual information. The authors propose an innovative approach that reformulates classic self-supervised vision tasks—such as rotation prediction and color matching—into image-instruction-answer triplets, integrating them into the visual instruction tuning process via natural language instructions. By merely adjusting the training data distribution with just 3%–10% visually grounded instructions, the method effectively steers models to base their responses on visual evidence. This strategy requires no architectural modifications or additional training stages, yet consistently yields significant performance gains on vision-centric benchmarks across multiple models.
Existing multimodal instruction-following evaluation benchmarks primarily focus on textual instructions and overlook implicit constraints embedded in the visual modality, thereby failing to comprehensively assess models’ alignment capabilities under joint vision-language instructions. To address this gap, this work proposes VC-IFEval—the first vision-centric instruction-following evaluation framework—which systematically constructs an instruction dataset incorporating vision-dependent constraints to enable fine-grained assessment of multimodal large language models (MLLMs). Experiments demonstrate that this benchmark effectively uncovers deficiencies of current models in handling vision-related instructions and that targeted fine-tuning significantly improves both instruction-following accuracy and output consistency, thereby filling a critical void in evaluating visual constraint adherence within multimodal instruction understanding.
This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.
This work investigates how visual instruction tuning integrates image information into the hierarchical architecture of large language models—a mechanism not well understood in prior studies. Through a systematic analysis employing probing, causal intervention, representational geometry comparison, and selective layer fine-tuning, the study reveals that visual features are predominantly embedded in the model’s intermediate semantic layers in a localized manner. Building on this insight, the authors demonstrate that fine-tuning only these intermediate layers achieves performance comparable to full-model fine-tuning across multiple vision-centric benchmarks, while substantially reducing training costs. These findings underscore the pivotal role of intermediate layers in cross-modal alignment and offer an efficient strategy for multimodal adaptation.
This work addresses the challenge of efficiently equipping vision-language models with continuously evolving domain-specific skills, a task hindered by the high cost of conventional fine-tuning. The authors propose injecting capabilities from domain-specialized large language models into vision-language models through model fusion, enabling cross-modal skill transfer without additional training data or substantial computational resources. The study presents the first systematic analysis of the method’s applicability, fusion strategies, and hyperparameter sensitivity, offering quantitative evaluations of techniques such as Task Arithmetic (TA) and DARE across heterogeneous architectures. Experimental results demonstrate strong performance in instruction-following and cross-lingual tasks, while revealing limitations in mathematical reasoning, thereby delineating the effective boundaries and critical tuning factors for cross-modal skill injection.