Score
Design and implement fine-tuning and adaptation pipelines for large multimodal LLMs that accept and jointly reason over images, audio/sensor signals and text (including vision–language and vision–language–action models), using techniques such as LoRA/low-rank adapters, parameter-efficient updates, two-stage or captioning-then-instruction finetuning, continual/multistage supervised training, and capability-preserving constraints. Produce the artifacts and analyses needed for deployment and evaluation—adapter checkpoints, training curricula and schedules, loss functions and regressors, and metrics to verify preservation of base-language capabilities while enabling new multimodal perception, scoring, or action-alignment for closed-loop inference.
This work addresses the parameter efficiency challenge in visual–language large models (VLLMs) for effective cross-modal fusion. We systematically analyze 34 state-of-the-art VLLMs and, for the first time, unify their training paradigms into three categories—single-stage fine-tuning, two-stage fine-tuning, and direct adaptation—establishing the first taxonomy of VLLM efficiency grounded in training methodology. Our study fills a critical gap by providing the first systematic analysis of direct adaptation, empirically demonstrating that it achieves over 90% of two-stage fine-tuning performance with less than 1% parameter overhead. We comprehensively examine core components—including LLM backbones, vision encoders, multimodal fusion architectures, parameter-efficient adaptation techniques (e.g., LoRA, Adapters), and evaluation protocols—and synthesize key benchmarks and metrics. The work delivers both a theoretical framework and empirical evidence to advance efficient multimodal modeling.
To address weak fine-grained visual understanding, inter-task data conflicts, and low efficiency in visual instruction tuning of multimodal large language models (MLLMs), this paper proposes EVIT, an efficient visual instruction tuning framework. Methodologically, EVIT introduces (1) a novel Visual Cue Enhancement (VCE) mechanism that fuses multi-level visual features to strengthen fine-grained perception, and (2) a Dual-LoRA architecture—comprising two decoupled low-rank adaptation modules—that separately parameterizes skill space and task space, enabling parameter-efficient and precisely controllable cross-task transfer. By enhancing the visual projector and integrating hierarchical visual features, EVIT significantly improves visual understanding accuracy and cross-task generalization on both general benchmarks and downstream tasks, achieving state-of-the-art performance with lightweight implementation.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
Training video large language models (Video LLMs) incurs prohibitively high computational costs and relies heavily on massive video datasets. Method: This paper proposes RED-VILLM—a resource-efficient development paradigm that leverages off-the-shelf image large language models (Image LLMs) and introduces a lightweight temporal adaptation module to enable efficient knowledge transfer, eliminating redundant architectural design and costly large-scale video pretraining. It achieves, for the first time, efficient cross-modal transfer from Image LLMs to Video LLMs via a plug-and-play temporal modeling structure, integrated with instruction tuning and multi-stage alignment training. Contribution/Results: RED-VILLM attains superior performance over conventional Video LLMs under extremely limited instruction data and computational resources, significantly reducing both computational overhead and dependency on video data. Additionally, this work releases the first open-source Video LLM tailored for the Chinese research community.
To address slow convergence and suboptimal performance in multimodal large language model (MLLM) instruction tuning—caused by imbalanced learning between the language backbone and feature encoders—this paper proposes a collaborative optimization framework. Methodologically, it introduces: (1) a novel learning balance quantification metric that explicitly models gradient dynamics of both components; (2) a state-aware generative distribution regularization loss to enhance cross-modal semantic alignment; and (3) a dynamic learning rate scheduler with a model-agnostic architecture, supporting instruction tuning across visual and audio modalities. Evaluated across diverse LLM backbones and multimodal encoders, the framework significantly accelerates training convergence and achieves state-of-the-art performance on both vision and audio understanding benchmarks.
Existing multimodal large language models (MLLMs) employ only unidirectional mapping of visual information into the language space, failing to effectively leverage visual knowledge to enhance holistic reasoning. This work introduces Vision-Augmented Large Language Models (VA-LLMs), breaking from this paradigm by enabling LLMs to actively store, share, and retrieve visual knowledge. Our contributions are threefold: (1) Modular Visual Memory (MVM), a structured, long-term storage mechanism for visual knowledge; (2) Soft Multimodal Mixture of Experts (MoME), a dynamic architecture that coordinates vision and language experts during autoregressive generation; and (3) a cross-modal knowledge injection and collaborative reasoning framework. Experiments demonstrate substantial improvements in physical commonsense understanding and spatial reasoning, achieving state-of-the-art performance across multiple multimodal benchmarks, including MMMU, ScienceQA, and POPE.
This work addresses the deployment challenges of existing parameter-efficient fine-tuning methods—such as LoRA and Soft Prompting—which require modifications to the model’s computation graph and are thus incompatible with high-throughput inference engines like vLLM. The authors propose ART, a novel approach that treats visual inputs as learnable “computational art.” By optimizing only pixel-level visual prompts while keeping all parameters of the multimodal large language model frozen, ART injects task-specific information without altering the model architecture or computation graph. Consequently, it seamlessly integrates with precompiled inference systems and supports arbitrary fine-tuning objectives. Experimental results demonstrate that ART achieves performance on par with LoRA on Qwen-series models across multiple textual benchmarks, particularly excelling in mathematical reasoning and structured tool-use tasks.
This work addresses the limitations of current vision–language foundation models and evaluation benchmarks, which are predominantly English-centric and computationally expensive, thereby hindering their applicability to low-resource languages in multimodal settings. To overcome these challenges, we propose a lightweight text–speech–vision tri-modal fusion framework that leverages a stack of efficient adapters for cross-modal alignment, coupled with cost-effective data construction and fine-tuning strategies to enable collaborative multimodal modeling under constrained computational resources. Our contributions include an open-source compact multilingual multimodal model, an end-to-end speech–text–LLM pipeline, a culturally aware evaluation benchmark, and a reproducible toolchain, collectively yielding significant improvements in multimodal understanding and generation for non-English languages.
To address the challenge of adapting vision-language models (VLMs) downstream while keeping pretrained weights frozen, this paper proposes Masked Fine-Tuning (MFT): a parameter-efficient paradigm that updates no model parameters. Instead, it introduces learnable gating scores into the language module and image–text projector to dynamically mask and reconfigure existing weight connections, thereby enabling structured subnetwork reconfiguration. This is the first work to apply masked fine-tuning to VLMs, demonstrating that merely “rewiring” frozen backbone connections—without modifying the visual encoder—suffices to unlock strong task-specific adaptability. MFT is compatible with multilingual backbones and consistently outperforms LoRA variants and full fine-tuning across multiple VLM benchmarks, while preserving the visual encoder entirely frozen. The method achieves superior performance with zero trainable parameters in the vision tower, offering a novel perspective on structural adaptation of frozen multimodal foundations. Code is publicly available.
This work addresses two key challenges in few-shot visual detection (<1,000 images): the strong data dependency of CNNs and the low fine-tuning efficiency of multimodal large language models (MLLMs). We propose a specialized fine-tuning paradigm for text-annotated detection tasks, explicitly modeling the synergy between language guidance and visual localization. Our approach integrates prompt engineering, dynamic context reasoning, and lightweight end-to-end fine-tuning. Under extreme data scarcity, the method achieves a 36% mAP improvement over conventional CNN baselines—matching or even surpassing their fully supervised performance. To our knowledge, this is the first work to empirically demonstrate the strong generalization capability of MLLMs in domain-specific few-shot visual detection. It establishes a novel cross-modal few-shot learning paradigm grounded in task-aware textual supervision. The implementation is publicly available.