fine-tune multimodal llms

Design and implement fine-tuning and adaptation pipelines for large multimodal LLMs that accept and jointly reason over images, audio/sensor signals and text (including vision–language and vision–language–action models), using techniques such as LoRA/low-rank adapters, parameter-efficient updates, two-stage or captioning-then-instruction finetuning, continual/multistage supervised training, and capability-preserving constraints. Produce the artifacts and analyses needed for deployment and evaluation—adapter checkpoints, training curricula and schedules, loss functions and regressors, and metrics to verify preservation of base-language capabilities while enabling new multimodal perception, scoring, or action-alignment for closed-loop inference.

fine-tunemultimodalllms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address weak fine-grained visual understanding, inter-task data conflicts, and low efficiency in visual instruction tuning of multimodal large language models (MLLMs), this paper proposes EVIT, an efficient visual instruction tuning framework. Methodologically, EVIT introduces (1) a novel Visual Cue Enhancement (VCE) mechanism that fuses multi-level visual features to strengthen fine-grained perception, and (2) a Dual-LoRA architecture—comprising two decoupled low-rank adaptation modules—that separately parameterizes skill space and task space, enabling parameter-efficient and precisely controllable cross-task transfer. By enhancing the visual projector and integrating hierarchical visual features, EVIT significantly improves visual understanding accuracy and cross-task generalization on both general benchmarks and downstream tasks, achieving state-of-the-art performance with lightweight implementation.

Adapt MLLMs to tasks with minimal computational overheadEnhance vision-language projection with local detailsResolve data conflicts in diverse complex tasks

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

From Image to Video, what do we need in multimodal LLMs?

Apr 18, 2024
SH
Suyuan Huang
🏛️ Beihang University | Xiaohongshu

Training video large language models (Video LLMs) incurs prohibitively high computational costs and relies heavily on massive video datasets. Method: This paper proposes RED-VILLM—a resource-efficient development paradigm that leverages off-the-shelf image large language models (Image LLMs) and introduces a lightweight temporal adaptation module to enable efficient knowledge transfer, eliminating redundant architectural design and costly large-scale video pretraining. It achieves, for the first time, efficient cross-modal transfer from Image LLMs to Video LLMs via a plug-and-play temporal modeling structure, integrated with instruction tuning and multi-stage alignment training. Contribution/Results: RED-VILLM attains superior performance over conventional Video LLMs under extremely limited instruction data and computational resources, significantly reducing both computational overhead and dependency on video data. Additionally, this work releases the first open-source Video LLM tailored for the Chinese research community.

Efficiently develop Video LLMs using Image LLMsLeverage Image LLMs' knowledge for temporal video understandingMinimize training data and resources for Video LLMs

CoMMIT: Coordinated Instruction Tuning for Multimodal Large Language Models

Jul 29, 2024
JW
Junda Wu
🏛️ University of California, San Diego | Adobe Research | The University of New South Wales | CSIRO’s Data61

To address slow convergence and suboptimal performance in multimodal large language model (MLLM) instruction tuning—caused by imbalanced learning between the language backbone and feature encoders—this paper proposes a collaborative optimization framework. Methodologically, it introduces: (1) a novel learning balance quantification metric that explicitly models gradient dynamics of both components; (2) a state-aware generative distribution regularization loss to enhance cross-modal semantic alignment; and (3) a dynamic learning rate scheduler with a model-agnostic architecture, supporting instruction tuning across visual and audio modalities. Evaluated across diverse LLM backbones and multimodal encoders, the framework significantly accelerates training convergence and achieves state-of-the-art performance on both vision and audio understanding benchmarks.

Addresses unbalanced learning in multimodal instruction tuningCoordinates LLM and feature encoder learning dynamicallyMitigates oscillation and biased learning for better convergence

Existing multimodal large language models (MLLMs) employ only unidirectional mapping of visual information into the language space, failing to effectively leverage visual knowledge to enhance holistic reasoning. This work introduces Vision-Augmented Large Language Models (VA-LLMs), breaking from this paradigm by enabling LLMs to actively store, share, and retrieve visual knowledge. Our contributions are threefold: (1) Modular Visual Memory (MVM), a structured, long-term storage mechanism for visual knowledge; (2) Soft Multimodal Mixture of Experts (MoME), a dynamic architecture that coordinates vision and language experts during autoregressive generation; and (3) a cross-modal knowledge injection and collaborative reasoning framework. Experiments demonstrate substantial improvements in physical commonsense understanding and spatial reasoning, achieving state-of-the-art performance across multiple multimodal benchmarks, including MMMU, ScienceQA, and POPE.

Enhances LLMs with multimodal knowledge storage and sharingIntroduces Modular Visual Memory for open-world visual informationUses Mixture of Multimodal Experts to improve reasoning capabilities

Latest Papers

What's happening recently
View more

This work addresses the deployment challenges of existing parameter-efficient fine-tuning methods—such as LoRA and Soft Prompting—which require modifications to the model’s computation graph and are thus incompatible with high-throughput inference engines like vLLM. The authors propose ART, a novel approach that treats visual inputs as learnable “computational art.” By optimizing only pixel-level visual prompts while keeping all parameters of the multimodal large language model frozen, ART injects task-specific information without altering the model architecture or computation graph. Consequently, it seamlessly integrates with precompiled inference systems and supports arbitrary fine-tuning objectives. Experimental results demonstrate that ART achieves performance on par with LoRA on Qwen-series models across multiple textual benchmarks, particularly excelling in mathematical reasoning and structured tool-use tasks.

Computational Graph ModificationHigh-throughput InferenceMulti-modal LLMs

This work addresses the limitations of current vision–language foundation models and evaluation benchmarks, which are predominantly English-centric and computationally expensive, thereby hindering their applicability to low-resource languages in multimodal settings. To overcome these challenges, we propose a lightweight text–speech–vision tri-modal fusion framework that leverages a stack of efficient adapters for cross-modal alignment, coupled with cost-effective data construction and fine-tuning strategies to enable collaborative multimodal modeling under constrained computational resources. Our contributions include an open-source compact multilingual multimodal model, an end-to-end speech–text–LLM pipeline, a culturally aware evaluation benchmark, and a reproducible toolchain, collectively yielding significant improvements in multimodal understanding and generation for non-English languages.

compute-efficient AIlow-resource languagesmultilingual multimodal LLMs

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

Dec 28, 2025
MZ
Mingyuan Zhang
🏛️ Northeastern University

To address the challenge of adapting vision-language models (VLMs) downstream while keeping pretrained weights frozen, this paper proposes Masked Fine-Tuning (MFT): a parameter-efficient paradigm that updates no model parameters. Instead, it introduces learnable gating scores into the language module and image–text projector to dynamically mask and reconfigure existing weight connections, thereby enabling structured subnetwork reconfiguration. This is the first work to apply masked fine-tuning to VLMs, demonstrating that merely “rewiring” frozen backbone connections—without modifying the visual encoder—suffices to unlock strong task-specific adaptability. MFT is compatible with multilingual backbones and consistently outperforms LoRA variants and full fine-tuning across multiple VLM benchmarks, while preserving the visual encoder entirely frozen. The method achieves superior performance with zero trainable parameters in the vision tower, offering a novel perspective on structural adaptation of frozen multimodal foundations. Code is publicly available.

Achieving high performance without altering frozen backboneReorganizing internal subnetworks for task adaptationUnlocking hidden capabilities in vision-language models

Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes

Oct 03, 2025
NE
Nirmal Elamon
🏛️ Artificial Creative intelligence | Expedia Group

This work addresses two key challenges in few-shot visual detection (<1,000 images): the strong data dependency of CNNs and the low fine-tuning efficiency of multimodal large language models (MLLMs). We propose a specialized fine-tuning paradigm for text-annotated detection tasks, explicitly modeling the synergy between language guidance and visual localization. Our approach integrates prompt engineering, dynamic context reasoning, and lightweight end-to-end fine-tuning. Under extreme data scarcity, the method achieves a 36% mAP improvement over conventional CNN baselines—matching or even surpassing their fully supervised performance. To our knowledge, this is the first work to empirically demonstrate the strong generalization capability of MLLMs in domain-specific few-shot visual detection. It establishes a novel cross-modal few-shot learning paradigm grounded in task-aware textual supervision. The implementation is publicly available.

Bridging vision and language through efficient cross-modal learning strategiesFine-tuning multi-modal LLMs for object detection with limited dataImproving performance on specialized visual tasks using language-guided models

Hot Scholars

ZY

Zhiyuan You

MMLab, The Chinese University of Hong Kong
Deep LearningComputer VisionLow-level Vision
FS

Fahad Shahbaz Khan

MBZUAI, Linköping University Sweden
Computer VisionObject RecognitionGenerative AIAI for Science
YZ

Yuhan Zhu

Nanjing University, Shanghai AI Lab
Computer VisionVision-Language ModelsVideo Understanding
JZ

Jun Zhang

Master Student, Nanjing University
Computer VisionVision and Language
XZ

Xiangyu Zeng

Nanjing University; Shanghai AI Laboratory
Computer VisionMLLM