Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

📅 2025-12-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Visual language models (VLMs) exhibit severe deficiencies in generative capabilities for low-resource languages such as Romanian. Method: We introduce RoVQA—the first Romanian-language Flickr30k visual question answering dataset—and propose a LoRA-based efficient fine-tuning framework for multimodal large language models, enabling the first joint modeling of Romanian visual QA and image captioning. Contribution/Results: Our approach is validated on Qwen2-VL, LLaVA-1.6, and LLaMA-3.2. The Qwen2-VL-RoVQA variant achieves +6.05% and +2.61% absolute gains in BERTScore F1 on VQA and image captioning, respectively, alongside significant reductions in grammatical errors. Notably, the model demonstrates strong cross-task generalization. This work establishes the first reproducible benchmark—comprising data, methodology, and baselines—for multimodal generative AI in low-resource languages.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30k dataset into Romanian and further extend it for visual question answering by leveraging open-source LLMs. We demonstrate the usefulness of our datasets by fine-tuning open-source VLMs on Romanian visual question answering. We select VLMs from three widely used model families: LLaMA 3.2, LLaVA 1.6, and Qwen2. For fine-tuning, we employ the parameter-efficient LoRA method. Our models show improved Romanian capabilities in visual QA, as well as on tasks they were not trained on, such as Romanian image description generation. The seven-billion-parameter Qwen2-VL-RoVQA obtains top scores on both tasks, with improvements of +6.05% and +2.61% in BERTScore F1 over its original version. Finally, the models show substantial reductions in grammatical errors compared to their original forms, indicating improvements not only in language understanding but also in Romanian fluency.
Problem

Research questions and friction points this paper is trying to address.

Develops Romanian vision-language models via multimodal instruction tuning
Translates and extends Flickr30k for Romanian visual question answering
Enhances Romanian fluency and task performance using parameter-efficient fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Translate Flickr30k dataset into Romanian for VQA
Fine-tune open-source VLMs using LoRA for efficiency
Extend model capabilities to untrained tasks like image description
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.