Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address pervasive visual hallucinations—textual responses inconsistent with image content—in multimodal large language models (MLLMs), this paper proposes Semantic Curriculum Preference Optimization (SCPO). SCPO unifies fine-grained semantic contrast, bidirectional symmetric preference learning, and progressive curriculum learning within a multimodal alignment framework. It introduces a dynamic reference model and a semantic difficulty scheduler that progresses from easy to hard instances, mitigating myopic training behavior. Evaluated on multiple LLaVA variants, SCPO reduces visual hallucination rates by up to 62.9%. Crucially, it maintains or even improves factual accuracy and overall performance on standard multimodal benchmarks—including MMBench and OCRBench—demonstrating robust generalization without compromising capability.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evidence. While Direct Preference Optimization(DPO) is widely used for alignment, its application to MLLMs often fails to capture fine-grained semantic differences and encourages shortcut learning. To address these challenges, we propose Semantic Curriculum Preference Optimization (SCPO), a novel framework for MLLM alignment. SCPO employs a progressive, easy-to-hard curriculum built upon our Semantic Curriculum Preference Pairs dataset, which provides fine-grained semantic contrasts sorted by difficulty. This curriculum is trained with a dynamic reference model and a novel symmetric, bidirectional objective to facilitate simultaneous learning from both textual and visual preferences. To our knowledge, SCPO is the first framework to unify semantics, symmetry, and curriculum for MLLMs alignment, effectively mitigating visual hallucinations. Extensive experiments on LLaVA models across various scales and versions validate that SCPO demonstrates superior performance compared to baseline models on multiple hallucination benchmarks, reducing the hallucination rate by up to 62.9%. Moreover, evaluations on generalized benchmarks show that SCPO improves factuality while preserving general capabilities, with its performance remaining stable across general vision-language benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Mitigating visual hallucinations in Multimodal Large Language Models
Addressing fine-grained semantic differences in model alignment
Preventing shortcut learning while improving response factuality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Curriculum Preference Optimization for MLLM alignment
Dynamic reference model with bidirectional learning objectives
Fine-grained semantic contrasts sorted by difficulty curriculum
Y
Yuanshuai Li
School of Engineering, Westlake University, Hangzhou, China
Y
Yuping Yan
School of Engineering, Westlake University, Hangzhou, China
J
Junfeng Tang
School of Engineering, Westlake University, Hangzhou, China
Yunxuan Li
Yunxuan Li
Google, California Institute of Technology
PhysicsArtificial IntelligenceNatural Language Processing
Z
Zeqi Zheng
School of Engineering, Westlake University, Hangzhou, China
Y
Yaochu Jin
School of Engineering, Westlake University, Hangzhou, China