Offline-Online Curriculum RL for Multimodal Reasoning

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of multimodal large language models to rely on non-causal shortcuts during reasoning, often producing incorrect intermediate steps while arriving at correct answers, thereby compromising interpretability and reliability. To mitigate this issue, the authors propose the O²-CritiCuRL framework, which introduces a novel critical-step-aware mechanism. In the offline phase, it distills essential reasoning steps through multi-trajectory replay analysis and step-level importance estimation. In the online phase, it employs curriculum reinforcement learning combined with truncated chain training to progressively guide the model in completing missing steps and refining its reasoning trajectory. By moving beyond static supervision, this approach effectively filters redundant steps and dynamically focuses on critical reasoning paths, achieving state-of-the-art performance on multimodal reasoning benchmarks while significantly enhancing both training and inference efficiency.
📝 Abstract
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, $O^2$-CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at https://github.com/kk0013/CritiCuRL.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
step-level supervision
critical steps
reliability
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Curriculum Reinforcement Learning
Multimodal Reasoning
Step-level Importance
Offline-Online Learning
Critical-step Awareness
🔎 Similar Papers
No similar papers found.
W
Wendi Deng
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
H
Hang Du
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
Guoshun Nan
Guoshun Nan
Professor of Beijing University of Posts and Telecommunications
Multimodal LearningVideo LLM6G SecuritySemantic Communications
H
Haokun Tian
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
J
Jiaqi Yu
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
X
Xinlei Cao
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
J
Jiale Li
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
J
Jingfeng Chen
Carnegie Mellon University, Pittsburgh, USA.
L
Ling Deng
Beijing University of Posts and Telecommunications, Beijing, 100876, China.
T
Ting Li
China Telecom Corporation Limited Sichuan Branch, Chengdu, Sichuan, China.
H
Hao Yang
Changsha University of Science & Technology, Changsha, Hunan, China.
Jun Liu
Jun Liu
Professor, Lancaster University
Computer VisionDigital HealthMachine LearningHuman Behavior Analysis
Xudong Jiang
Xudong Jiang
IEEE Fellow, Nanyang Technological University, Singapore
Pattern RecognitionComputer VisionMachine LearningImage ProcessingBiometrics
Sicong Leng
Sicong Leng
Nanyang Technological University
Multi-modal Learning