CLGRPO: Reasoning Ability Enhancement for Small VLMs

📅 2025-06-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the weak reasoning capability of small vision-language models (SVLMs, ≤2B parameters) stemming from limited model capacity, this paper proposes a four-stage progressive training framework: (1) self-supervised generation of chain-of-thought (CoT) data, followed by (2) supervised fine-tuning (SFT), (3) grouped relative policy optimization (GRPO), and (4) ClipLow gradient clipping–enhanced training with dual constraints on output format and answer accuracy. Notably, ClipLow-GRPO is the first reinforcement learning method specifically designed for small models, effectively mitigating both capacity bottlenecks and challenges in modeling complex reasoning patterns. Evaluated on EMOSet-118K, our 1B-parameter SVLM achieves a 2.77-percentage-point gain in reasoning accuracy and a 0.69-percentage-point improvement in recall—approaching the performance of an 8B-parameter baseline. This advancement significantly enhances the practical utility of SVLMs.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Small Vision Language Models (SVLMs) generally refer to models with parameter sizes less than or equal to 2B. Their low cost and power consumption characteristics confer high commercial value. However, their reasoning abilities are limited by the number of parameters. To address this issue, this paper proposes a post-training optimization paradigm called the Incremental Training Strategy to enhance the reasoning ability of SVLMs. Firstly, we constructed a Self-Supervised Chain-of-Thought (COT) Data Construction System, which leverages multiple LVLMs with 7B parameters or more to transform original data into COT data in a self-supervised manner. Our proposed Incremental Training Strategy consists of four stages. Stage 1 injects domain knowledge by performing Supervised Fine-Tuning (SFT) to the pretrained model on the COT data. Stage 2 aligns the COT data format by conducting a small amount of Group Relative Policy Optimization (GRPO) training constrained only by format rewards on the COT data. Stage 3 enhances reasoning ability by applying GRPO training on the COT data with constraints on both format and accuracy rewards. The resulting model shows significant improvement compared to the baseline. Stage 4 addresses the limited capacity of the SVLMs and the weak ability to capture complex patterns by proposing ClipLow GRPO (CLGRPO) to constrain the capture space of the training process. We conducted extensive comparative and ablation experiments on the abstract semantic recognition dataset EMOSet-118K. Experimental results demonstrate that our method significantly improves the reasoning ability of 1B SVLM. Compared to the baseline model fine-tuned on the original data, accuracy increased by 2.77 and recall by 0.69, achieving performance comparable to that of 8B models.
Problem

Research questions and friction points this paper is trying to address.

Enhancing reasoning ability of small VLMs
Overcoming parameter limitations in SVLMs
Improving accuracy and recall in SVLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised COT Data Construction System
Four-stage Incremental Training Strategy
ClipLow GRPO for constrained capture space