Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of vision-language models to hallucinations regarding character identities and actions in video captioning, particularly during camera transitions. To this end, it introduces FlexBench, a benchmark accompanied by a Graded Physical Alignment (GPA) scoring system designed to precisely distinguish between information omissions and factual errors. Furthermore, the GM-DPO algorithm is proposed, which incorporates weighted preference margins to impose higher penalties on severe action hallucinations and integrates automated evaluation checklists for fine-grained training. Experimental results demonstrate that GM-DPO improves GPA scores by 2.02–3.40 points over standard DPO and reduces the weighted hallucination rate of the Qwen3-8B model by 21.3%, while effectively preserving its long-text generation capabilities.
📝 Abstract
Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Limb-Motion Captioning
Hallucination
Video Captioning
Preference Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graded Preference Optimization
GM-DPO
Limb-Motion Captioning
Hallucination Mitigation
Vision-Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yanan Wang
1Zhejiang University, 2Alibaba Token Hub, Alibaba Group
T
Tingsong Li
2Alibaba Token Hub, Alibaba Group, 3University of Science and Technology of China
Kaixun Jiang
Kaixun Jiang
Fudan University
Computer VisionAdversarial Examples
C
Chongyang Zhong
2Alibaba Token Hub, Alibaba Group
C
Chenwei Xie
2Alibaba Token Hub, Alibaba Group
Z
Zhaohe Liao
2Alibaba Token Hub, Alibaba Group, 5Shanghai Jiao Tong University