🤖 AI Summary
This study addresses the issue of high error-label rates caused by majority voting and judge mechanisms in self-evolving models. To this end, it proposes the VQS framework, which introduces a programmatic verification mechanism to replace conventional paradigms. Specifically, images are parsed into structured records from which fixed programs generate question-answer pairs, requiring the model only to verify fine-grained factual claims. This approach ensures data quality and drives continuous model optimization under unsupervised conditions. Experimental results demonstrate that human-evaluated accuracy improves from 76% to 94%, with performance gains of up to 3.84 points across multiple scales of Qwen3-VL, significantly outperforming existing baselines.
📝 Abstract
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS