🤖 AI Summary
This study addresses the bottlenecks of homogeneous supervision signals and cumulative error propagation across multiple iterations in the self-improvement of unified multimodal models. To overcome these limitations, this work proposes a recursive cross-capability self-improvement framework wherein textual and visual capabilities mutually serve as data sources for cross-generative training. Furthermore, external program execution is introduced as an objective ground-truth verification mechanism to transcend unidirectional supervision constraints and effectively interrupt error propagation. Empirical evaluations demonstrate that this approach elevates accuracy on chart-related tasks from 45.7% to 60.2%, achieving a verification pass rate of 95.2%. These results significantly outperform conventional continual training paradigms, offering a novel pathway for the autonomous evolution of multimodal models.
📝 Abstract
Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.