arXiv · 2610.03002
Recursive Self-Improvement in Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huijuan Wang, Chufan Shi, Cheng Yang, Yaokang Wu, Taylor Berg-Kirkpatrick, Xuezhe Ma. 2026-10-02. Recursive Self-Improvement in Unified Multimodal Models. https://arxiv.org/abs/2610.03002
Cite the original work for its findings. Save a collection to share your selection of sources.