arXiv Science⌕ Search

arXiv · 2609.38721

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Abstract

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi. 2026-09-30. UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement. https://arxiv.org/abs/2609.38721

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks

Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or exceed solver-reference scores on several families, while weaker systems often fail to produce high-value valid plans and even strong systems lose quality on geometric, product-level, or design-heavy tasks. Trace analyses separate two failure points: task-contract misformulation and weak solution construction. Successful runs instead calibrate agent-written implementations against verifier feedback and adapt search to case-specific structure. Ablations show that procedure injection and memory accumulation help selectively, when they supply the missing formulation, calibration, or search support.

cs.AI↗

LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

Computer-use agents (CUAs) execute multi-stage tasks through tools, where an early unsafe decision can propagate to consequential actions. Evaluating only final outcomes can miss such decisions, while constructing executable environments for new tasks can make benchmark expansion costly. We present LPS-Bench, a benchmark of long-horizon planning safety in MCP-style tool workflows under benign requests and adversarial steering. A template-guided multi-agent pipeline generates user instructions, simulated toolkits, and case-specific safety criteria, followed by human review. This design supports scalable case expansion without building a separate application environment for every test case. LPS-Bench comprises 570 cases derived from 65 scenarios across 7 task domains and 9 planning-risk types, with representative cases additionally adapted to reusable skills. An LLM-based evaluator applies case-specific criteria to complete interaction records, examining tool choices, arguments, and responses to environmental feedback throughout execution. Evaluations of 13 LLM agents reveal persistent failures in both benign and adversarial settings. Prompt-based interventions yield model-dependent gains, but substantial safety failures remain.

cs.AI↗

On the Tip of the Tongue: Why LLMs Hallucinate Answers They Can Decode

A language model can give the wrong answer even when the correct answer is decodable from its intermediate states. To study this gap between decodability and selection, we distinguish \textit{read} from \textit{write} at the first answer token. Read asks whether the gold token can be decoded from intermediate residual states under same-relation decoy controls. Write asks whether the final readout ranks that token first among content tokens. Under three different readers, with a randomized-label control, a substantial fraction of failures remain readable while another content token is selected. We explain this through the selection margin at the final readout, the difference between the answer logit and the logit of its strongest alternative, which is answer support minus alternative support, and can also be split into a context-averaged baseline linked to token frequency and an item-specific term. Setting the answer support to the level typical of successful generations is sufficient to recover first-token selection for the majority of failures in most of the models we study; the original alternative remains ahead in most remaining failures under this edit, and this outcome follows directly from the readout geometry. Removing the frequency direction alone shifts selection but rarely recovers the answer. Prompt variants of the same fact that succeed supply support that transfers to failing variants through the residual stream and through late MLP outputs, with less consistent effects through late attention. First-token recovery leaves most full answers wrong, which limits the recovery achieved by these edits and separates three things that are easily conflated, decodability, recoverability, and generation.

cs.AI↗