Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

2. Document Verification & Archival Data
- Contributing Researchers: Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu
- Submission Date: September 28, 2026
- Full Preprint Document: Download Official PDF
- Permanent Archive Record: arXiv:2609.35767v1
3. Academic Citation Reference
Standard Reference (APA Format):
Yijia Fan, et al. (2026). Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning. arXiv:2609.35767v1. https://arxiv.org/abs/2609.35767v1
Academic Field: Computer Vision | Document Identifier: arXiv:2609.35767v1
BibTeX Entry:
@article{arxiv_2609.35767v1,
author = {Yijia Fan and Ziqi Huang and Zhongang Cai and Yan Li and Zimo Wen and Wanqi Yin and Haiwen Diao and Ziwei Liu},
title = {{Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning}},
journal = {arXiv preprint arXiv:2609.35767v1},
year = {2026},
url = {https://arxiv.org/abs/2609.35767v1}
}
Henüz yorum yapılmamış. İlk yorumu siz yapın!