Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning - Technical Figure
Figure 1: Architectural diagram and empirical setup from the original research paper (arXiv:2609.35767v1).

2. Document Verification & Archival Data

  • Contributing Researchers: Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu
  • Submission Date: September 28, 2026
  • Full Preprint Document: Download Official PDF
  • Permanent Archive Record: arXiv:2609.35767v1

3. Academic Citation Reference

Standard Reference (APA Format):

Yijia Fan, et al. (2026). Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning. arXiv:2609.35767v1. https://arxiv.org/abs/2609.35767v1

Academic Field: Computer Vision | Document Identifier: arXiv:2609.35767v1


BibTeX Entry:

Kod
@article{arxiv_2609.35767v1,
  author    = {Yijia Fan and Ziqi Huang and Zhongang Cai and Yan Li and Zimo Wen and Wanqi Yin and Haiwen Diao and Ziwei Liu},
  title     = {{Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning}},
  journal   = {arXiv preprint arXiv:2609.35767v1},
  year      = {2026},
  url       = {https://arxiv.org/abs/2609.35767v1}
}

Yorumlar (0)

Henüz yorum yapılmamış. İlk yorumu siz yapın!

Yorum Bırakın