Unified multimodal models · Reinforcement learning · Test-time reflection

UMM-Reflection

Learning native reflection in unified models with interleaved reinforcement learning

Yijia Fan1,*,†, Ziqi Huang1,*, Zhongang Cai1, Yan Li2, Zimo Wen2, Wanqi Yin3, Haiwen Diao1, Ziwei Liu1,‡

1Nanyang Technological University2Shanghai Jiao Tong University3The University of Tokyo

*Equal contribution   †Work done during an internship at NTU   ‡Corresponding author

Nanyang Technological University Shanghai Jiao Tong University The University of Tokyo
Stage 1, data construction: a prompt pool feeds an image generator; a VLM critic either accepts the image or writes an edit, an image editor applies it, and accepted trajectories are kept as one-shot, natural repair or planned progression.
Stage 2, supervised initialization: interleaved text and image trajectories train unified BAGEL with text cross-entropy and image flow matching, giving the reflection policy pi SFT.
Stage 3, reinforcement learning: from a shared first image, 16 complete trajectories are rolled out, scored by a frozen verifier into trajectory rewards, and a group-relative advantage updates text and flow jointly.
From borrowed critique to native reflection. External models are used only to build the SFT data. After RL, a single model inspects and revises its own images, with no critic or verifier at inference.
+12.05 GenEval
0.72 → 0.84 over reflection SFT; the RL training domain
+10.97 WISE
0.63 → 0.74 on world-knowledge prompts, held out from RL
+3.48 OneIG
0.79 → 0.83 on OneIG-Bench, held out from RL
+4.63 T2I-CompBench++
0.50 → 0.55 on T2I-CompBench++, held out from RL
Overview

Abstract

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped.

We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

Video

A three-minute tour of UMM-Reflection

Key finding

Reflection is only useful when it changes the image

Writing a plausible critique is easy. Turning that critique into a better image is the hard part. We compare three ways to give BAGEL a second look at its own output, each with up to three edits, on GenEval and on three benchmarks never used in RL. Numbers below each score are the change from BAGEL-Base.

External critic

BAGEL + GPT-5.5 critic

A frontier model reads each image and writes the edit instruction. BAGEL performs the edit.

GenEval0.79+0.08
WISE*0.70+0.12
OneIG0.83+0.02
T2I-CompBench++0.52+0.03
Wrong GenEval images repaired27.6%

The critic declared 67 of the 163 wrong first images finished at round 0.

Imitation

BAGEL-SFT reflection

The model is fine-tuned on interleaved reflection trajectories and reflects on its own.

GenEval0.72+0.00
WISE0.63+0.08
OneIG0.79−0.01
T2I-CompBench++0.50+0.01
Wrong GenEval images repaired20.6%

It writes reflections in the right format, but the edits rarely fix what was wrong.

Learned from consequences

UMM-Reflection

Same model and protocol, trained with whole-trajectory RL on the images its reflections produce.

GenEval0.84+0.12
WISE0.74+0.19
OneIG0.83+0.02
T2I-CompBench++0.55+0.06
Wrong GenEval images repaired64.9%

Reflection and generation are trained together, so a diagnosis is rewarded only when the edit it leads to is better.

* The critic's WISE score comes from a separate GPT-4o scoring run; its change is measured against BAGEL-Base in that same run (0.58).

Same GenEval prompt, a dog right of a tie. SFT keeps moving the tie and the dog back and forth and ends where it started. After RL the revisions move steadily into the correct-image region and end with the tie left of the dog.
Same prompt, same first-round quality, different revisions. SFT re-rolls the same mistake across rounds. After RL each edit builds on the previous one and the trajectory ends in the correct-image region. The state-space panel is schematic.
The gain is in the revisionsFirst-round accuracy is 70–73 for base, SFT and RL models alike. With three reflection rounds, UMM-Reflection reaches 84, while base and SFT stay at 73 and 72.
It fixes what resampling cannotOn the 62 prompts where four independent BAGEL samples all fail, reflection after RL repairs 60% of wrong first images. SFT repairs 12%.
It transfersRL sees only GenEval-style prompts. The same checkpoint improves WISE by 0.19, T2I-CompBench++ by 0.06 and OneIG-Bench by 0.02 over the base model.
Method

Reinforcement learning on whole reflection trajectories

SFT teaches the format of reflection: inspect, diagnose, write an edit, render it. It cannot tell the model which reflections actually lead to better images, because that is only known after the edit is drawn. RL closes that loop. The model rolls out complete reflection trajectories, a verifier scores the images they produce, and the outcome is credited to both the words and the pixels that led to it.

UMM-Reflection RL. (1) Sixteen rollouts share one detached initial image for the prompt a dog right of a tie. (2) Each rollout interleaves the model's own reflection text with its renders for three rounds. (3) A frozen verifier scores every image; trajectory rewards are 1.21, 1.19 and 0.30 for the three rollouts shown. (4) Group normalization gives one advantage per trajectory, (5) which updates both the text head and the flow head of the unified backbone.
One RL step on one prompt. Sixteen rollouts share one first image; each interleaves the model's own reflection (verbatim) with its renders for up to three rounds. A frozen verifier scores every image (q) and every trajectory (R(τ)). Group normalization gives one advantage per trajectory, which updates both the text and flow heads. Green and red frames mark verifier pass and fail; dashed arrows are used in training only.

Image scores along each trajectory

Verifier scores q for the first image and the three revisions of the rollouts in the figure. Green bars pass the verifier.

How the trajectory reward is assembled

R(τ) = qT + α Σ[Δt]+ + β Smulti − λ Σ[−Δt]+ − p·1[premature DONE], with α = β = 0.3 and λ = p = 0.5.

One group of sixteen rollouts, ranked by reward

Illustrative group

All sixteen rollouts start from the same first image (q = 0.30). τ₁, τ₂ and τ₁₆ are the rollouts in the figure; the other thirteen are constructed to show each reward term at work: steady progress, a partial fix, damage after a pass, and stopping with DONE before the image passes. The advantage is Ai = clip((R − μ) / max(σ, 0.1), −1, 1). Select a row to see its reward breakdown above.

RankRolloutq per roundWhat happenedR(τ)AdvantageAi
4,096

Why one advantage per trajectory

A group-relative estimate for every round would need sibling groups at every round: branching 16 ways at each of three rounds is 16³ = 4,096 rollouts per prompt. The trajectory-level advantage keeps the 16-sample cost of GRPO, while the progress terms in R(τ) still reward each round that improves the image.

1 shared Ai

Why text and flow share it

The text policy and the flow renderer are not normalized separately. A reflection that leads to a better image raises the advantage of both the diagnostic tokens and the rendering steps that followed, so the model learns which reflections lead to which visual outcomes.

0 verifiers at test time

Nothing extra at inference

The verifier and the frozen SFT reference are used only to compute rewards and KL during training. At inference only the unified model runs: it looks at its own image, writes the reflection and renders the revision.

Benchmarks

One checkpoint, four benchmarks, three of them never seen in RL

All rows use the same prompts, resolution and scorer, with up to three reflection rounds for the models that reflect. RL trains on GenEval-style prompts only; WISE, OneIG-Bench and T2I-CompBench++ measure transfer.

All five systems at a glance

Hover or tap an axis to compare on that benchmark.

GenEval accuracy by reflection budget

Official 553 prompts and detector, 512 px. R0 is the first image; +k allows k edits.

Repair rate during RL training

Share of wrong first images that are correct after three edits, along one RL run.

ModelGenEvalWISEOneIGT2I-CompBench++
BAGEL-Base0.710.550.800.49
BAGEL-SFT0.72+0.000.63+0.080.79−0.010.50+0.01
BAGEL-Self-Agentic0.77+0.050.61+0.060.81+0.000.52+0.03
BAGEL-T2I-RL (1k updates)0.76+0.040.54−0.010.80−0.010.49+0.00
UMM-Reflection0.84+0.120.74+0.190.83+0.020.55+0.06

Differences are relative to BAGEL-Base. Self-Agentic runs the untuned base model through the same three inspect-and-edit rounds. T2I-RL applies 1,000 RL updates to single-shot generation, with no reflection.

Analysis

Where the gain comes from

The first image barely changes. What RL changes is what happens after it: which errors get repaired, on which kinds of prompts, and which half of the model has to learn for that to work.

Sub-scores on every benchmark

GenEval families: first image vs. after reflection Official 553 prompts. Hollow dot: UMM-Reflection's first image; filled dot: its image after three reflection rounds; grey ring: reflection SFT after three rounds.

Paired vs SFT109: 38prompts won vs lost against reflection SFT on GenEval (p < 10⁻⁸, McNemar)
Position+42ptsspatial-relation accuracy over SFT, 0.47 → 0.89; the largest family gain
Hard prompts60%vs 12%repaired on the 62 prompts where four independent base samples all fail
Same image budget84vs 80reflection vs best-of-4 sampling from the stronger T2I-RL renderer (p ≤ 0.04)

Both heads must be trained

Trajectory RL with one head frozen. GenEval final accuracy and repair rate of wrong first images.

What the numbers say

A better renderer alone does not repairTraining only the flow head leaves the model close to SFT (73, repair 22.8%): the renderer improves, but the reflections that drive it do not.
Learning what to write is the larger partTraining only the text head recovers most of the gain (78, 49.4%).
Joint training adds six more pointsWith both heads, the renderer also learns to execute the reflections the text head now writes: 84 and 64.9%. This joint optimization is what a single unified model makes possible.
Qualitative examples

Repairs, round by round

Selected prompts from all four benchmarks on which UMM-Reflection's first image fails and its final image passes. Each reflection is the model's own output, verbatim; the badge under each image is that benchmark's own verdict. Select an image to enlarge it.

Qualitative comparison

Same prompt, every variant

GenEval prompts on which UMM-Reflection's final image is the only one that passes. Each column is one model's final image; the number under each name is its GenEval score. The inset in the last column is UMM-Reflection's own first image.

Badges are the official GenEval detector verdicts. BAGEL-Base generates one image; every other column is the image after three rounds.

Citation

BibTeX

@article{ummreflection2026,
  title   = {Learning Native Reflection in Unified Models
             with Interleaved Reinforcement Learning},
  author  = {Fan, Yijia and Huang, Ziqi and Cai, Zhongang and Li, Yan and
             Wen, Zimo and Yin, Wanqi and Diao, Haiwen and Liu, Ziwei},
  journal = {arXiv preprint},
  year    = {2026}
}