Unified multimodal models that can both perceive and render images present an intriguing possibility: they can repair their own generations without external intervention. In principle, such a model diagnoses what an image gets wrong, revises it, observes the result, and diagnoses again. Whether a revision helps becomes known only after it is rendered, which means the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning on reflection trajectories provides a cold start but does not discover the high-success repair paths. Naive reinforcement learning that optimizes only the renderer or only one head of the model leaves most of the potential gain untapped. This article introduces UMM-Reflection, which applies reinforcement learning to complete reflection trajectories inside a single unified model. Sibling trajectories that share one initial image enable a group-relative advantage that compares reflection strategies. A single trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines that rely on an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Why unified models still need explicit reflection training
Recent multimodal models can encode images and generate images from text prompts. They can also perform in-context editing tasks such as object removal or style transfer. But these capabilities are typically inherited from pre-training on image-caption pairs and instruction-following data. When a model generates an image that contains an error, a distorted object, an inconsistent spatial relationship, or a missing detail there is no built-in mechanism to detect the error propose a fix and verify the fix without additional engineering. Prior work has explored external verification, where a separate critic model evaluates generated revisions, and single-round editing, where the model makes one revision attempt based on a prompt. These approaches fail to scale because error detection and revision are fundamentally different operations and stitching them together requires hand-crafted pipelines that do not generalize across error types. The key challenge is that a model needs to learn two correlated skills simultaneously: producing a useful reflection that identifies what is wrong and executing a revision that corrects it. Supervised data can provide a starting point but reflection trajectories that lead to successful repairs are sparse and the space of possible reflection-revision pairs is combinatorially large.
The cold-start problem with supervised fine-tuning
Supervised fine-tuning (SFT) on reflection trajectories has been the most common starting point for training reflection capabilities. A dataset of image-revision-reflection triplets is collected, typically by pairing an generated image with a human-written or model-generated explanation of what was wrong and what change was made. SFT on this data gives the model a basic ability to produce reflection text and apply revisions. But the learned policy is limited to the specific reflection-revision pairs seen during fine-tuning. When the model encounters a new type of error not represented in the fine-tuning set, it either produces a generic reflection that does not pinpoint the issue, or it applies a revision that does not fix the underlying problem. The paper notes that SFT alone finds a useful initialization but does not locate the high-success repair paths that require multi-step reasoning. This limitation motivated the use of reinforcement learning to explore beyond the coverage of the supervised dataset.
Limitations of naive reinforcement learning
Naive reinforcement learning methods that optimize only the renderer or only one component of the unified model say only the reflection head or only the image generation head suffer from two related problems First the credit assignment problem when the final outcome such as a human rating of the repaired image is observed it is difficult to determine which intermediate action the reflection text the revision stroke or both contributed most to the success Second optimizing one head while keeping the other fixed creates an unbalanced policy If the reflection head learns to produce critiques that are easy for the current renderer to address the renderer never learn to handle harder cases Conversely if the renderer becomes very good at masking errors the reflection head may learn to produce critiques that are unnecessary failure modes result in most of the potential gain from reflection remaining untapped because the model does not learn to coordinate the two skills jointly over complete revision loops.
UMM-Reflection: interleaved RL across complete trajectories
UMM-Reflection addresses these limitations by applying reinforcement learning to entire reflection-revision loops within a single unified model. The model simultaneously learns to produce a reflection (a sequence of tokens describing what is wrong and how to fix it) and to render a revised image based on that reflection. The training framework operates on sibling trajectories: for a given initial generated image, multiple reflection-revision attempts are sampled, producing a set of sibling trajectories that share the same starting point. This grouping is essential for the advantage computation that follows.
The group-relative advantage compares reflection strategies by evaluating each trajectory's outcome relative to the average outcome of its siblings. If a particular reflection leads to a revision that scores higher than the average of sibling revisions, the reflection receives a positive advantage signal. This approach does not require an external verifier; the reward signal can be as simple as a human rating, a learned perception score, or a task-specific metric such as CLIP similarity to a target image. Because the comparison is relative to siblings that share the same initial error, the advantage isolates the effect of the reflection strategy from the effect of the initial image content.
Crucially, a single trajectory-level advantage updates both the reflection tokens and the flow-based revisions. In a flow-based revision system, the image is modified through a series of invertible operations, and the gradient of the reward can be backpropagated through the flow. This means the reflection tokens and the revision parameters receive gradients from the same reward signal, enabling joint optimization. The update rule effectively says: "adjust the reflection and the revision together so that the next loop produces a better outcome." By updating both roles of the same model, the method avoids the combinatorial blow-up that would occur if per-round credit assignment were attempted independently for each step in the loop.
Credit flow across rounds and model roles
In single-round editing pipelines or pipelines with an external critic, credit is typically assigned after a fixed number of rounds, and the reflection and revision modules are updated separately. UMM-Reflection differs in that credit flows across rounds and to both roles of the same model. When a revision in round t leads to a better outcome, the gradient flows back through the revision operation and through the reflection that prompted it. This gradient then influences the reflection parameters for round t+1 and also retroactively adjusts the reflection parameters for earlier rounds. The same mechanism applies to the revision (flow) parameters: if a particular sequence of flow operations consistently leads to higher rewards when prompted by effective reflections, those flow parameters are strengthened.
This cross-round credit flow means that the model learns which types of reflections are most useful for which types of revisions, and it learns which revision sequences are most compatible with particular reflection styles. Over many training trajectories, the model develops a coordinated repertoire of reflection-revision pairs that generalize to new error types. The paper emphasizes that no external verifier is needed at inference time: the model has learned to produce reflections that are self-consistent, and the revision operation is already embedded in the model's weights.
BAGEL benchmark results and transfer across composition sets
The primary experimental evaluation is on the BAGEL benchmark, which measures how well text-to-image models can follow editing instructions across multiple rounds. The benchmark provides a starting generated image and an editing prompt, and the model must produce a revised image that satisfies the prompt. Evaluation uses GenEval, a metric that combines automatic similarity measures with human judgment of edit fidelity.
On BAGEL, UMM-Reflection improves GenEval by 12.05 points over strong fine-tuning (SFT) baselines. This is a substantial gain, especially since the baseline already uses reflection-capable fine-tuning. The improvement indicates that the interleaved RL framework discovers reflection-revision strategies that were not captured by supervised data alone. The gains transfer to three additional compositional benchmarks without any further training. On WISE, a benchmark of visual editing tasks, UMM-Reflection achieves a +10.97 point improvement over SFT. On OneIG-Bench, which evaluates one-shot image generation and editing, the improvement is +3.48 points. On T2I-CompBench++, a comprehensive text-to-image composition benchmark, the improvement is +4.63 points. The fact that none of these benchmarks was used in training means the method learns reflection-revision skills that generalize across compositional settings.
Audio and qualitative analyses
The paper also reports qualitative results and user studies. Human evaluators were shown pairs of revisions one from the SFT baseline and one from UMM-Reflection and asked to choose the one that better preserves the original image content while satisfying the editing prompt. UMM-Reflection revisions were preferred in 58% of comparisons with evaluators noting that the reflections produced by UMM-Reflection were more specific about what was wrong for example the object is tilted rather than the image looks off and the revisions were more precise for example applying a small rotation rather than regenerating the entire object.
An ablation study reported in the appendix isolates the contribution of the group-relative advantage versus a flat baseline where all trajectories receive the same advantage, and the contribution of the trajectory-level update versus updating only the reflection head or only the revision head. The full UMM-Reflection objective outperforms all ablations, confirming that both the group-relative comparison and the joint update are necessary for the observed gains.
Implementation details and compute
UMM-Reflection is implemented on top of a frozen base text-to-image model. The model's existing cross-attention and U-Net components serve as the image revision mechanism, and the language modeling head is used to produce reflection tokens. No additional modules are introduced. The reinforcement learning objective is added on top of the existing supervised fine-tuning, so the training starts from a reflection-capable initialization and then extends it with RL. The reward signal is computed at the end of each reflection-revision loop and is normalized using the group-relative advantage formula. Training uses Proximal Policy Optimization (PPO) with clipping, but the clipping range is applied to the group-relative advantage rather than to per-token importance weights. The number of sibling trajectories sampled per initial image is four, and the reflection length is limited to 32 tokens to keep the optimization tractable. The total training compute is roughly 0.5% of the compute used to train the base text-to-image model, making the incremental cost modest.
Why the one-model approach matters
The decision to keep everything inside a single unified model, rather than routing reflections to a separate critic or using an external verification pipeline, has several practical implications. First, the model can be deployed like any other text-to-image model: no additional network, no extra inference pass, and no change to the serving infrastructure. Second, because the same weights produce both reflections and revisions, there is no mismatch between the reflection language and the revision operations both are expressed in the model's native parameter space. Third, the absence of an external verifier means that the reflection-revision loop can run at inference time if desired, enabling interactive editing where a user can prompt the model to "try again" based on its own critique. Fourth, the method inherits the base model's capabilities in rendering, lighting, and object geometry, so the revisions are not limited to simple pixel-level operations but can involve coherent changes in pose, lighting, and composition.
Conclusion
UMM-Reflection demonstrates that reinforcement learning applied to complete reflection-revision loops inside a unified multimodal model can significantly improve the model's ability to self-correct. The key technical innovations are the group-relative advantage, which compares sibling trajectories that share an initial image, and the trajectory-level update, which adjusts both reflection tokens and flow-based revisions simultaneously. The result is a model that learns to produce more specific and useful reflections and to apply more precise revisions, all without needing an external verifier. The gains transfer across compositional benchmarks, suggesting that the method learns a general skill for image editing rather than overfitting to a particular task distribution. For developers working on multimodal model training, UMM-Reflection provides a minimal-code addition to existing fine-tuning pipelines and a path toward fully autonomous image editing capabilities.
The method's design choices warrant further discussion. The group-relative advantage depends on having sufficiently diverse sibling trajectories; if all sampled revisions produce similar outcomes, the advantage signal becomes uninformative. The paper ablates this by reporting results with two sibling counts, two and four, and shows that four siblings produce more stable advantage estimates without a proportional increase in training cost. The trajectory-level update that simultaneously adjusts reflection tokens and flow-based revisions is the other critical design decision. An ablation that updates only the reflection head while keeping the revision fixed, or vice versa, achieves smaller gains 6.2 points on GenEval versus the full 12.05 confirming that coordination between the two roles is essential. The ablation also shows that using a separate critic for the reward signal halves the GenEval improvement to 5.8 points, further confirming the benefit of native reflection within the unified model.
Read the paper on arXiv