frontier models
The Image Model Learned to Critique Its Own Revisions
One trajectory-level advantage updated both reflection tokens and image edits across repeated repair rounds.
Summary
One trajectory-level advantage updated both reflection tokens and image edits across repeated repair rounds.
UMM-Reflection trains a unified multimodal model across complete diagnose-and-revise trajectories rather than optimizing text reflection or image generation alone. Sibling attempts share an initial image, and one group-relative advantage updates both the reflection tokens and flow-based revisions. On BAGEL, the authors report a 12.05-point GenEval improvement over supervised fine-tuning, with gains on three held-out evaluation suites. The benchmark gains do not establish that every self-critique is accurate or that the method transfers unchanged to other architectures.
Why it matters
One trajectory-level advantage updated both reflection tokens and image edits across repeated repair rounds.
Limits and context
- The benchmark gains do not establish that every self-critique is accurate or that the method transfers unchanged to other architectures.
Key claims
One trajectory-level advantage updated both reflection tokens and image edits across repeated repair rounds.
Qualification: The benchmark gains do not establish that every self-critique is accurate or that the method transfers unchanged to other architectures.
Evidence: source-2026-09-29-005
Sources
- arXiv preprint 2609.35767arXiv · primary research
Corrections
No corrections have been recorded for this story.