benchmarks evals
Image Quality Rose While the Subject's Identity Drifted
A benchmark separates fidelity from polish across generation, editing, restoration and multi-subject scenes, then tests identity as persistent knowledge.
Summary
A benchmark separates fidelity from polish across generation, editing, restoration and multi-subject scenes, then tests identity as persistent knowledge.
The study compares identity supplied in prompt context, encoded in subject-specific parameters and maintained through a persistent identity layer. Drift worsened during repeated edits, at small subject scales, under severe restoration and when several subjects shared a scene. The persistent representation improved identity scores across tested foundation models while preserving comparable instruction adherence and perceptual quality. The benchmark is produced alongside one of the compared approaches, so the result should be read as reported evaluation rather than neutral product certification.
Why it matters
A benchmark separates fidelity from polish across generation, editing, restoration and multi-subject scenes, then tests identity as persistent knowledge.
Limits and context
No additional limitation was separately recorded.
Key claims
A benchmark separates fidelity from polish across generation, editing, restoration and multi-subject scenes, then tests identity as persistent knowledge.
Evidence: source-2026-09-06-015
Sources
- arXiv preprint 2609.04151arXiv · primary research
Corrections
No corrections have been recorded for this story.