media creative tools
One Video Editor Learned Six Kinds of Change Without Training
EditVid combines sparse causal memory, token correspondence and latent blending for instruction- and reference-guided edits.
Summary
EditVid combines sparse causal memory, token correspondence and latent blending for instruction- and reference-guided edits.
The framework supports style transfer, attribute changes, object insertion, part edits and subject replacement without task-specific training. On FiVE, the authors report 78.16 FiVE-Acc versus 58.95 for the strongest evaluated training-free baseline, with competitive IVEBench results. A user study preferred EditVid overall in 51.8 percent of comparisons against seven methods. Those numbers reflect the chosen benchmarks and comparisons, not a blanket claim of identity-safe or artifact-free editing.
Why it matters
EditVid combines sparse causal memory, token correspondence and latent blending for instruction- and reference-guided edits.
Limits and context
- Those numbers reflect the chosen benchmarks and comparisons, not a blanket claim of identity-safe or artifact-free editing.
Key claims
EditVid combines sparse causal memory, token correspondence and latent blending for instruction- and reference-guided edits.
Qualification: Those numbers reflect the chosen benchmarks and comparisons, not a blanket claim of identity-safe or artifact-free editing.
Evidence: source-2026-09-06-007
Sources
- arXiv preprint 2609.04190arXiv · primary research
Corrections
No corrections have been recorded for this story.