research
A Video Model Learned Depth as the Next Frame
GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.
Summary
GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.
The method adapts a video generative model to jointly represent images and geometry targets rather than training separate task-specific diffusion systems. Its authors report stronger zero-shot monocular depth and normal estimation than prior generative competitors with substantially less training data, and performance near discriminative systems trained on more than 100 times as much data. Those comparisons remain benchmark results from the proposing team.
Why it matters
GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.
Limits and context
- Those comparisons remain benchmark results from the proposing team.
Key claims
GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.
Qualification: Those comparisons remain benchmark results from the proposing team.
Evidence: source-2026-08-31-009
Sources
- arXiv preprint 2608.28549arXiv · primary research
Corrections
No corrections have been recorded for this story.