TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

A Video Model Learned Depth as the Next Frame

GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.

Published Updated Story ID: mp-2026-08-31-009
Read the complete editionStory JSON

Summary

GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.

The method adapts a video generative model to jointly represent images and geometry targets rather than training separate task-specific diffusion systems. Its authors report stronger zero-shot monocular depth and normal estimation than prior generative competitors with substantially less training data, and performance near discriminative systems trained on more than 100 times as much data. Those comparisons remain benchmark results from the proposing team.

Why it matters

GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.

Limits and context

  • Those comparisons remain benchmark results from the proposing team.

Key claims

  1. GeoNeXt reframes depth and surface-normal estimation as next-frame prediction inside a pretrained video generator.

    Qualification: Those comparisons remain benchmark results from the proposing team.

    Evidence: source-2026-08-31-009

Sources

  1. arXiv preprint 2608.28549arXiv · primary research

Corrections

No corrections have been recorded for this story.