robotics
The Depth Stayed Sound While the Camera Pose Drifted Away
Scal3R froze the geometry backbone, added multi-reference pose tokens and cut average trajectory error by more than 60 percent on KITTI.
Summary
Scal3R froze the geometry backbone, added multi-reference pose tokens and cut average trajectory error by more than 60 percent on KITTI.
Long online reconstructions often collapse because every pose is extrapolated from the first frame even when per-frame depth remains stable. Scal3R adds lightweight tokens—about one percent of the model's parameters—to query pose against multiple past keyframes, then closes loops with online pose-graph optimization. The authors report convergence in eight hours on one GPU and state-of-the-art results across six evaluated datasets. The evidence concerns benchmark reconstruction, not safety certification for deployed navigation.
Why it matters
Scal3R froze the geometry backbone, added multi-reference pose tokens and cut average trajectory error by more than 60 percent on KITTI.
Limits and context
- The evidence concerns benchmark reconstruction, not safety certification for deployed navigation.
Key claims
Scal3R froze the geometry backbone, added multi-reference pose tokens and cut average trajectory error by more than 60 percent on KITTI.
Qualification: The evidence concerns benchmark reconstruction, not safety certification for deployed navigation.
Evidence: source-2026-09-06-004
Sources
- arXiv preprint 2609.04201arXiv · primary research
Corrections
No corrections have been recorded for this story.