robotics
The Robot's Rich Vision Still Forgot What Control Needed
GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.
Summary
GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.
The authors call the mismatch between visually rich representations and control-useful structure the action-sufficiency gap. Their training constraints align geometry, predict instruction-relevant affordances and reconstruct goal regions inside a VLA policy and two world-action models. On zero-shot LIBERO-Plus transfer, the three variants gained 4.6, 12.6 and 5.2 points over matched counterparts; on RoboCasa the reported gains were 12.6, 9.0 and 8.4 points. Those results support the tested training principle, not a general guarantee for unseen robots.
Why it matters
GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.
Limits and context
- Those results support the tested training principle, not a general guarantee for unseen robots.
Key claims
GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.
Qualification: Those results support the tested training principle, not a general guarantee for unseen robots.
Evidence: source-2026-09-04-004
Sources
- arXiv preprint 2609.04193arXiv · primary research
Corrections
No corrections have been recorded for this story.