TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

The Robot's Rich Vision Still Forgot What Control Needed

GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.

Published Updated Story ID: mp-2026-09-04-004
Read the complete editionStory JSON

Summary

GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.

The authors call the mismatch between visually rich representations and control-useful structure the action-sufficiency gap. Their training constraints align geometry, predict instruction-relevant affordances and reconstruct goal regions inside a VLA policy and two world-action models. On zero-shot LIBERO-Plus transfer, the three variants gained 4.6, 12.6 and 5.2 points over matched counterparts; on RoboCasa the reported gains were 12.6, 9.0 and 8.4 points. Those results support the tested training principle, not a general guarantee for unseen robots.

Why it matters

GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.

Limits and context

  • Those results support the tested training principle, not a general guarantee for unseen robots.

Key claims

  1. GIFT supervises intermediate features for geometry, affordances and goal regions while leaving three different action formulations intact.

    Qualification: Those results support the tested training principle, not a general guarantee for unseen robots.

    Evidence: source-2026-09-04-004

Sources

  1. arXiv preprint 2609.04193arXiv · primary research

Corrections

No corrections have been recorded for this story.