TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

A 512-Byte Message Beat the Reconstructed Grasp Geometry

The outcome-centered representation was 288 times smaller than one RGB-D frame and reached 0.876 AUC on 11,979 simulated grasps.

Published Updated Story ID: mp-2026-09-15-014
Read the complete editionStory JSON

Summary

The outcome-centered representation was 288 times smaller than one RGB-D frame and reached 0.876 AUC on 11,979 simulated grasps.

Networked robot arms often transmit dense geometry even when the action only needs a compact prediction of outcomes. This work learns an action-conditioned stochastic bottleneck that preserves outcome distributions rather than reconstructing the whole scene. Across 11,979 simulated grasps on 13 objects, its representation reached 0.876 AUC for lift success versus 0.542 for a reconstructed-geometry wrench score. The 512-byte interface was 288× smaller than one RGB-D frame and ran in 16 ms per CPU decision. Performance on unseen objects weakened before a feedback update, underscoring the remaining generalization gap.

Why it matters

The outcome-centered representation was 288 times smaller than one RGB-D frame and reached 0.876 AUC on 11,979 simulated grasps.

Limits and context

  • Networked robot arms often transmit dense geometry even when the action only needs a compact prediction of outcomes.

Key claims

  1. The outcome-centered representation was 288 times smaller than one RGB-D frame and reached 0.876 AUC on 11,979 simulated grasps.

    Qualification: Networked robot arms often transmit dense geometry even when the action only needs a compact prediction of outcomes.

    Evidence: source-2026-09-15-014

Sources

  1. arXiv preprint 2609.13235arXiv · primary research

Corrections

No corrections have been recorded for this story.