TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

A Mobile Manipulator Trained on More Than 5,000 Hours

MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.

Published Updated Story ID: mp-2026-09-29-010
Read the complete editionStory JSON

Summary

MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.

MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams. The pretraining mix spans more than 5,000 hours, 400,000 episodes, 12 datasets and 17 embodiments. The paper reports 44.71% success on EBench, 61.2% on RoboCasa365 and an 83% mean across five real-world tasks. These figures come from different benchmarks with different protocols and should not be compared as one common score.

Why it matters

MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.

Limits and context

  • MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams.
  • These figures come from different benchmarks with different protocols and should not be compared as one common score.

Key claims

  1. MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.

    Qualification: MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams.

    Evidence: source-2026-09-29-010

Sources

  1. arXiv preprint 2609.35652arXiv · primary research

Corrections

No corrections have been recorded for this story.