robotics
A Mobile Manipulator Trained on More Than 5,000 Hours
MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.
Summary
MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.
MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams. The pretraining mix spans more than 5,000 hours, 400,000 episodes, 12 datasets and 17 embodiments. The paper reports 44.71% success on EBench, 61.2% on RoboCasa365 and an 83% mean across five real-world tasks. These figures come from different benchmarks with different protocols and should not be compared as one common score.
Why it matters
MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.
Limits and context
- MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams.
- These figures come from different benchmarks with different protocols and should not be compared as one common score.
Key claims
MM-ABC coordinated separate arm and base streams while future prediction strengthened the training signal.
Qualification: MM-ABC combines sparse multilevel vision-language features, a training-only future branch and masked joint attention between arm and base action streams.
Evidence: source-2026-09-29-010
Sources
- arXiv preprint 2609.35652arXiv · primary research
Corrections
No corrections have been recorded for this story.