robotics
Fifteen Hundred Hours Taught Two Arms to Keep Improving
XR-2 combines a large household-manipulation corpus with correction data gathered while the policy is actually running.

Summary
XR-2 combines a large household-manipulation corpus with correction data gathered while the policy is actually running.
The team releases 1,500 hours of two-arm demonstrations spanning everyday household tasks and uses the corpus to train a vision-language-action model called XR-2. Their experiments vary the amount of expert data, then add DAgger corrections from real-time human interventions when the deployed policy starts to drift. Success improves steadily along both tested scaling axes, supporting the claim that demonstration volume and on-policy correction remain useful at the dataset's current size. The authors also open-source the dataset for reproducible work. These are results from the paper's robot setup and task suite, not evidence that the policy can safely perform arbitrary domestic work.
Why it matters
XR-2 combines a large household-manipulation corpus with correction data gathered while the policy is actually running.
Limits and context
- Success improves steadily along both tested scaling axes, supporting the claim that demonstration volume and on-policy correction remain useful at the dataset's current size.
- These are results from the paper's robot setup and task suite, not evidence that the policy can safely perform arbitrary domestic work.
Key claims
XR-2 combines a large household-manipulation corpus with correction data gathered while the policy is actually running.
Qualification: Success improves steadily along both tested scaling axes, supporting the claim that demonstration volume and on-policy correction remain useful at the dataset's current size.
Evidence: source-2026-09-05-002
Sources
- arXiv preprint 2609.03591arXiv · primary research
Corrections
No corrections have been recorded for this story.