benchmarks evals
Near-Perfect Robot Scores Hid Failure Recovery
LIBERO-RECOVER adds more than 1,000 scenarios spanning retries, adaptation, object repair and environmental recovery.
Summary
LIBERO-RECOVER adds more than 1,000 scenarios spanning retries, adaptation, object repair and environmental recovery.
Existing manipulation benchmarks can approach 100-percent success from clean initial states. LIBERO-RECOVER instead collects real execution failures and tests four levels of recovery plus spatial, structural, interaction and topological reasoning. It is a proposed benchmark, not evidence that evaluated systems already recover reliably.
Why it matters
LIBERO-RECOVER adds more than 1,000 scenarios spanning retries, adaptation, object repair and environmental recovery.
Limits and context
- It is a proposed benchmark, not evidence that evaluated systems already recover reliably.
Key claims
LIBERO-RECOVER adds more than 1,000 scenarios spanning retries, adaptation, object repair and environmental recovery.
Qualification: It is a proposed benchmark, not evidence that evaluated systems already recover reliably.
Evidence: source-2026-09-08-019
Sources
- arXiv preprint 2609.05178arXiv · primary research
Corrections
No corrections have been recorded for this story.