benchmarks evals
The Robot Judge Called Ambiguous Failures Success
The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.
Summary
The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.
FailBench brings together 2,197 manipulation attempts from fourteen public sources, with three-quarters of failures occurring naturally. General-purpose vision-language models beat systems fine-tuned specifically for failure detection, while the strongest detector reached 0.77 mean balanced accuracy. Performance approached saturation when object motion made outcomes visible but dropped below 0.60 on contact-intensive assembly tasks; ambiguous evidence systematically biased predictions toward success. Cropping outcome-relevant regions improved the top detector by 2.4 points without retraining.
Why it matters
The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.
Limits and context
No additional limitation was separately recorded.
Key claims
The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.
Evidence: source-2026-09-05-007
Sources
- arXiv preprint 2609.03611arXiv · primary research
Corrections
No corrections have been recorded for this story.