TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Robot Judge Called Ambiguous Failures Success

The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.

Published Updated Story ID: mp-2026-09-05-007
Read the complete editionStory JSON

Summary

The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.

FailBench brings together 2,197 manipulation attempts from fourteen public sources, with three-quarters of failures occurring naturally. General-purpose vision-language models beat systems fine-tuned specifically for failure detection, while the strongest detector reached 0.77 mean balanced accuracy. Performance approached saturation when object motion made outcomes visible but dropped below 0.60 on contact-intensive assembly tasks; ambiguous evidence systematically biased predictions toward success. Cropping outcome-relevant regions improved the top detector by 2.4 points without retraining.

Why it matters

The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. The best of thirteen vision-language detectors reached 0.77 balanced accuracy and fell near chance on contact-heavy assembly.

    Evidence: source-2026-09-05-007

Sources

  1. arXiv preprint 2609.03611arXiv · primary research

Corrections

No corrections have been recorded for this story.