TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

Failure Detection Watched Relationships, Not Whole Scenes

RAFAIL reached 73.4% balanced accuracy across three real-world manipulation tasks without failure training data.

Published Updated Story ID: mp-2026-09-17-011
Read the complete editionStory JSON

Summary

RAFAIL reached 73.4% balanced accuracy across three real-world manipulation tasks without failure training data.

Whole-scene anomaly detectors can react to harmless visual changes, while runtime vision-language models add computation. RAFAIL learns representations of task-relevant relationships—such as gripper to object or object to target—from successful demonstrations annotated offline. At runtime, relationship-specific out-of-distribution detectors run without VLM inference or failure examples. Across three real-world manipulation tasks, the method reached 73.4% balanced accuracy and outperformed the strongest evaluated uncertainty and OOD baselines. The result supports narrowing failure detection to the geometry that matters for task progress.

Why it matters

RAFAIL reached 73.4% balanced accuracy across three real-world manipulation tasks without failure training data.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. RAFAIL reached 73.4% balanced accuracy across three real-world manipulation tasks without failure training data.

    Evidence: source-2026-09-17-011

Sources

  1. arXiv preprint 2609.18324arXiv · primary research

Corrections

No corrections have been recorded for this story.