TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

A Graph Checked the Robot’s Plan Before It Moved

GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.

Published Updated Story ID: mp-2026-09-20-010
Read the complete editionStory JSON

Summary

GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.

GAVEL represents object relations, action preconditions, effects and beliefs about hidden object locations in an explicit graph world model. It predicts the consequences of language-model actions, repairs violations that follow from the graph and reserves language-model replanning for semantic errors. On BEHAVIOR-1K with Qwen3-8B, single-task success rose from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%; belief-aware reordering cut travel about 5.4% versus a static variant. These are embodied-simulation results, not household-robot deployment evidence.

Why it matters

GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.

Limits and context

  • These are embodied-simulation results, not household-robot deployment evidence.

Key claims

  1. GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.

    Qualification: These are embodied-simulation results, not household-robot deployment evidence.

    Evidence: source-2026-09-20-010

Sources

  1. arXiv preprint 2609.19315arXiv · primary research

Corrections

No corrections have been recorded for this story.