robotics
A Graph Checked the Robot’s Plan Before It Moved
GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.
Summary
GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.
GAVEL represents object relations, action preconditions, effects and beliefs about hidden object locations in an explicit graph world model. It predicts the consequences of language-model actions, repairs violations that follow from the graph and reserves language-model replanning for semantic errors. On BEHAVIOR-1K with Qwen3-8B, single-task success rose from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%; belief-aware reordering cut travel about 5.4% versus a static variant. These are embodied-simulation results, not household-robot deployment evidence.
Why it matters
GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.
Limits and context
- These are embodied-simulation results, not household-robot deployment evidence.
Key claims
GAVEL lifted one compact model from 41.2% to 91.8% success on long-horizon simulated tasks.
Qualification: These are embodied-simulation results, not household-robot deployment evidence.
Evidence: source-2026-09-20-010
Sources
- arXiv preprint 2609.19315arXiv · primary research
Corrections
No corrections have been recorded for this story.