robotics
The Robot Chose Which Layer to Remember Mid-Action
LayerRoute dynamically mixes visual-language layers and rereads earlier action states, adding as little as 0.31 percent in one tested policy.
Summary
LayerRoute dynamically mixes visual-language layers and rereads earlier action states, adding as little as 0.31 percent in one tested policy.
Vision-language-action policies commonly expose fixed visual-language layers to each action layer and leave intermediate action states implicit. LayerRoute adds two interfaces: a router that mixes cached visual-language representations according to the current action state, and a reread path for earlier action representations. Across simulation and real-robot benchmarks, the authors report consistent gains for two base policies, including up to 7.2 points on LIBERO Long with 0.31 percent and 3.87 percent additional parameters in the respective systems. The results support adaptive representation access, but do not establish a universal routing recipe for every robot or task.
Why it matters
LayerRoute dynamically mixes visual-language layers and rereads earlier action states, adding as little as 0.31 percent in one tested policy.
Limits and context
- LayerRoute adds two interfaces: a router that mixes cached visual-language representations according to the current action state, and a reread path for earlier action representations.
- The results support adaptive representation access, but do not establish a universal routing recipe for every robot or task.
Key claims
LayerRoute dynamically mixes visual-language layers and rereads earlier action states, adding as little as 0.31 percent in one tested policy.
Qualification: LayerRoute adds two interfaces: a router that mixes cached visual-language representations according to the current action state, and a reread path for earlier action representations.
Evidence: source-2026-09-09-006
Sources
- arXiv preprint 2609.06079arXiv · primary research
Corrections
No corrections have been recorded for this story.