robotics
The Robot Learned When One Subtask Was Actually Over
A small vision-language model distilled teacher explanations into real-time stage-transition decisions.
Summary
A small vision-language model distilled teacher explanations into real-time stage-transition decisions.
Long-horizon robot systems must decide when to stop one skill and start the next, but hand-built completion checkers are brittle and large cloud models are slow. StageGuard combines teacher-model reasoning with demonstrations to generate explanations of policy switching, then trains a lightweight vision-language model to emit compact self-explanations and transition decisions. Evaluation covered two trajectory benchmarks, BEHAVIOR-1K closed-loop control and real robots. The abstract reports substantial prediction gains without enough figures to support a broader quantitative claim.
Why it matters
A small vision-language model distilled teacher explanations into real-time stage-transition decisions.
Limits and context
No additional limitation was separately recorded.
Key claims
A small vision-language model distilled teacher explanations into real-time stage-transition decisions.
Evidence: source-2026-09-19-007
Sources
- arXiv preprint 2609.20791arXiv · primary research
Corrections
No corrections have been recorded for this story.