robotics
The Robot Compressed Its History Before Deployment
A distilled workspace token replaced live vision-language queries on memory-intensive manipulation tasks.
Summary
A distilled workspace token replaced live vision-language queries on memory-intensive manipulation tasks.
Full observation histories can introduce spurious correlations, while repeatedly asking a vision-language model what matters adds deployment cost. This method uses the expensive model during training to identify task-relevant present and historical information, then distills that set into a lightweight workspace token with a reconstruction objective. In simulation and hardware, the token acted as a drop-in observation replacement for memory-intensive policies and removed in-loop VLM reasoning. The authors report not only lower deployment overhead but better policy performance than the compared history representations.
Why it matters
A distilled workspace token replaced live vision-language queries on memory-intensive manipulation tasks.
Limits and context
- The authors report not only lower deployment overhead but better policy performance than the compared history representations.
Key claims
A distilled workspace token replaced live vision-language queries on memory-intensive manipulation tasks.
Qualification: The authors report not only lower deployment overhead but better policy performance than the compared history representations.
Evidence: source-2026-09-18-005
Sources
- arXiv preprint 2609.20820arXiv · primary research
Corrections
No corrections have been recorded for this story.