robotics
Touch Needed a Faster Clock Than Vision
Agile-WAM predicted next-frame touch and slower visual change, reaching 11.9-millisecond inference.
Summary
Agile-WAM predicted next-frame touch and slower visual change, reaching 11.9-millisecond inference.
Adjacent camera frames often change gradually while tactile readings can jump at first contact. Agile-WAM encodes both streams into one latent space but supervises visual prediction at a longer offset and tactile prediction at the next frame, then generates future latents and action chunks together. Across nine simulated and five physical contact-rich tasks, the compact model beat the strongest reported baseline while keeping inference latency to 11.9 milliseconds. The five real-world experiments showed a 29.4% relative gain in overall success rate.
Why it matters
Agile-WAM predicted next-frame touch and slower visual change, reaching 11.9-millisecond inference.
Limits and context
No additional limitation was separately recorded.
Key claims
Agile-WAM predicted next-frame touch and slower visual change, reaching 11.9-millisecond inference.
Evidence: source-2026-09-18-007
Sources
- arXiv preprint 2609.20761arXiv · primary research
Corrections
No corrections have been recorded for this story.