robotics
One Action Step Replaced Ten—and the Robot Moved More Smoothly
IMLE-VLA raised inference from 15 to 55 hertz while reporting 98% average success across forty LIBERO tasks.
Summary
IMLE-VLA raised inference from 15 to 55 hertz while reporting 98% average success across forty LIBERO tasks.
Diffusion and flow-matching action heads often sample through several iterative steps, creating stop-and-go robot motion. IMLE-VLA replaces that loop with a single-step conditional generator trained to preserve multiple plausible actions. Applied to the reported π0.5 backbone, it raised inference frequency 3.67×, from 15 to 55 Hz, and achieved 98.0% average success across the 40-task LIBERO benchmark. Four real-world Franka tasks showed 2.2–3.0× lower jerk and 3.9–6.6× less VLA inference time per episode. These are the authors’ benchmark and hardware results, not a universal latency guarantee.
Why it matters
IMLE-VLA raised inference from 15 to 55 hertz while reporting 98% average success across forty LIBERO tasks.
Limits and context
- These are the authors’ benchmark and hardware results, not a universal latency guarantee.
Key claims
IMLE-VLA raised inference from 15 to 55 hertz while reporting 98% average success across forty LIBERO tasks.
Qualification: These are the authors’ benchmark and hardware results, not a universal latency guarantee.
Evidence: source-2026-09-13-009
Sources
- arXiv preprint 2609.10915arXiv · primary research
Corrections
No corrections have been recorded for this story.