robotics
Driving Policies Improved With Three Orders Fewer Simulator Interactions
A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.

Summary
A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.
OPTED separates reinforcement learning from camera-policy post-training. A privileged teacher learns from vectorized maps and boxes, then supervises pretrained TransFuser and VaVAM students in closed loop on neural reconstructions of real driving logs. The reported driving scores increased 1.6 times and 9.5 times. In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.
Why it matters
A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.
Limits and context
- In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.
Key claims
A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.
Qualification: In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.
Evidence: source-2026-09-19-008
Sources
- arXiv preprint 2609.20756arXiv · primary research
Corrections
No corrections have been recorded for this story.