TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

Driving Policies Improved With Three Orders Fewer Simulator Interactions

A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.

Published Updated Story ID: mp-2026-09-19-008
Read the complete editionStory JSON

Summary

A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.

OPTED separates reinforcement learning from camera-policy post-training. A privileged teacher learns from vectorized maps and boxes, then supervises pretrained TransFuser and VaVAM students in closed loop on neural reconstructions of real driving logs. The reported driving scores increased 1.6 times and 9.5 times. In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.

Why it matters

A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.

Limits and context

  • In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.

Key claims

  1. A vector-input teacher lifted two camera-policy driving scores by 1.6× and 9.5× in AlpaSim.

    Qualification: In controlled experiments, the approach matched direct reinforcement-learning post-training with roughly one-thousandth as many simulator interactions while remaining closer to the human demonstration prior; this is simulation evidence, not public-road validation.

    Evidence: source-2026-09-19-008

Sources

  1. arXiv preprint 2609.20756arXiv · primary research

Corrections

No corrections have been recorded for this story.