benchmarks evals
The Driving Simulator Preserved What Policies Notice
A new metric ranked scenes by policy-relevant fidelity instead of appearance alone.
Summary
A new metric ranked scenes by policy-relevant fidelity instead of appearance alone.
DreamStream uses a simulator-grounded video model to vary visual appearance while preserving traffic layout and dynamic-object continuity for closed-loop driving tests. Its FDπ metric measures scene similarity through features used by public driving policies; under that metric, the system improved over the strongest evaluated simulator by 1.6 times on nuScenes and 4.7 times on NAVSIM. A new adversarial benchmark exposed scorer bias and weak recovery behavior. These are simulation and metric results, not evidence of safe road deployment.
Why it matters
A new metric ranked scenes by policy-relevant fidelity instead of appearance alone.
Limits and context
- These are simulation and metric results, not evidence of safe road deployment.
Key claims
A new metric ranked scenes by policy-relevant fidelity instead of appearance alone.
Qualification: These are simulation and metric results, not evidence of safe road deployment.
Evidence: source-2026-09-23-014
Sources
- arXiv preprint 2609.26792arXiv · primary research
Corrections
No corrections have been recorded for this story.