robotics
Failed Rollouts Became a Steering Field
TraceFlow lifted ordered packing from 21 to 39 successes in 50 trials without changing policy weights.
Summary
TraceFlow lifted ordered packing from 21 to 39 successes in 50 trials without changing policy weights.
TraceFlow stores time-ordered robot states and actions with only a terminal success bit, then turns the densities of retrieved successes and failures into a bounded correction for a frozen flow-matching policy. On a physical ordered-packing task, completions rose from 21 of 50 to 39; one stacking round reached 47 and eliminated wrong-sequence episodes in that test. Simulation gains were selective rather than universal, with two task subsets declining slightly and stacking branches peaking before round ten. The result makes failure traces useful without claiming unlimited self-improvement.
Why it matters
TraceFlow lifted ordered packing from 21 to 39 successes in 50 trials without changing policy weights.
Limits and context
- TraceFlow stores time-ordered robot states and actions with only a terminal success bit, then turns the densities of retrieved successes and failures into a bounded correction for a frozen flow-matching policy.
Key claims
TraceFlow lifted ordered packing from 21 to 39 successes in 50 trials without changing policy weights.
Qualification: TraceFlow stores time-ordered robot states and actions with only a terminal success bit, then turns the densities of retrieved successes and failures into a bounded correction for a frozen flow-matching policy.
Evidence: source-2026-09-18-011
Sources
- arXiv preprint 2609.20646arXiv · primary research
Corrections
No corrections have been recorded for this story.