robotics
Three Bridge Demonstrations Connected All 72 Tasks
AALT selected demonstrations for the new routes they enabled, not only the information they contained.

Summary
AALT selected demonstrations for the new routes they enabled, not only the information they contained.
Active imitation learning usually asks which demonstration would reveal the most about an expert policy. AALT instead values demonstrations that connect reusable behaviors across many start-goal tasks. In a simulated UR5e ordered-retrieval domain, its latent topology rose from 42 of 72 successful tasks to 72 of 72 after three demonstrations totaling five transitions beyond the initial set. After 20 demonstrations, the strongest baseline averaged 88.6% success with 98 transitions. The evidence is simulation-only, but it isolates compositional reachability as a useful acquisition objective.
Why it matters
AALT selected demonstrations for the new routes they enabled, not only the information they contained.
Limits and context
- The evidence is simulation-only, but it isolates compositional reachability as a useful acquisition objective.
Key claims
AALT selected demonstrations for the new routes they enabled, not only the information they contained.
Qualification: The evidence is simulation-only, but it isolates compositional reachability as a useful acquisition objective.
Evidence: source-2026-09-17-008
Sources
- arXiv preprint 2609.18004arXiv · primary research
Corrections
No corrections have been recorded for this story.