robotics
The Robot Practiced the Future Before It Moved
Action-free endoscopic video raised a simulated surgical controller's average task success from 63.5 to 77.8 percent under a fixed demonstration budget.

Summary
Action-free endoscopic video raised a simulated surgical controller's average task success from 63.5 to 77.8 percent under a fixed demonstration budget.
Surgical WAM first learns visual dynamics from video that contains no robot-action labels, then fine-tunes a single generative model to predict both future endoscopic observations and executable action chunks. During control it executes only a short prefix before observing and planning again. Across four simulated manipulation tasks, the authors report an average success increase from 63.5 to 77.8 percent, including a 20-point gain on PegTransfer. The work is a simulation study, not evidence of clinical safety or deployment, but it tests whether abundant video can reduce dependence on scarce synchronized demonstrations.
Why it matters
Action-free endoscopic video raised a simulated surgical controller's average task success from 63.5 to 77.8 percent under a fixed demonstration budget.
Limits and context
- During control it executes only a short prefix before observing and planning again.
- The work is a simulation study, not evidence of clinical safety or deployment, but it tests whether abundant video can reduce dependence on scarce synchronized demonstrations.
Key claims
Action-free endoscopic video raised a simulated surgical controller's average task success from 63.5 to 77.8 percent under a fixed demonstration budget.
Qualification: During control it executes only a short prefix before observing and planning again.
Evidence: source-2026-08-12-001
Sources
- arXiv preprint 2608.11204arXiv · primary research
Corrections
No corrections have been recorded for this story.