research
Ten Video Sources Became One Long-Horizon World
SolarWM unifies 1.43 million clips and adapts four 5B-to-33B video backbones behind shared training and inference interfaces.

Summary
SolarWM unifies 1.43 million clips and adapts four 5B-to-33B video backbones behind shared training and inference interfaces.
SolarWM converts clips from ten datasets into a frame-aligned record carrying observations, metric camera geometry, captions, quality metadata, selection decisions and provenance. A backbone-native layer then applies one three-stage recipe across four models based on Wan2.2, LTX-2.5 and MiniMax-H3 without erasing their native representations. The authors say causal models trained on five-second sequences can support interactive rollouts from minutes to hours, and they are releasing the data, pipeline, recipes, weights and framework; those are preprint claims, not independent replication.
Why it matters
SolarWM unifies 1.43 million clips and adapts four 5B-to-33B video backbones behind shared training and inference interfaces.
Limits and context
- The authors say causal models trained on five-second sequences can support interactive rollouts from minutes to hours, and they are releasing the data, pipeline, recipes, weights and framework; those are preprint claims, not independent replication.
Key claims
SolarWM unifies 1.43 million clips and adapts four 5B-to-33B video backbones behind shared training and inference interfaces.
Qualification: The authors say causal models trained on five-second sequences can support interactive rollouts from minutes to hours, and they are releasing the data, pipeline, recipes, weights and framework; those are preprint claims, not independent replication.
Evidence: source-2026-09-03-002
Sources
- arXiv preprint 2609.02886arXiv · primary research
Corrections
No corrections have been recorded for this story.