frontier models
The Dense Video View Taught the Sparse One What Changed
A model used denser temporal sampling as its own teacher, improving state tracking without labels, a separate teacher or added inference cost.

Summary
A model used denser temporal sampling as its own teacher, improving state tracking without labels, a separate teacher or added inference cost.
S3T gives the same model a dense view of a clip as privileged training information and asks a sparse-view student to match its next-token distribution. On LLaVA-OneVision-2-8B, the authors report VSTAT gains from 1.74 to 2.70 points depending on configuration. Training on unlabeled synthetic clips also transferred to real video, adding 7.95 points on VSTAT-YouTube and 4.50 on MVBench Action Count. Those gains are benchmark results for the tested model, not proof of general video understanding.
Why it matters
A model used denser temporal sampling as its own teacher, improving state tracking without labels, a separate teacher or added inference cost.
Limits and context
- Those gains are benchmark results for the tested model, not proof of general video understanding.
Key claims
A model used denser temporal sampling as its own teacher, improving state tracking without labels, a separate teacher or added inference cost.
Qualification: Those gains are benchmark results for the tested model, not proof of general video understanding.
Evidence: source-2026-09-06-003
Sources
- arXiv preprint 2609.04203arXiv · primary research
Corrections
No corrections have been recorded for this story.