benchmarks evals
The Video Looked Real Until the Second Object Moved
Six leading generators scored near 0.8 on a familiar video benchmark, yet none cleared 0.42 when paired motions had to obey the same physical law.

Summary
Six leading generators scored near 0.8 on a familiar video benchmark, yet none cleared 0.42 when paired motions had to obey the same physical law.
Principia replaces camera-dependent absolute measurements with relationships between two objects in one controlled scene. Its eight tests span gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulums and mass-spring oscillation, allowing violations to be measured directly in image space without knowing the frame rate, object scale or camera calibration. Across thousands of generations from six systems, the authors report that no model exceeded 0.42 despite scores around 0.8 on VBench; vision-language models also struggled to identify the violations, with the best reaching 67% accuracy and most near chance. These are benchmark results reported in a new preprint, not an independent audit of every video model.
Why it matters
Six leading generators scored near 0.8 on a familiar video benchmark, yet none cleared 0.42 when paired motions had to obey the same physical law.
Limits and context
- These are benchmark results reported in a new preprint, not an independent audit of every video model.
Key claims
Six leading generators scored near 0.8 on a familiar video benchmark, yet none cleared 0.42 when paired motions had to obey the same physical law.
Qualification: These are benchmark results reported in a new preprint, not an independent audit of every video model.
Evidence: source-2026-09-04-001
Sources
- arXiv preprint 2609.04200arXiv · primary research
Corrections
No corrections have been recorded for this story.