research
The Camera Saw the Blink. The Model Lost the Count
A controlled benchmark separated event frequency from visual complexity and found video-language models failing first on brief events, then on faithful timelines.

Summary
A controlled benchmark separated event frequency from visual complexity and found video-language models failing first on brief events, then on faithful timelines.
Researchers generated 2,190 controlled videos of bouncing-ball contacts, blinks and categorical state changes, each paired with an executable event trace. At an 80 percent reliability threshold, Gemini 3.6 Flash counted persistent state transitions up to 12 events at 0.5 and 1 hertz, yet showed no reliable positive-count region for transient blinks; in the high-count, high-frequency regime only 0.2 percent of final counts were correct. Sampling more frames improved bouncing-ball answer accuracy from 19.6 to 29.3 percent, but the reported event sequence matched ground truth only 3.7 percent of the time. These are author-reported preprint results from controlled and benchmark videos, not a universal audit of every video model or real-world deployment.
Why it matters
A controlled benchmark separated event frequency from visual complexity and found video-language models failing first on brief events, then on faithful timelines.
Limits and context
- At an 80 percent reliability threshold, Gemini 3.6 Flash counted persistent state transitions up to 12 events at 0.5 and 1 hertz, yet showed no reliable positive-count region for transient blinks; in the high-count, high-frequency regime only 0.2 percent of final counts were correct.
- Sampling more frames improved bouncing-ball answer accuracy from 19.6 to 29.3 percent, but the reported event sequence matched ground truth only 3.7 percent of the time.
- These are author-reported preprint results from controlled and benchmark videos, not a universal audit of every video model or real-world deployment.
Key claims
A controlled benchmark separated event frequency from visual complexity and found video-language models failing first on brief events, then on faithful timelines.
Qualification: At an 80 percent reliability threshold, Gemini 3.6 Flash counted persistent state transitions up to 12 events at 0.5 and 1 hertz, yet showed no reliable positive-count region for transient blinks; in the high-count, high-frequency regime only 0.2 percent of final counts were correct.
Evidence: source-2026-08-07-001
Sources
- arXiv preprint 2608.06361arXiv · primary research
Corrections
No corrections have been recorded for this story.