research
The Model Cropped the Image Without Looking
A causal audit found many visual-tool calls either irrelevant to the answer or informative but scheduled without a coherent plan.

Summary
A causal audit found many visual-tool calls either irrelevant to the answer or informative but scheduled without a coherent plan.
Researchers intervened at policy, trajectory and individual-step levels to ask whether crop-and-zoom observations causally changed multimodal-model answers. Across six models and five fine-grained perception benchmarks, they identify calls whose returned image had no causal effect and cases where useful evidence arrived through an incoherent call schedule. Aggregate gains were concentrated in a calibrated minority, making this a diagnosis of benchmark rollouts rather than proof that visual tools are never useful.
Why it matters
A causal audit found many visual-tool calls either irrelevant to the answer or informative but scheduled without a coherent plan.
Limits and context
No additional limitation was separately recorded.
Key claims
A causal audit found many visual-tool calls either irrelevant to the answer or informative but scheduled without a coherent plan.
Evidence: source-2026-08-07-008
Sources
- arXiv preprint 2608.06270arXiv · primary research
Corrections
No corrections have been recorded for this story.