developer tools
The GUI Agent Stopped Inventing Coordinates
A frozen multimodal model describes the target, then a separate matcher chooses among layout-aware candidates.
Summary
A frozen multimodal model describes the target, then a separate matcher chooses among layout-aware candidates.
The proposed regression-free pipeline separates semantic interpretation from precise localization. A frozen multimodal model expands an instruction into a structured visual description, while a grounding model matches that description against layout-prior candidates using only text-versus-icon labels. By avoiding direct coordinate generation, the authors aim to suppress hallucinated click locations without costly end-to-end fine-tuning. Reported benchmark gains remain preprint results and do not eliminate interface ambiguity or unsafe actions.
Why it matters
A frozen multimodal model describes the target, then a separate matcher chooses among layout-aware candidates.
Limits and context
- A frozen multimodal model expands an instruction into a structured visual description, while a grounding model matches that description against layout-prior candidates using only text-versus-icon labels.
- Reported benchmark gains remain preprint results and do not eliminate interface ambiguity or unsafe actions.
Key claims
A frozen multimodal model describes the target, then a separate matcher chooses among layout-aware candidates.
Qualification: A frozen multimodal model expands an instruction into a structured visual description, while a grounding model matches that description against layout-prior candidates using only text-versus-icon labels.
Evidence: source-2026-08-11-010
Sources
- arXiv preprint 2608.09654arXiv · primary research
Corrections
No corrections have been recorded for this story.