developer tools
The Web Agent Predicted Differences, Not Just Pages
A matching objective trains predicted states to separate the true result of an action from the states produced by alternatives.

Summary
A matching objective trains predicted states to separate the true result of an action from the states produced by alternatives.
Most web-agent world models learn to reproduce the next HTML or accessibility-tree snapshot, even though the downstream ranker needs to tell candidate actions apart. The new objective uses branching WebArena Go-Browse trajectories with multiple actions and resulting states at each decision point. The authors report better predicted-state matching, action ranking on WebPRMBench and end-to-end success on WebArena-Lite than action-only or supervised-next-state comparisons.
Why it matters
A matching objective trains predicted states to separate the true result of an action from the states produced by alternatives.
Limits and context
- The authors report better predicted-state matching, action ranking on WebPRMBench and end-to-end success on WebArena-Lite than action-only or supervised-next-state comparisons.
Key claims
A matching objective trains predicted states to separate the true result of an action from the states produced by alternatives.
Qualification: The authors report better predicted-state matching, action ranking on WebPRMBench and end-to-end success on WebArena-Lite than action-only or supervised-next-state comparisons.
Evidence: source-2026-09-03-003
Sources
- arXiv preprint 2609.02885arXiv · primary research
Corrections
No corrections have been recorded for this story.