developer tools
One Vision Plan Drove Several Phone Actions
Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.
Summary
Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.
Instead of asking a vision-language model to plan and ground every tap, Jev-Mobile uses infrequent VLM goals, the accessibility tree as an executable action space and a lightweight typed model for repeated local choices. On AndroidWorld it reached 79% task success, compared with 78% for SeeAct-V and 84% for the step-wise VLM baseline. The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.
Why it matters
Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.
Limits and context
- The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.
Key claims
Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.
Qualification: The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.
Evidence: source-2026-09-26-014
Sources
- arXiv preprint 2609.30186arXiv · primary research
Corrections
No corrections have been recorded for this story.