TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

One Vision Plan Drove Several Phone Actions

Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.

Published Updated Story ID: mp-2026-09-26-014
Read the complete editionStory JSON

Summary

Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.

Instead of asking a vision-language model to plan and ground every tap, Jev-Mobile uses infrequent VLM goals, the accessibility tree as an executable action space and a lightweight typed model for repeated local choices. On AndroidWorld it reached 79% task success, compared with 78% for SeeAct-V and 84% for the step-wise VLM baseline. The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.

Why it matters

Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.

Limits and context

  • The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.

Key claims

  1. Jev-Mobile cut successful-run time by 32.7% and model API cost by 73.4% against a step-wise VLM baseline.

    Qualification: The efficiency comparison includes only successful trajectories and depends on the tested mobile environment and serving prices.

    Evidence: source-2026-09-26-014

Sources

  1. arXiv preprint 2609.30186arXiv · primary research

Corrections

No corrections have been recorded for this story.