TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

consumer products

The Agent Could Tap the App. It Could Not Hear the Ask

ElderBench collected 249 smartphone tasks from older adults across 20 apps—and found that indirect, ambiguous and underspecified instructions sharply reduced agent performance.

Published Updated Story ID: mp-2026-09-08-001
Read the complete editionStory JSON

Summary

ElderBench collected 249 smartphone tasks from older adults across 20 apps—and found that indirect, ambiguous and underspecified instructions sharply reduced agent performance.

ElderBench asks mobile agents to work from naturally elicited requests rather than the explicit, goal-shaped commands common in GUI benchmarks. The researchers collected 249 smartphone tasks from older adults across 20 applications, then characterized how their wording differed syntactically, semantically and pragmatically from existing test instructions. Mainstream GUI agents and vision-language models performed substantially worse on the older-adult-oriented requests in both online and offline settings. Controlled normalization and failure analysis tied part of that decline to indirect speech, referential ambiguity and missing details. The result is a benchmark finding, not a measure of any older person's ability: it identifies a design gap in agents that expect users to phrase needs like test cases.

Why it matters

ElderBench collected 249 smartphone tasks from older adults across 20 apps—and found that indirect, ambiguous and underspecified instructions sharply reduced agent performance.

Limits and context

  • The result is a benchmark finding, not a measure of any older person's ability: it identifies a design gap in agents that expect users to phrase needs like test cases.

Key claims

  1. ElderBench collected 249 smartphone tasks from older adults across 20 apps—and found that indirect, ambiguous and underspecified instructions sharply reduced agent performance.

    Qualification: The result is a benchmark finding, not a measure of any older person's ability: it identifies a design gap in agents that expect users to phrase needs like test cases.

    Evidence: source-2026-09-08-001

Sources

  1. arXiv preprint 2609.04850arXiv · primary research

Corrections

No corrections have been recorded for this story.