TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Perfect at Normal Speed Hid a Fragile Imitator

An imitation policy matched its expert at nominal pace, then lost far more insertions as the task accelerated.

Published Updated Story ID: mp-2026-09-02-005
Read the complete editionStory JSON

Summary

An imitation policy matched its expert at nominal pace, then lost far more insertions as the task accelerated.

In a parcel stowing task, both the scripted expert and an ACT learner achieved 100% success at nominal speed. At the fastest demonstrated pace, expert success fell to 84% while the learner reached 53%; 35 of the learner's 47 failures were insertion misalignments. The controlled comparison shows that nominal task success does not establish temporal robustness.

Why it matters

An imitation policy matched its expert at nominal pace, then lost far more insertions as the task accelerated.

Limits and context

  • The controlled comparison shows that nominal task success does not establish temporal robustness.

Key claims

  1. An imitation policy matched its expert at nominal pace, then lost far more insertions as the task accelerated.

    Qualification: The controlled comparison shows that nominal task success does not establish temporal robustness.

    Evidence: source-2026-09-02-005

Sources

  1. arXiv preprint 2609.01453arXiv · primary research

Corrections

No corrections have been recorded for this story.