TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Real Session Brought Its Mess With It

DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

Published Updated Story ID: mp-2026-08-30-011
Read the complete editionStory JSON

Summary

DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

The benchmark spans eight broad scenarios and 17 capability categories, with most tasks requiring several capabilities at once. Isolated containers add insufficient, unstable and noisy environmental conditions. Five agent frameworks paired with four models showed substantial gaps in strict completion, and both the model and harness shaped robustness. The source platform supplies the sessions and evaluation design, so the benchmark still needs broader replication across organizations.

Why it matters

DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

    Evidence: source-2026-08-30-011

Sources

  1. arXiv preprint 2608.26546arXiv · primary research

Corrections

No corrections have been recorded for this story.