benchmarks evals
The Real Session Brought Its Mess With It
DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.
Summary
DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.
The benchmark spans eight broad scenarios and 17 capability categories, with most tasks requiring several capabilities at once. Isolated containers add insufficient, unstable and noisy environmental conditions. Five agent frameworks paired with four models showed substantial gaps in strict completion, and both the model and harness shaped robustness. The source platform supplies the sessions and evaluation design, so the benchmark still needs broader replication across organizations.
Why it matters
DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.
Limits and context
No additional limitation was separately recorded.
Key claims
DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.
Evidence: source-2026-08-30-011
Sources
- arXiv preprint 2608.26546arXiv · primary research
Corrections
No corrections have been recorded for this story.