TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

The Skill Had to Beat the Same Agent Without It

ACES measures capability packages through paired live trials under a fixed model, workspace, sandbox and scorer.

Published Updated Story ID: mp-2026-08-24-009
Read the complete editionStory JSON

Summary

ACES measures capability packages through paired live trials under a fixed model, workspace, sandbox and scorer.

Across 947 paired cases from 58 production skills and four harnesses, the preprint reports a mean composite Skill Lift of 0.2134 and positive lift in 72.8 percent of cases. Static scans and runtime outcomes were only weakly correlated, suggesting that document quality and actual agent benefit are complementary gates.

Why it matters

ACES measures capability packages through paired live trials under a fixed model, workspace, sandbox and scorer.

Limits and context

  • Static scans and runtime outcomes were only weakly correlated, suggesting that document quality and actual agent benefit are complementary gates.

Key claims

  1. ACES measures capability packages through paired live trials under a fixed model, workspace, sandbox and scorer.

    Qualification: Static scans and runtime outcomes were only weakly correlated, suggesting that document quality and actual agent benefit are complementary gates.

    Evidence: source-2026-08-24-009

Sources

  1. arXiv preprint 2608.20614arXiv · primary research

Corrections

No corrections have been recorded for this story.