benchmarks evals
The Physics Changed. The Agent Kept Editing the Old Machine
PACE-Bench mutates a simulator after a code-driven design succeeds, forcing agents to redesign mechanisms rather than tune yesterday’s parameters.

Summary
PACE-Bench mutates a simulator after a code-driven design succeeds, forcing agents to redesign mechanisms rather than tune yesterday’s parameters.
Across 144 source-to-target pairs in six physics domains, the interface and goal stayed fixed while the target environment changed underneath the design. Ten self-evolving methods remained far from saturation: Reflexion with Qwen3-14B solved 35.9 percent overall, and GPT-5.5 solved 66.7 percent of the statics subset under the full budget. Simulator-grounded reflection beat unverified revision, while memory often anchored agents to obsolete designs.
Why it matters
PACE-Bench mutates a simulator after a code-driven design succeeds, forcing agents to redesign mechanisms rather than tune yesterday’s parameters.
Limits and context
No additional limitation was separately recorded.
Key claims
PACE-Bench mutates a simulator after a code-driven design succeeds, forcing agents to redesign mechanisms rather than tune yesterday’s parameters.
Evidence: source-2026-08-17-003
Sources
- arXiv preprint 2608.14441arXiv · primary research
Corrections
No corrections have been recorded for this story.