research
Training Tasks Learned to Disagree With the Solver
CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.
Summary
CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.
The system generated 5,431 executable tasks using multi-solver disagreement or strong-pass/weak-fail calibration. Models trained on the full set posted author-reported gains as large as 24.71 points on Terminal-Bench 2.0, 27.68 on SWE-Bench Pro and 30.04 on Doc2Repo.
Why it matters
CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.
Limits and context
No additional limitation was separately recorded.
Key claims
CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.
Evidence: source-2026-08-07-017
Sources
- arXiv preprint 2608.06352arXiv · primary research
Corrections
No corrections have been recorded for this story.