TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

Training Tasks Learned to Disagree With the Solver

CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.

Published Updated Story ID: mp-2026-08-07-015
Read the complete editionStory JSON

Summary

CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.

The system generated 5,431 executable tasks using multi-solver disagreement or strong-pass/weak-fail calibration. Models trained on the full set posted author-reported gains as large as 24.71 points on Terminal-Bench 2.0, 27.68 on SWE-Bench Pro and 30.04 on Doc2Repo.

Why it matters

CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. CalibForge revised terminal tasks until solver behavior placed them inside a learnable difficulty zone.

    Evidence: source-2026-08-07-017

Sources

  1. arXiv preprint 2608.06352arXiv · primary research

Corrections

No corrections have been recorded for this story.