TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Branching Beat Deeper Thought in Twelve of Fourteen Settings

A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.

Published Updated Story ID: mp-2026-08-26-027
Read the complete editionStory JSON

Summary

A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.

Across 49,327 graded items and 151,876 model calls, branching improved accuracy in all 14 model-benchmark settings by an average 5.98 percentage points and ranked best in 12. Growing one trace averaged 2.18 points and pruning/recomposition 0.94; paired scoring also showed how pipeline failures could reverse comparative conclusions.

Why it matters

A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.

    Evidence: source-2026-08-26-016

Sources

  1. arXiv preprint 2608.23956arXiv · primary research

Corrections

No corrections have been recorded for this story.