benchmarks evals
Branching Beat Deeper Thought in Twelve of Fourteen Settings
A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.
Summary
A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.
Across 49,327 graded items and 151,876 model calls, branching improved accuracy in all 14 model-benchmark settings by an average 5.98 percentage points and ranked best in 12. Growing one trace averaged 2.18 points and pruning/recomposition 0.94; paired scoring also showed how pipeline failures could reverse comparative conclusions.
Why it matters
A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.
Limits and context
No additional limitation was separately recorded.
Key claims
A shared harness compares three test-time recursion operators under identical prompts, budgets and grading.
Evidence: source-2026-08-26-016
Sources
- arXiv preprint 2608.23956arXiv · primary research
Corrections
No corrections have been recorded for this story.