TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

Execution Feedback Raised Scientific Coding Accuracy 9.9 Points

SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.

Published Updated Story ID: mp-2026-09-27-010
Read the complete editionStory JSON

Summary

SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.

SciWalker contains 8,178 scientific coding problems spanning five domains and 32 subdomains, with solutions represented as operator graphs. The framework uses execution feedback to revise intermediate programs rather than grading only a final answer. In the reported experiment, Qwen3.5-9B accuracy rose from 29.3% to 39.2%; that gain applies to this benchmark and training recipe, not scientific software in general.

Why it matters

SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.

Limits and context

  • The framework uses execution feedback to revise intermediate programs rather than grading only a final answer.
  • In the reported experiment, Qwen3.5-9B accuracy rose from 29.3% to 39.2%; that gain applies to this benchmark and training recipe, not scientific software in general.

Key claims

  1. SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.

    Qualification: The framework uses execution feedback to revise intermediate programs rather than grading only a final answer.

    Evidence: source-2026-09-27-010

Sources

  1. arXiv preprint 2609.30054arXiv · primary research

Corrections

No corrections have been recorded for this story.