developer tools
Execution Feedback Raised Scientific Coding Accuracy 9.9 Points
SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.
Summary
SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.
SciWalker contains 8,178 scientific coding problems spanning five domains and 32 subdomains, with solutions represented as operator graphs. The framework uses execution feedback to revise intermediate programs rather than grading only a final answer. In the reported experiment, Qwen3.5-9B accuracy rose from 29.3% to 39.2%; that gain applies to this benchmark and training recipe, not scientific software in general.
Why it matters
SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.
Limits and context
- The framework uses execution feedback to revise intermediate programs rather than grading only a final answer.
- In the reported experiment, Qwen3.5-9B accuracy rose from 29.3% to 39.2%; that gain applies to this benchmark and training recipe, not scientific software in general.
Key claims
SciWalker organized 8,178 problems across five domains into operator graphs that could be tested while solving.
Qualification: The framework uses execution feedback to revise intermediate programs rather than grading only a final answer.
Evidence: source-2026-09-27-010
Sources
- arXiv preprint 2609.30054arXiv · primary research
Corrections
No corrections have been recorded for this story.