safety
The Model Stopped Doing the Arithmetic Itself
A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.

Summary
A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.
Program-Solve asks a clinical language model to write case-specific Python for a restricted local executor instead of calculating directly. On 1,100 cases covering 55 calculators, the 32B model rose from 83.47% to 90.53%, a paired gain of 7.05 points whose reported interval cleared zero. The 7B model’s 3.29-point increase was not statistically reliable. The study supplied formulas and gold variables, and its audit flagged version, use or coefficient concerns in 16 calculators, so deterministic execution does not replace verified formulas or extraction.
Why it matters
A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.
Limits and context
- The 7B model’s 3.29-point increase was not statistically reliable.
- The study supplied formulas and gold variables, and its audit flagged version, use or coefficient concerns in 16 calculators, so deterministic execution does not replace verified formulas or extraction.
Key claims
A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.
Qualification: The 7B model’s 3.29-point increase was not statistically reliable.
Evidence: source-2026-09-12-003
Sources
- arXiv preprint 2609.10728arXiv · primary research
Corrections
No corrections have been recorded for this story.