TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

The Model Stopped Doing the Arithmetic Itself

A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.

Published Updated Story ID: mp-2026-09-12-003
Read the complete editionStory JSON

Summary

A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.

Program-Solve asks a clinical language model to write case-specific Python for a restricted local executor instead of calculating directly. On 1,100 cases covering 55 calculators, the 32B model rose from 83.47% to 90.53%, a paired gain of 7.05 points whose reported interval cleared zero. The 7B model’s 3.29-point increase was not statistically reliable. The study supplied formulas and gold variables, and its audit flagged version, use or coefficient concerns in 16 calculators, so deterministic execution does not replace verified formulas or extraction.

Why it matters

A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.

Limits and context

  • The 7B model’s 3.29-point increase was not statistically reliable.
  • The study supplied formulas and gold variables, and its audit flagged version, use or coefficient concerns in 16 calculators, so deterministic execution does not replace verified formulas or extraction.

Key claims

  1. A restricted Python solver improved one 32-billion-parameter clinical model by seven points, but did not reliably help the smaller one.

    Qualification: The 7B model’s 3.29-point increase was not statistically reliable.

    Evidence: source-2026-09-12-003

Sources

  1. arXiv preprint 2609.10728arXiv · primary research

Corrections

No corrections have been recorded for this story.