benchmarks evals
The P-Value Passed. The Assumption Did Not
P-Bench targets statistical mistakes that survive correctly executed code but invalidate an agent's scientific conclusion.
Summary
P-Bench targets statistical mistakes that survive correctly executed code but invalidate an agent's scientific conclusion.
P-Bench contains 425 open-ended hypothesis-testing tasks across economics, biology and medicine, requiring an agent to choose a method, compute a p-value and interpret the result. The authors say current agents often execute code correctly while violating assumptions behind the test. Their open-weight Fisher-R1-14B model, trained with verified statistical rewards, improved single-trial success by 21 percent relative to DeepSeek-V4-Pro on average. The benchmark and model are preprint contributions, not validation for unsupervised science.
Why it matters
P-Bench targets statistical mistakes that survive correctly executed code but invalidate an agent's scientific conclusion.
Limits and context
- The benchmark and model are preprint contributions, not validation for unsupervised science.
Key claims
P-Bench targets statistical mistakes that survive correctly executed code but invalidate an agent's scientific conclusion.
Qualification: The benchmark and model are preprint contributions, not validation for unsupervised science.
Evidence: source-2026-08-10-014
Sources
- arXiv preprint 2608.07437arXiv · primary research
Corrections
No corrections have been recorded for this story.