benchmarks evals
Bigger Models Did Not Follow Scientific Constraints Better
SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.
Summary
SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.
Experiments across closed and open multimodal models found large discipline gaps, with chemistry especially difficult. Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard. The authors say data and code will be released.
Why it matters
SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.
Limits and context
- Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard.
Key claims
SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.
Qualification: Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard.
Evidence: source-2026-08-27-016
Sources
- arXiv preprint 2608.25973arXiv · primary research
Corrections
No corrections have been recorded for this story.