TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Bigger Models Did Not Follow Scientific Constraints Better

SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.

Published Updated Story ID: mp-2026-08-27-027
Read the complete editionStory JSON

Summary

SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.

Experiments across closed and open multimodal models found large discipline gaps, with chemistry especially difficult. Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard. The authors say data and code will be released.

Why it matters

SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.

Limits and context

  • Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard.

Key claims

  1. SciMIF tests ten groups of general and discipline-specific constraints across 22 tasks in five sciences.

    Qualification: Scaling the model did not reliably improve adherence, and fine-grained constraints requiring disciplinary application remained hard.

    Evidence: source-2026-08-27-016

Sources

  1. arXiv preprint 2608.25973arXiv · primary research

Corrections

No corrections have been recorded for this story.