benchmarks evals
Many Models Couldn’t Switch the Meaning of ‘Not’
NAFBench changed formal negation rules while holding the natural-language problem steady.
Published Updated Story ID: mp-2026-09-24-017
Summary
NAFBench changed formal negation rules while holding the natural-language problem steady.
The strongest open models scored 59% to 74% across four specified semantics and remained sensitive to rule order. Two frontier models reached 100% on the fixed-complexity main set.
Why it matters
NAFBench changed formal negation rules while holding the natural-language problem steady.
Limits and context
No additional limitation was separately recorded.
Key claims
NAFBench changed formal negation rules while holding the natural-language problem steady.
Evidence: source-2026-09-24-019
Sources
- arXiv preprint 2609.27517arXiv · primary research
Corrections
No corrections have been recorded for this story.