benchmarks evals
The Hardest Compositional Test Fell to 15.8%
TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.
Summary
TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.
TranSGrid tests whether models can recombine known pieces when the task requires deductive, inductive and abductive reasoning together. Seven Transformers performed much worse than on an ordinary held-out set: the largest solved 79.6% of that set, 55.3% of TranSGrid and 15.8% of its hardest subset. Making actions nearly linear or making goals action-explicit returned scores toward the simpler held-out level. The result argues that familiar systematic-generalization tasks can hide the reasoning demands they claim to measure.
Why it matters
TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.
Limits and context
No additional limitation was separately recorded.
Key claims
TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.
Evidence: source-2026-09-20-006
Sources
- arXiv preprint 2609.19212arXiv · primary research
Corrections
No corrections have been recorded for this story.