TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Hardest Compositional Test Fell to 15.8%

TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.

Published Updated Story ID: mp-2026-09-20-006
Read the complete editionStory JSON

Summary

TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.

TranSGrid tests whether models can recombine known pieces when the task requires deductive, inductive and abductive reasoning together. Seven Transformers performed much worse than on an ordinary held-out set: the largest solved 79.6% of that set, 55.3% of TranSGrid and 15.8% of its hardest subset. Making actions nearly linear or making goals action-explicit returned scores toward the simpler held-out level. The result argues that familiar systematic-generalization tasks can hide the reasoning demands they claim to measure.

Why it matters

TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. TranSGrid combined deduction, induction and abduction instead of simplifying two of them away.

    Evidence: source-2026-09-20-006

Sources

  1. arXiv preprint 2609.19212arXiv · primary research

Corrections

No corrections have been recorded for this story.