TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Benchmark Hid the Question

EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.

Published Updated Story ID: mp-2026-09-27-002
Read the complete editionStory JSON

Summary

EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.

EnigmaForge replaces the usual explicit question with a packet of documents whose facts imply both what must be solved and the one admissible answer. Its generator uses satisfiability checks to create puzzles with a unique solution and to verify that every clue is load-bearing. Across 25 frontier models, 600 instances and 17,400 evaluation records, the authors report a 22-fold spread in reconstructing the hidden problem, compared with a 1.6-fold spread in recovering stated facts. They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly. This is a synthetic benchmark preprint, but it isolates a practical failure mode: gathering facts can be much easier than discovering the question those facts were meant to answer.

Why it matters

EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.

Limits and context

  • They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly.

Key claims

  1. EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.

    Qualification: They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly.

    Evidence: source-2026-09-27-002

Sources

  1. arXiv preprint 2609.30144arXiv · primary research

Corrections

No corrections have been recorded for this story.