benchmarks evals
The Benchmark Hid the Question
EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.

Summary
EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.
EnigmaForge replaces the usual explicit question with a packet of documents whose facts imply both what must be solved and the one admissible answer. Its generator uses satisfiability checks to create puzzles with a unique solution and to verify that every clue is load-bearing. Across 25 frontier models, 600 instances and 17,400 evaluation records, the authors report a 22-fold spread in reconstructing the hidden problem, compared with a 1.6-fold spread in recovering stated facts. They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly. This is a synthetic benchmark preprint, but it isolates a practical failure mode: gathering facts can be much easier than discovering the question those facts were meant to answer.
Why it matters
EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.
Limits and context
- They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly.
Key claims
EnigmaForge asked models to infer both the puzzle and its answer from documents, producing a 22-fold spread in intuition.
Qualification: They also filter refusals separately so a model that declines the task is not confused with one that reasons incorrectly.
Evidence: source-2026-09-27-002
Sources
- arXiv preprint 2609.30144arXiv · primary research
Corrections
No corrections have been recorded for this story.