TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

AI Exploration Sometimes Got Worse With More Searching

Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.

Published Updated Story ID: mp-2026-09-25-027
Read the complete editionStory JSON

Summary

Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.

ExplorationBench creates AlienCode and AlienLogic environments whose executable rules conflict with familiar knowledge. Together they contain 55 discovery targets and 140 tasks, each with a flawed manual, environmental feedback and a tool schema. Across ten systems, the strongest could acquire and apply unfamiliar rules, but results varied substantially by trajectory and continued exploration sometimes stalled or reversed earlier gains. The benchmark tests synthetic worlds, not open-ended scientific discovery.

Why it matters

Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.

Limits and context

  • The benchmark tests synthetic worlds, not open-ended scientific discovery.

Key claims

  1. Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.

    Qualification: The benchmark tests synthetic worlds, not open-ended scientific discovery.

    Evidence: source-2026-09-25-016

Sources

  1. arXiv preprint 2609.30199arXiv · primary research

Corrections

No corrections have been recorded for this story.