benchmarks evals
AI Exploration Sometimes Got Worse With More Searching
Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.
Summary
Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.
ExplorationBench creates AlienCode and AlienLogic environments whose executable rules conflict with familiar knowledge. Together they contain 55 discovery targets and 140 tasks, each with a flawed manual, environmental feedback and a tool schema. Across ten systems, the strongest could acquire and apply unfamiliar rules, but results varied substantially by trajectory and continued exploration sometimes stalled or reversed earlier gains. The benchmark tests synthetic worlds, not open-ended scientific discovery.
Why it matters
Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.
Limits and context
- The benchmark tests synthetic worlds, not open-ended scientific discovery.
Key claims
Two executable alien-world sandboxes separate learning unfamiliar rules from recalling familiar ones.
Qualification: The benchmark tests synthetic worlds, not open-ended scientific discovery.
Evidence: source-2026-09-25-016
Sources
- arXiv preprint 2609.30199arXiv · primary research
Corrections
No corrections have been recorded for this story.