Skip to main content
Watch: I Asked Claude to Build Me a Business

Evaluation / Agents / Reasoning

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Max Pan · Tencent Hunyuan

arXiv:2609.3019912 upvotes

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Authors: Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Max Pan (Tencent Hunyuan)

arXiv ID: 2609.30199

Problem: Evaluating whether an AI system can do genuine scientific exploration is hard for two reasons: you have to verify that a genuinely new hypothesis holds, and you have to prove the system discovered it through exploration rather than recalling knowledge from pretraining. Most benchmarks only exercise recall on familiar problem shapes.

Key Methodology:

  • Builds "Alien Worlds" whose rules are executable (every answer is exactly checkable) and deliberately conflict with familiar knowledge, so pretraining recall cannot solve the tasks.
  • Two sandboxes: AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems explore the sandbox, then solve held-out tasks.
  • Evaluates 10 AI systems on the resulting framework.

Key Results:

  • The strongest systems can acquire and apply the unfamiliar rules from exploration, but performance varies substantially across trajectories.
  • Continued exploration can stall or reverse earlier gains - more exploration does not monotonically improve the discovered rule set.

Applied Context: ExplorationBench is the eval design that separates "can explore genuinely new rules from interaction" from "can recall related knowledge," which is exactly the distinction that matters for research agents, tool-using assistants, and any system that must learn the rules of workflow it has never seen. The verifiable-sandbox shape (executable rules, flawed manual, held-out tasks) is a reusable template for testing agents against local knowledge rather than world knowledge, and the stall/reversal finding is a direct caution for long-running agent loops: keep exploring until the held-out accuracy stops responding to new budget, not until the loop feels done.