CruxBench: A Benchmark of Information Discovery
This study addresses the challenge of evaluating "information discovery" in large language models (LLMs), specifically their capacity to identify pivotal sub-questions for decomposing complex tasks. To this end, it proposes an evaluation benchmark grounded in the Value of Information (VOI), which quantifies the extent to which posing questions updates predictive beliefs. This framework introduces a novel contamination-resistant, open-ended design that leverages future events to generate ground truth and compute VOI, thereby enabling scalable model assessment. The findings demonstrate that VOI correlates strongly with overall model capability. Furthermore, the analysis reveals that frontier LLMs perform only marginally better than random baselines on information discovery, underscoring the substantial difficulty of this task.